# The Four Wars of the AI Stack - Dec 2023 Recap

Latent Space · 2024-01-26

<https://addtry.com/bce82bee-9298-4bfe-9f59-c3d381bb7e4d>

Swyx and Alessio recap December 2023 by framing the AI landscape as four wars: the Data War (NYT lawsuit demanding destruction of GPTs, OpenAI's data partnerships, synthetic data from DeepMind); the GPU/Inference War (Mixtral price dropping 90% to $0.27/million tokens, benchmark drama between Anyscale and Together, alternative architectures like Mamba); the Multimodality War (Midjourney reaching $200M ARR, ElevenLabs hitting unicorn status, OpenAI and Google building god models); and the RAG/Ops War (LangChain vs LlamaIndex, Pinecone's $750M valuation, Qdrant serving OpenAI and Anthropic internally). They argue agents and open source are not active wars, and predict 2024 will see AI finally reach production, with hardware like Rabbit R1 and Tab capturing unique context.

## Questions this episode answers

### What are the four wars of the AI stack according to the Latent Space podcast?

Hosts Alessio and Swyx identify the four wars as: the Data War over training data quality and attribution rights; the GPU/Inference War between those with abundant compute and those optimizing efficiency; the Multimodality War pitting generalist models like Gemini against specialized point solutions; and the RAG/Ops War involving database, framework, and LLM operations tools competing in the retrieval-augmented generation stack.

[1:34](https://addtry.com/bce82bee-9298-4bfe-9f59-c3d381bb7e4d?t=94000)

### Are AI inference providers losing money on Mixtral, and what is the break-even price?

Swyx analyzed Dylan Patel's data and calculated that serving Mixtral breaks even at around 51 cents per million output tokens, assuming 50% GPU utilization. Providers like Anyscale (50 cents), OctoAI (50 cents), Abacus AI (30 cents), and DeepInfra (27 cents) are pricing below cost, likely operating at a loss, while Perplexity's 56 cents is roughly at cost.

[26:57](https://addtry.com/bce82bee-9298-4bfe-9f59-c3d381bb7e4d?t=1617000)

### How much revenue does Midjourney generate and how many employees does it have?

Swyx noted that as of late 2023, Midjourney had reached at least $200 million in annual recurring revenue, bootstrapped, with a team of only 15 to 30 employees, highlighting its efficiency compared to venture-backed AI startups.

[42:25](https://addtry.com/bce82bee-9298-4bfe-9f59-c3d381bb7e4d?t=2545000)

## Key moments

- **[0:00] Intro**
  - [0:47] Swyx started the Latent Space monthly recap in August 2023, sorting AI news into thematic categories.
- **[1:48] Four Wars**
  - [1:48] The December 2023 recap introduced the 'Four Wars of the AI Stack': Data, GPU/Inference, Multimodality, and RAG/Ops.
  - [3:06] Swyx predicts no new vector DBs will emerge in 2024, except potentially serverless TurboPuffer.
- **[3:36] Selection Criteria**
  - [3:36] Swyx excluded 'Agents' and 'Open Source AI' as standalone wars because agents cooled down and open source lacks an opposing side.
  - [4:58] Code models haven't become a major battlefront yet because tool fragmentation slows experimentation, according to Alessio and Swyx.
  - [7:39] Alessio explains low-background tokens: pre-AI data is now as precious as radiation-free steel, per Chroma's Jeff Huber.
- **[10:10] Data War**
  - [11:49] The data war spectrum ranges from creators happy to be in training data to the NYT demanding all GPTs be destroyed.
  - [12:59] Swyx predicts the NYT v. OpenAI case will go to the Supreme Court and define fair use in AI.
  - [14:24] Alessio recalls Rap Genius proving Google scraping with diacritics in 2014, a precursor to today's data attribution fights.
  - [16:08] If NYT blocks its data, AI trainers will use other quality sources, diminishing NYT's exclusivity, Swyx argues.
  - [17:58] Synthetic data dominated NeurIPS discussions; it's the top theme as human-generated data becomes scarce.
  - [19:03] Andrej Karpathy at NeurIPS highlighted DeepMind's verifiable synthetic data for math and code, Swyx reports.
- **[21:41] GPU War**
  - [21:41] Mixtral's MoE release sparked the GPU/Inference War, with prices crashing 90% in one week.
  - [22:32] Anyscale's self-published Mixtral benchmark drew rebuke from Together and Meta's Sumith for testing demo endpoints.
  - [24:05] Alessio notes that AI benchmarks carry more weight than software benchmarks because end users cannot easily replicate results.
  - [25:17] Swyx emphasizes that inference cost isn't everything — latency, uptime, and throughput are critical for production apps.
  - [26:08] Swyx endorses Artificial Analysis, an independent benchmark site that pings production APIs for transparent metrics.
  - [26:57] Swyx estimates Mixtral break-even at 51 cents per million tokens; providers charging under that are losing money.
- **[41:36] Multimodality**
  - [42:12] Midjourney reached $200M–$300M ARR fully bootstrapped with only 15–30 employees.
  - [46:41] ElevenLabs is rumored to be a unicorn; Swyx admits he underestimated voice synthesis as a business.
  - [52:02] The Multimodality War pits OpenAI's and Google's 'God Models' against point solutions like Midjourney.
- **[53:57] RAG & Ops**
  - [54:40] LangChain's first commercial product is LangSmith, an ops tool, blurring lines between frameworks and ops.
  - [1:00:21] Swyx reveals that both OpenAI and Anthropic use Qdrant for internal RAG, suggesting it beat competitors in evaluations.
  - [1:02:42] Swyx questions what event would mark the true 'year of AI in production,' comparing it to 'Linux on the desktop'.
- **[1:05:18] Code & Agents**
  - [1:05:22] Alessio outlines 'semantics over syntax': AI lets non-coders express intent, but warns of codebase maintenance risks.
- **[1:11:56] Needle & Hardware**
  - [1:13:16] Swyx lists absurd prompt tricks: tip the model, take a deep breath, tell it your grandma is dying.
  - [1:15:11] Swyx says Gemini's marketing was dishonest, but it is a credible alternative to OpenAI, which is good for market diversity.
  - [1:17:34] Swyx argues that great consumer tech companies like Uber and Airbnb are initially provocative, and AI wearables follow that pattern.

## Speakers

- **Alessio** (host)
- **Swyx** (host)

## Topics

Image Generation, Inference, Open Source Tools

## Mentioned

AnyScale (company), Chroma (company), DeepInfra (company), ElevenLabs (company), Midjourney (company), New York Times (company), OpenAI (company), Perplexity (company), Pinecone (company), Qdrant (company), Replicate (company), Together (company), TurboPuffer (company), DALL-E (product), LangChain (product), LangSmith (product), LlamaIndex (product), Mixtral (product), Stable Diffusion (product), pgvector (product)

## Transcript

### Intro

**Alessio** [0:01]
Hey everyone. Welcome to the Late in Space Podcast. This is Alessio, partner and CTO in residence at Decibel Partners, and today I'm joined just by my co-host, Swyx, for a new podcast format.

**Swyx** [0:13]
Yeah. And, uh, it's a bit uncomfortable because we have to just stare into each other's eyes lovingly. Um, but, uh, in our end of year survey last year, a lot of listeners were asking us for more one-on-one time, more, uh, opinions from the both of us as hosts on what's going on in AI.

You know, both of us are very actively involved. And, uh, I, I think-- I don't think this year will be any different.

**Alessio** [0:35]
Mm-hmm.

**Swyx** [0:35]
This, this year there's lots more excitement to come, and I think, you know, um, uh, we're trying to grow Late in Space in terms of the types of formats and the, the amount of value that we deliver to our subscribers.

Um, so one thing that we've been trying and experimenting with is this monthly recap that I started doing around August of last year, where I basically just take the notable news items of the month, and then I sort them and categorize them according to some order that makes sense and, uh, write them down in the newsletter.

And, um, this last December recap was, uh, particularly exciting 'cause, uh, it seemed like it popped off in, uh, in a number of areas, particularly with the AI breakdown. Um, our friend, um, NLW featured it on his podcast, and I figured we can just kinda go over that as a way of setting the stage for 2024, but also recapping what happened in 2023.

**Alessio** [1:25]
Yeah. Um, and people always ask me if December is like a slow month. Uh, but I think-

**Swyx** [1:30]
No

**Alessio** [1:30]
... you almost, you almost broke Substack-

**Swyx** [1:32]
No

**Alessio** [1:32]
... with how many links we had in the-

**Swyx** [1:34]
Yes

**Alessio** [1:34]
... in the thing.

**Swyx** [1:34]
No, we actually did. So a lot of people, uh, talk- commented to me about the formatting issues, um, in, within the newsletter that I sent out, and I know that they are there, but I couldn't fix it because Substack was broken by us with how long it was.

**Alessio** [1:48]
Oh, man. Uh, but so y- we had this kind of like four main buckets called the, the, the four wars of the AI stack. Uh, data quality, and I guess like data quantity as well, in a way. Uh, the GPU rich versus poor, which we have a whole episode about with Dylan Patel.

### Four Wars

**Alessio** [2:06]
Uh, multimodality. Uh, we're actually recording tomorrow with, uh, Luma Labs about their new 3D models. So we went from text to image to 3D video. Uh, I wonder what's, what's next. And then-

**Swyx** [2:17]
And we're gonna release Hugging Face as well as-

**Alessio** [2:20]
Mm-hmm

**Swyx** [2:20]
... I guess I've been thinking about calling it Multimodality 101 because, uh, the first modality beyond text that you should really pay attention to is vision.

**Alessio** [2:27]
Right. Yeah. Um, yeah, and then the RAG ops war.

**Swyx** [2:31]
Yeah.

**Alessio** [2:32]
I think that's, uh, uh, maybe-

**Swyx** [2:33]
I don't know what to call it. I don't know what-

**Alessio** [2:35]
Yeah. No

**Swyx** [2:35]
... if you would've called it anything else. That is my-

**Alessio** [2:36]
The tooling war. I don't know. Um, but I, I think beginning of last year that was like kind of the hottest space because there wasn't much open source model work, and I think over the last maybe like four or five months, everybody's still focused on, uh, fine-tuning Llama 2 and like, uh, DPO to improve these models, Mixtral and all these things and people forgot about our friends at LangChain, Llama Index and some of the things that b- were maybe top of mind.

Vector DBs, you know, it seemed like everybody was releasing a Vector DB early in the year.

**Swyx** [3:06]
Yeah. I think that I, I'm, I'll be very surprised if any new Vector DBs come out this year. Uh, with one exception, which is something I'm keeping my, an, an eye on, which is TurboPuffer.

**Alessio** [3:17]
Mm-hmm.

**Swyx** [3:17]
I don't know if you've seen-

**Alessio** [3:18]
Yep

**Swyx** [3:18]
... uh, them going around. Um, yeah. Uh, all the smart people seem to be adopting TurboPuffer as the first serverless Vector DB.

**Alessio** [3:24]
Yeah.

**Swyx** [3:25]
Uh, which could be interesting.

**Alessio** [3:26]
Yeah, no, and we're gonna have definitely Jeff and Anton on the podcast at, at some point.

**Swyx** [3:32]
Yes.

**Alessio** [3:32]
Uh, I know they're gonna be, uh, they're gonna be fun, fun guests. But, um-

### Selection Criteria

**Swyx** [3:36]
I should also mention, um, I think it's interesting. So the, the reason I selected these four wars was a process of elimination of wars that I think didn't, ended up not mattering. So, uh, for the, for those who don't know, inside of my writing, I often include footnotes that are in, in themselves, um, just essays in footnotes.

Um, and so I think it's also notable the things that people thought were hot that were less hot than expected. Um, so it was agents, uh, definitely less hot, uh, than, uh, at, than at the start of 2023.

Um, and then this one is a very controversial s- non-selection by me, I think. Open source AI is-

**Alessio** [4:17]
Mm

**Swyx** [4:17]
... not a battle in, in the sense that I don't think there's anyone against open source AI. Everyone is, like, on one side. There's no, like, opposing side apart from regulators. But, um, in, in my mind, when I think about, like, for engineers, everyone, everyone, ev- en- engineers are all universally in favor of open source models.

So there's, there's no battle here as everyone just wants it to, to improve. So it's not, like, interesting-

**Alessio** [4:39]
Mm-hmm

**Swyx** [4:40]
... to write about. We just want o- more open source.

**Alessio** [4:42]
Yeah. The, the only battle is people offering inference on it and-

**Swyx** [4:46]
Yes. It is

**Alessio** [4:46]
... killing each other in the process.

**Swyx** [4:47]
Yes. And-

**Alessio** [4:48]
But the-

**Swyx** [4:48]
... yes, I, I cl- I classified that as a, a GPU rich versus poor war. Um, but maybe there's a, there's a better way to classify that, and you can, you can, um, give me some feedback on that because-

**Alessio** [4:58]
Mm-hmm

**Swyx** [4:58]
... um, I, it, it's a struggle to, uh, to, to try to ca- categorize the world. Um, code models as well. Um, I was very struck by, um, conversation I had with Poolside. I saw Kant from Poolside. Um, so they're, they haven't been on the podcast yet.

They, they, they're kinda stealth still, but they had a very, very notable fundraise. I think they had, like, $50 million raised. Uh-

**Alessio** [5:18]
I think even more, yeah

**Swyx** [5:20]
... for, for a seed. Uh, spending most of it on GPUs. And, uh, my conversation with Isael, he was like, "Hey, you know, like, Replit was, like, one of our podcast's early biggest winners?" Replit didn't really follow up with...

Like, they, they, they, they announced, they announced, like, their 1.5 model, but it's not really widely used, uh, beyond Replit, you know? Uh, there's StarCoder, there is CodeLlama.

**Alessio** [5:44]
Mm-hmm.

**Swyx** [5:44]
But, like, it's not really, for how important code is, it doesn't seem, like, as big of a battlefront as just general function calling, reasoning, um, you know, th- these, these other kinds of domains. Um, and so I thought it was j- just interesting to note that even though we have-- w- we, as a, as a podcast, try to pay particular attention to developer tooling, to code models.

We interviewed Cursor, Fine, Replit, Codium, Co- um, and, and Hugging Face. Um- Um, I, I, th-these all seem, like, very small, um, c-compared to the other... The, the amount of money being thrown, the amount of, uh, heat in, in the other domains.

**Alessio** [6:24]
Mm-hmm.

**Swyx** [6:24]
And I, I don't know why that is.

**Alessio** [6:25]
Yeah. I, I think it's maybe the fragmentation of the tooling. You know? Like, most people in code are using VS Code, Cursor, GitHub-

**Swyx** [6:35]
Yeah

**Alessio** [6:35]
... one of the three, so there's maybe not as much experimentation, versus with text, people are just trying everything. It's, it's hard to try a code model, you know? I see code models being released, but, like, it's not super easy to just plug it into your workflow.

So I think engineers like myself are just lazy, and- ... it's like, "Hey, I'm having great success with whatever I'm using."

**Swyx** [6:55]
Yeah, yeah.

**Alessio** [6:56]
"I don't really wanna, wanna go there." Um-

**Swyx** [6:58]
Um, special f- special case form of code is SQL, and, uh, the semantic layer data engineering type, type things. Uh, we also had two guests on there, uh, from Seek and Cube. Um, and we also talked to a bit of Databricks, a bit of Julius.

**Alessio** [7:12]
Yeah. And we have Brian from, from Hex.

**Swyx** [7:13]
And Brian. And Brian from Hex.

**Alessio** [7:14]
That's right. Yeah.

**Swyx** [7:15]
Does he count? I don't know.

**Alessio** [7:17]
Yeah, I don't know.

**Swyx** [7:17]
I guess... Yeah, yeah, yeah. I guess the Hex notebooks, yes. Uh, Hex Magic, yes. Um, Rexus is, is, uh, that's a different beast. Um, anyway, but yeah, I, I think, I think, um, um,

people who, like, come to AI engineering for the AI might actually end up finding themselves in data engineering in the end.

**Alessio** [7:38]
Mm-hmm.

**Swyx** [7:39]
In, in like, traditional ML engineering in the end. They might have to discover that they're, you know, doing Rexus. They, um, and, uh, all the, all the, the stuff that just gets swept u- swept under the rug in a demo, uh, becomes their job.

And, um, I think, I probably, I would probably say, like, you know, this, just because we didn't select a theme for last year, uh, doesn't mean it wasn't important. It just wasn't top of mind yet. And maybe I think that would be an emerging theme this year.

**Alessio** [8:04]
Mm-hmm. Yeah. I, I think that's kind of the consequence of the low-background tokens, like the end of the low-background tokens. Once we know-

**Swyx** [8:11]
Can you explain what, what you think are low back- low-background tokens? We, this is our November recap, by the way.

**Alessio** [8:15]
Yeah. It, it, well, the comparison that our friend Jeff Huber at Chroma, uh, brought up is steel before the atomic, uh, bomb creation. So steel before had no radiation in it. After all the testing, a lot of steel had radiation embedded in it, so it was really precious to get low-background steel, meaning with no radiation.

And same with tokens. You can assume that any internet content from three years ago, it's just internet. It doesn't have... It's, like, people writing. It's not models writing. Instead now, anything we're gonna get on Common Crawl updates and things like that, you never know if it's human written or not, and I think that will put more work on data engineering, right?

Because even basic stuff like, uh, checking if a text says, "As a model created by OpenAI," you know-

**Swyx** [9:02]
Mm-hmm

**Alessio** [9:02]
... um, it's gonna be important. So people had just been blindly taking all the data sets offered by Eleuther and Common Crawl and all these different things, assuming that all the data in it is good. I think now, how do you build on top of it?

And we've seen the New York Times lawsuit, uh, against OpenAI. We've seen, um, data partnerships starting to rise in different companies. Um, I think that's gonna be one of the bigger challenges, and maybe we'll see more of the, the work that Databricks has done to build the DALI 5K, uh, instruction tuning.

Just first-party creation of data. It's like you got people sitting at their desk every day. If everybody wrote five, you know, Q&A pairs or, or things like that, you would have a massive unique data set for, for your model.

So, um, yeah.

**Swyx** [9:51]
Yeah. For people who missed that episode, that was one of our early episodes as well, and, um, Mike Conover since left and to start Brightwave.

**Alessio** [9:58]
Yeah.

**Swyx** [9:58]
Uh, which, uh, I'm sure we'll have him back this year at some point.

**Alessio** [10:00]
Yeah, yeah. They're doing a lot of interesting stuff. I think the next episode will be, um, very cool.

**Swyx** [10:05]
Awesome. Uh, so how do you want to tackle this? Do you want to just kind of go through the four wars?

**Alessio** [10:10]
Yeah, let's do it. Um, you have, uh, you, you created this, like, Wikipedia-like- ... uh, infographic for, for each of them.

### Data War

**Swyx** [10:18]
So yeah. I should say, like, the, the inspiration for this actually was, um, s- during the Sam Altman, uh, leadership s- um, battle, uh, people were making mock Wikipedia entries for-

**Alessio** [10:32]
Mm-hmm

**Swyx** [10:32]
... the debate and for, like, who's on the side of the, you know, the DCLs and who was on the side of the EX. And, uh, so I, I like that format because it's very concise. It ha- it has the, it lists the key players, and it's, uh, kind of, uh, fun to think about, like, who's on what side and, um, think about what's is important and what people are, are battling over.

Um, I... and I think, um, it is, it is important to focus on key battlegrounds as a, as a concept because there's so many interesting things you could be talking about in AI, and they're not all equally interesting.

**Alessio** [11:01]
Mm-hmm.

**Swyx** [11:02]
So how do you decide what is interesting? Um, I think it's money. It's power. It's, uh, people. It's, um, you know, like, impact, that kind of stuff. Um, and so yeah. That, that's, that's what I ended up doing.

And, um, so fun fact, uh, the way I did this was I actually edited the HTML on Wikipedia, and then I just screenshotted it, uh, just to get the, the formatting.

**Alessio** [11:22]
Good old developer tools. Developer tools is all you need. Um, so the, the data war, belligerents.

**Swyx** [11:31]
Yeah.

**Alessio** [11:32]
On one side you have journalists, writers, artists. On the other side you have researchers, startups, um, synthetic data researchers. Um, I guess, like, maybe we wanna talk about what are the axes of war. So, like, one of them is attribution, right?

Like, I, I think there's a, a varying spectrum of how comfortable people are about this data going into a model. So some people are happy to have your model trained on it no, no matter what. Some people are happy to have your model trained on as long as you disclose that it's in the model.

Uh, some people just hate that, uh, you trained on their model, and some people, uh, like the New York Times, want you to destroy any artifact that it might, might have touched your article. So, um, that's kind of what we're fighting on.

It's not always... I just want to make it clear that it's not just like you should never use the data or you should always use the data. I think people are just trying to figure out What's the right form of attribution, and how do I get paid as somebody whose data ended up being in this training?

Uh, I think we're giving everybody a lot of great tokens with Lead in Space because we do full transcripts on, on everything- ... and, uh, we're happy for people to, to train models on-

**Swyx** [12:44]
Oh, yeah, yeah

**Alessio** [12:45]
... on us. Um-

**Swyx** [12:45]
Please train a Lead in Space model.

**Alessio** [12:47]
Yeah, we would love it. Um, so that's kinda what we're fighting on. Anything that people should keep in mind about this war and, like, maybe some of the campaigns, uh, that are going on?

**Swyx** [12:59]
Yeah. Um, so I, I, I think the New York Times one is probably going to go to Supreme Court. Um, it is very, very critical. It is, it is a, a landmark, um, war that will probably decide what fair use means in, in ter- in context of AI.

Um, so I think it's, uh-- And, and, um, I recommend, um, I think The Verge did a good analysis of this. Uh, Platformer maybe did a good analysis of this. Um, there are, like, four criteria for what fair use is, and everyone basically converges onto, um, the last criteria, which is does your use, does your transformative use of my copyrighted, uh, material, um, diminish the, the market for my content?

**Alessio** [13:41]
Mm-hmm.

**Swyx** [13:42]
Um, and, um, it's, it's very hard to say. Um, I, I, I, I would suspect that yes, in some capacity, in some amount, uh, but, uh, good luck proving that in a court of law.

**Alessio** [13:54]
Right.

**Swyx** [13:54]
And, and I think, um, a negative ruling on, on OpenAI would stall-- would seriously stall the progress of AI. Um, and, and that's bad for the pr- for humanity, but good for content creators and writers, so obviously, which we, we want, we want them to be adequately compensated and, and re-recognized for their work.

Um, so there's, like, no good... Th-there's, like, no easy outcome here apart from the existing copyright system, which is also, uh, somewhat broken.

**Alessio** [14:23]
Mm-hmm.

**Swyx** [14:24]
Um, and, uh, it, it's, it's just a very, very tricky, challenging, uh, case, uh, I think. Um, yeah, so-

**Alessio** [14:32]
It-

**Swyx** [14:32]
... uh, that's it

**Alessio** [14:32]
... it's funny because we had something... I was a community moderator at a website called Rap Genius, which was, um, lyric sanitation, and there was, like, a similar thing in maybe, like, 2014 where, like, the, the music labels basically came to the website and said, "Hey, this is not fair use."

You know? Like, you cannot reuse the lyrics to the song. And eventually, uh, the website made deals with the record labels to, like, be able to, to do this. And then Google was stealing the transcripts to put in, like the-

**Swyx** [15:01]
Yes

**Alessio** [15:01]
... enhance thing, um, and-

**Swyx** [15:03]
And they proved it by-

**Alessio** [15:04]
Yeah, yeah. We did all the, like, uh, basically, like, the things on the I. Some Is, we put the dots-

**Swyx** [15:09]
Diacritics

**Alessio** [15:09]
... some I, we put like, uh, the-

**Swyx** [15:10]
Yeah

**Alessio** [15:10]
... accent, and that's how it, it made it all better.

**Swyx** [15:12]
Oh, I thought it was, uh... I thought they, they just, uh, varied the spacing, or they, like, used a different kind of spacing in the Unicode, uh-

**Alessio** [15:18]
I think it was-

**Swyx** [15:19]
You know

**Alessio** [15:19]
... the I thing, but maybe I'm... I mean, this is, like-

**Swyx** [15:21]
Yeah

**Alessio** [15:21]
... almost 10 years ago.

**Swyx** [15:22]
So, so, so Rap Genius proved it by injecting some data poison into their con- their corpus, and then Google reproduced it faithfully, and so therefore they proved that-

**Alessio** [15:30]
Mm-hmm

**Swyx** [15:30]
... Google was scraping Rap Genius.

**Alessio** [15:31]
Yeah.

**Swyx** [15:32]
Uh, did, did Google have to pay Rap Genius money in the end?

**Alessio** [15:35]
Um, I don't think so. Uh, but at the same... Th-there was also another issue with Rap Genius that we had that it got blacklisted by Google for, like... There-

**Swyx** [15:44]
Of course

**Alessio** [15:44]
... there was, like, a lot going on.

**Swyx** [15:45]
Of course.

**Alessio** [15:46]
Um-

**Swyx** [15:46]
Of course

**Alessio** [15:46]
... but anyway, this is not-

**Swyx** [15:48]
Yeah

**Alessio** [15:48]
... a Rap Genius special.

**Swyx** [15:49]
Yeah. I mean, ultimately, like, I think that, um, we do need quality data. I think that then if it's-- if this case is contained to The New York Times, The New York Times' worst outcome is that they will substitute it with Washington Post, and they substitute with The Economist, or, like, the, the second or third-ranked newspaper that is the most friendly to AI.

And then The New York Times will realize that actually their words are not as-

**Alessio** [16:12]
Mm-hmm

**Swyx** [16:12]
... not that much more valuable than other words. Uh, and then, then, then the, the value of the content comes down very, very dramatically. Um, so I, I think it'll, I think it will be interesting. But, uh, yeah, I, I do, I do think it's overstepping their bounds to call for the destruction of all GPTs.

**Alessio** [16:26]
Yeah.

**Swyx** [16:27]
That's, that's probably for sure. Um, then the, the bigger problem I have is with Stack Overflow and Reddit, which I named as, um, on the, on the side of The New York Times. Um, they have shut down their- effectively shut down their APIs, um, in order to try to train their own models.

Um, s- probably same as, uh, Twitter, actually. I should probably have put Twitter. I put Twitter on the wrong side maybe. I don't know. Twitter is on both sides.

**Alessio** [16:49]
Elon is, is on every side.

**Swyx** [16:51]
So-

**Alessio** [16:51]
The side of chaos

**Swyx** [16:52]
... yeah. What, what this is, is basically every UGC, user-generated p- content company of the 20-- 2000 to 2010s now has a giant pile of u- of user content that, uh, that is bec- that becomes valuable data that used to be open for researchers to scrape and train models.

Now, all of them are locking their walls, right, behind their, their wall gardens, and then trying to train their own models to, to boost their, their benefits. So this is a locally optimal outcome for them, but a globally suboptimal outcome-

**Alessio** [17:21]
Mm-hmm

**Swyx** [17:21]
... for humanity, because w- why should we care about the, the closed garden of Reddit, um, you know, the Reddit model, the Stack Overflow model, uh, the X model, as opposed to, like, it being a part of a data mix of, you know, 20% Reddit, 20% Stack Overflow, 20% X?

Um, that seems like a much better outcome for, for the world, but everyone is acting in their very narrow self-interest in, in trying to make their own model, which is probably going to suck.

**Alessio** [17:49]
Right.

Um, so next war, after you get data-

**Swyx** [17:55]
Oh, uh, we, we should mention synthetic data.

**Alessio** [17:57]
Oh. Oh, yeah.

**Swyx** [17:58]
Yeah, yeah. So, so, um, what, what, what happens when you run out of human data? You make your own.

**Alessio** [18:03]
Right.

**Swyx** [18:06]
Uh, so I, I would say, like, that, that is... When I went to NeurIPS, that was the number one discussion out of every single researcher's mouth. Um, uh, there is a lot of, uh, research coming from, uh, both, I guess, the, uh, the big labs as well as the, um, the, the aca- academic labs, um, on, like, what good synthetic data looks like.

Um, I don't know if you've, like, talked to any startups around that. I just talked to Louis Castricado-

**Alessio** [18:29]
Yeah, yeah

**Swyx** [18:30]
... the other day. Um, and he is promising a very, very, um- .prompt interesting approach to synthetic data generation. Uh, like he-- I think his phrase for it is like pre-train scale synthetic data as opposed to what has-- what the news research and the other open source communities have been doing, which is fine-tune scale synthetic data.

**Alessio** [18:51]
Mm-hmm.

**Swyx** [18:51]
Um, and so, uh, he wants to create like trillion token data sets that are all synthetic. Um, and I'm like, "Okay, that's interesting," but also at the same time, these are all just downloads from GPT-4

**Alessio** [19:03]
Right

**Swyx** [19:03]
...uh, or something else. So, uh, Lewis is very, very aware of that, and he's-- he has a way around it. I don't really understand it, but he claims that that's, that's a, that's a good way around it. Um, Andrej Karpathy, um, at NeurIPS highlighted this paper from DeepMind where, um, they were, they were bootstrapping synthetic data that could be verifiably, uh, proven correct, um, so specifically in math and in code-

**Alessio** [19:29]
Mm-hmm

**Swyx** [19:30]
...uh, where there is a correct answer. So, um, yeah, that, that makes sense. You, you know, you can, you can solve the synthetic data problem that way. But like what about, you know, beyond that?

**Alessio** [19:38]
Right.

**Swyx** [19:38]
There's just no answer.

**Alessio** [19:39]
Yeah. And isn't-- wasn't part of the issue also that the way that the, the phrases are constructed and like all of that in synthetic data ends up kind of like making mode collapse even worse? Um, because-

**Swyx** [19:52]
Yeah. Yeah

**Alessio** [19:52]
...one thing is like right or wrong, right? The other thing is like every g- sample is written in the same way, you know? Or like as a similar-- since it comes from a certain model, it kind of has a similar root-

**Swyx** [20:03]
Yeah, so the-

**Alessio** [20:04]
...of structure

**Swyx** [20:05]
...you already have, um... Yeah, so I mentioned this in the, in the best papers, um, discussion with Jon Frankel. Um, s-so the, the basic argument is you already have a flawed distribu- flaw distribution from a da- from a language model.

You are resampling that flawed distribution to, to double down on, on that flawed distribution. Like there's no extra information from humans. So on principle, how can this work? Um, and so the, the only, the only conclusion there is, uh, you don't need it to emulate a human.

You need it to emulate a, a useful assistant, however you define it. So, uh, I think the, the, the, the goal of synthetic data is less to emulate human speech, because that is basically solved. It is now more to spike the distribution in, in useful ways.

And s- and that's a phrase I borrow from Kan Jin. Uh, but anyway, so I, I think that synthetic data will be a giant theme for this year, uh, and not least because, uh, the human data is being locked up behind walls.

Um, so like it's a very, very clear trend. This is probably the, the most amount of money after GPUs will be spent here-

**Alessio** [21:06]
Mm-hmm

**Swyx** [21:06]
...on data. So one war I did not put here was the talent war, right? Like the, the war for PhDs and smart people. And but when you break down what the talent people do, uh, one is they, they make models and they, they, they, you know, they run inference on GPUs.

Um, or, or they run, they, they, they run training runs on GPUs. But the other is they clean data. They find data, clean data, and format data. Um, and like, so yeah, these are all just proxies for like the kind of talent that is flowing back and forth.

Uh, and ultimately, I, I think you have to focus on like what they're being, what they're working on, the, like the, the visible output of what they're b- working on, which is, uh, data.

### GPU War

**Alessio** [21:41]
All right, let's talk about the GPU inference war. I think this is one that has been heating up.

**Swyx** [21:47]
Yeah.

**Alessio** [21:47]
And we actually have a bunch of these folks coming on the podcast in the next few days.

**Swyx** [21:52]
Yeah, yeah. Are we-

**Alessio** [21:52]
So-

**Swyx** [21:52]
...calling it Compute Month or-

**Alessio** [21:54]
Yeah, we, we gotta figure out a name-

**Swyx** [21:56]
...Inference Month?

**Alessio** [21:56]
...but we have Model, Together, Replicate. Um, there's a, there's a lot coming up. Um, but basically the, the Mixtral release, the MoE model was kind of the, the spark-

**Swyx** [22:07]
Yes

**Alessio** [22:07]
...of, of the war. Um, I think in, uh, the price went down like 90% in one week.

**Swyx** [22:13]
Yeah.

**Alessio** [22:13]
It started at like-

**Swyx** [22:14]
I saw it, I wrote 222 times. Uh, but yeah, one divided by 222 is whatever the, the-

**Alessio** [22:18]
Yeah

**Swyx** [22:19]
...yeah, the, the price is.

**Alessio** [22:20]
Um, yeah, and then there was like the benchmark drama between Together and, uh, Anyscale-

**Swyx** [22:25]
Yes

**Alessio** [22:25]
...on like whether or not which one was faster and like whether or not their-- the benchmark was really reflective of performance.

**Swyx** [22:31]
Yeah.

**Alessio** [22:32]
Um-

**Swyx** [22:32]
And, uh, this was very, uh, surprisingly ugly, uh, in a way that I think usually people try to respect each other's work and play nice and say nice things when people release stuff. Even if it's a competitor, you say nice things or you don't say anything at all.

**Alessio** [22:45]
Mm-hmm.

**Swyx** [22:47]
Um, Anyscale, for some reason, they released a benchmark that, uh, on which of course Anyscale looks the best. Why would you release a benchmark the way you don't look the best? But then, uh, basically everyone featured in that benchmark, uh, didn't like it, of course.

Um, I, I do think there's some like methodological things. So, so for anyone doing benchmarks, you have to understand that there's a real, real, real difference between like a public benchmark that is meant for just limited testing-

**Alessio** [23:14]
Mm-hmm

**Swyx** [23:15]
...compared to, okay, if you're load testing us, or if you're, if you're seeing what a real enterprise customer would see, you, you have to give, you have to give them a heads up. You have to get a different API key, a different endpoint, and you test the real infrastructure, not the, not the demo one.

Um, this is very common for infra- infra companies, and I think Anyscale just, uh, uh, neglected that. Um, and it, it hurt their credibility. Like the-- Anyscale is not new at this game. Like they should have done that.

Uh, but what was interesting was this benchmark drama reached even beyond Anyscale.

**Alessio** [23:42]
Mm-hmm.

**Swyx** [23:42]
And we're, we're gonna have Sumith on, and he's gonna talk about like why he weighed in. 'Cause Sumith doesn't represent any inference for it.

**Alessio** [23:47]
Right.

**Swyx** [23:47]
He just works at Meta. But, uh, he felt like this was, uh, this is a very interesting, um, uh, debate. And I think we'll see more of this. Uh, you are-- you have been a data investor for a while.

Like database companies always do this.

**Alessio** [24:00]
Mm-hmm.

**Swyx** [24:00]
And I think now we're just seeing this kind of fight come into the inference space.

**Alessio** [24:05]
Yeah. Yeah, and the-- uh, I think the hardest thing is the end customer cannot replicate it, you know? So like if you give me like a Postgres benchmark, I can run Postgres on my MacBook, you know, and run similar ones.

I think with models, it's just impossible. So people tell you, "This is the benchmark," and you're like, "Okay, I have to go sign up to every single cloud now to try it." Um, it's just not, uh, easy. And we talked about this in Benchmarks 101, which is same with model benchmarks, right?

Just like, "Oh, this model is so much better than this," and then it's like Did you train on the questions? I was like, "What? Um, oh, I don't know." Uh, so and again, it's harder for people to just like run the models and test them, you know?

So like the- there's a lot more weight, I think, in AI on benchmarks than there is in traditional software because nobody buys, um, Upstash over Redis Cloud or whatever just based on a benchmark.

**Swyx** [24:59]
Yeah.

**Alessio** [24:59]
They try them and check performance and whatnot because they have real production scale workloads. Here it's like nobody's really doing anything with these models, so it's like whatever AnyScale says I guess is good. Um, but then customers are gonna go try it and just decide for them what the, what the right thing is.

**Swyx** [25:17]
Yeah. Yeah. And, and I think the-- it's important to understand it is not just about cost. I think, um, what the price war represented was a race to the bottom on cost. And you're like, okay, uh, DeepInfra, it- which is a company, it's-- we're not...

You know, the, the name of the company is DeepInfra. DeepInfra has promised to just always be the lowest cost provider. Like, okay, fine. That- that's a, that's a good value proposition. But, um, when you... That you- you're not o- only optimizing for that in a production application, right?

You're optimizing for latency. That's one thing. You're optimizing for uptime. That's something that you can only earn-

**Alessio** [25:47]
Mm-hmm

**Swyx** [25:47]
... over time. You're optimizing for throughput, uh, and like other forms of reliability. It- it starts to tail off beyond that, but there's, uh, three or four dimensions that really, really matter, right? Like if you're, if you're not table stakes on any of those things, you're out.

You're just out. Um, so the... So, um, actually, there was a really good website that was released, uh, just this week, um, called Artificial Analysis.

**Alessio** [26:08]
Mm-hmm.

**Swyx** [26:08]
Did you see it?

**Alessio** [26:08]
Yep.

**Swyx** [26:09]
Yeah, so um, this is what the industry needs, which is an independent third-party benchmark pinging the production API endpoints of all the providers and giving a third-party analysis of, of, um, what this is. Um, I actually built a, a prototype of this last year.

**Alessio** [26:23]
Yeah, I was gonna say.

**Swyx** [26:24]
But, uh, I didn't like, I didn't like the main- mai- maintaining it.

Um, I- I'm glad someone else is doing it just because, uh, I don't want to keep up to, to-- with all, with all these things. Um, but still, like I, I think some- I think it's a public service that somebody should do and, uh, so I'm glad that they, they did it, and I think they did it very well.

**Alessio** [26:44]
Yeah.

**Swyx** [26:44]
Um, so, so yeah. The-- I think, um, that is where the, I guess, the inference drama is, uh, is ending for now. I don't think, uh, you know... I haven't seen any, uh, continuing debate there. Uh, the only other thing that...

You know, I did some extra work on this for the, for the recap, um, which is like, are they losing money? Um, you know, are they pricing correctly their, their tokens f- from Mixtral? And, um, I actually managed to go into Dylan Patel's-

**Alessio** [27:11]
Mm-hmm

**Swyx** [27:11]
... uh, writeup of the Mixtral price war. And I think I reasonably worked out that you can serve Mixtral and the, the, the lowest you can possibly charge if you like take the most aggressive amortization of all your, uh, CapEx and all that is fifty to seventy-five cents per million tokens, which is what Perplexity prices-

**Alessio** [27:31]
Yeah

**Swyx** [27:31]
... their Mixtral at. And Perplexity is a very smart player. Um, they're not even an inference-

**Alessio** [27:36]
Right.

**Swyx** [27:36]
... infra provider. They're just like doing this for fun. Uh, but they're like, "Yeah, this like... We don't, we don't want to lose money on this. We will provide it at cost. This is what cost is to us."

So that means, um... So M- Mi- Perplexity provides it at fifty-six cents per million output tokens. That means Anyscale, which is fifty cents, OctoAI fifty cents, Abac- Abacus AI thirty cents, and DeepInfra twenty-seven cents, they're all losing money, because we think that the break-even is fifty-one cents.

**Alessio** [28:05]
Yeah. And that's-- and that-- even that is like a full batch size and-

**Swyx** [28:09]
Yeah

**Alessio** [28:09]
... kinda max-

**Swyx** [28:10]
No, no, no. I, I assumed-

**Alessio** [28:11]
Max utilization of the-

**Swyx** [28:12]
Uh, I assumed fifty percent usa- utilization.

**Alessio** [28:13]
Okay. Okay

**Swyx** [28:14]
So like very... Like when you talk to practitioners, very, very good is sixty percent.

**Alessio** [28:18]
Mm-hmm.

**Swyx** [28:19]
Uh, average is like thirty, forty. So I just-- I say fifty, right? You assume fifty percent, batch size sixteen, uh, a hundred tokens per second generation. That's also very, very high. These are all very favorable numbers. Like, probably the real number is closer to seventy-five cents per million than fifty cents per million.

**Alessio** [28:34]
Mm-hmm.

**Swyx** [28:34]
Anyway, um, anyone charging under fifty, definitely losing money. So then it's like, okay, you-- if you-- either you don't know what you're doing, which in which case good luck, um, or you know what you're doing, and you're purposely losing money for something.

**Alessio** [28:50]
Mm-hmm.

**Swyx** [28:50]
And what is that? Um, and I don't know, but I think it's an interesting aggressive strategy to pursue if you are doing it on purpose. So this is something that like the, the classical like Walmart would have a loss leader.

**Alessio** [29:03]
Right.

**Swyx** [29:03]
Like they, they, they really, really on purpose lose money on things so that they get you in the door s- to try things out. I-- like I don't know if that makes sense to you as a, as a VC.

**Alessio** [29:13]
Yeah, yeah. It's like the... Well, it's like all the, uh, you know, the candies are placed at the cash register because maybe you just went to get the thing on discount, and then you buy a Kit Kat, whatever, and they make money on the Kit Kat.

Um, the-- your ki- They all have the Pokémon trading cards at checkout now, so if you bring your kids to buy the discounted whatever for you, then you end up spending more. But I don't-- to me, the thing is like what-- where's the checkout register where you upsell people with these things, right?

**Swyx** [29:41]
Yeah, I don't know how you upsell, yeah.

**Alessio** [29:42]
It's like that's, that's really the, the big thing. Um, yeah. I don't know. I'm curious to see. I don't think, uh, Cloudflare still has a life. I wonder what they're gonna charge, um-

**Swyx** [29:53]
Mixtral

**Alessio** [29:53]
... for their own workers. Yeah. They cannot-

**Swyx** [29:55]
They cannot serve Mixtral.

**Alessio** [29:56]
Why?

**Swyx** [29:56]
Their GPUs are too underpowered. Cloudflare AI is like very good marketing for very, very underpowered inference, right?

**Alessio** [30:06]
Yeah. Well, well, I don't know. I, I, I think it all depends on like what is gonna be needed, right?

**Swyx** [30:12]
Yeah.

**Alessio** [30:12]
So they have, they have Mixtral 7B right now. I, I checked.

**Swyx** [30:15]
Yes.

**Alessio** [30:15]
Um, but yeah, I wonder if Mixtral-

**Swyx** [30:16]
But they can, they can serve Mixtral.

**Alessio** [30:17]
Yeah, yeah, yeah.

**Swyx** [30:18]
Okay. Yeah, yeah.

**Alessio** [30:18]
Um, I wonder... But, but I think they don't wanna get into this race right now probably.

**Swyx** [30:22]
No.

**Alessio** [30:23]
You know?

**Swyx** [30:23]
Yeah.

**Alessio** [30:23]
So, um, yeah. I'm curious, going back to the loss leading, it's like, is there gonna be a better model that comes next that they hope that you're already integrated their thing with? You know, if you, if you're using Together to serve Mixtral, and then something else comes in that you're gonna replace Mixtral with, hopefully you're still gonna use Together, and they're gonna get better unit economics on it.

Um- I don't know.

**Swyx** [30:48]
Yeah.

**Alessio** [30:49]
It's a good question.

**Swyx** [30:49]
It's a good question.

**Alessio** [30:50]
Thank you, VCs, for, uh- ... you know, paying for all of our inference.

**Swyx** [30:54]
No, no, no, I, I like I think these are... You know, everyone in here are grown adults. They're, they're smart investors. I'm sure there's some kind of long-term strategy here, and I'm, I'm trying to figure that out. Like, assume that people are smart and then what, what-

**Alessio** [31:05]
Yeah

**Swyx** [31:05]
... would smart people do?

**Alessio** [31:06]
Yeah.

**Swyx** [31:06]
Uh.

**Alessio** [31:06]
I, I think it's the same with, uh, Uber, right? It's like-

**Swyx** [31:09]
Yeah, yeah

**Alessio** [31:10]
... how could it have been so cheaper at the start?

**Swyx** [31:12]
Yeah, yeah.

**Alessio** [31:12]
You know? Like, you look back at, at all, DoorDash-

**Swyx** [31:14]
Yeah

**Alessio** [31:14]
... all these things, it's like...

**Swyx** [31:16]
And, like, last year was a great year for Uber.

**Alessio** [31:18]
Mm-hmm. Yeah, no, exactly.

**Swyx** [31:18]
All my, all my Uber friends are, like, suddenly very, very rich again. Um, one thing I, I will mention on, like, the engineering sort of technical detail side is, um, you know, the, the, the rise of Mixture of Experts is something that, you know, was, uh, we covered in our podcast with George and, and, uh, now with Mixtral.

And, um, it re- represents, um, the first successful, really, really commercially successful sparse model. And sparse in a very interesting way, in a sense that the divergence between the amount of, um, compute you need at training versus the amount of compute you need for inference, uh, continues to diver- diverge.

But also in a weird way where you need to keep all the weights of the, the, the MoE model, uh, loaded, even though you're not necessarily using them at all times.

**Alessio** [32:10]
Mm-hmm.

**Swyx** [32:10]
Um, so what I... I mean, basically what I think that is, is like I, I think that that is going to impose different needs on hardware, different needs on, uh, workload, different needs on, like, batching optimization. Like, uh, Fireworks, um, recently announced, uh, Fire Attention, where they wrote a custom CUDA kernel for Mixtral, uh, on H100.

It's, like, super, super-

**Alessio** [32:31]
Mm-hmm

**Swyx** [32:32]
... uh, domain-specific. Um, and they, they announced that they could, for example, quantize, uh, from, like, 16 bit down to eight bit with, like, no loss in, uh, performance. Like, all these, all these, all these magical details emerge when you, uh, take advantage of, like, very, very custom, uh, optimizations like that.

Um, and, uh, I think, like, the rise in MoEs this year, uh, is going to be, uh, going to have very meaningful impacts on the inference market and how... It's gonna shape how we think and price for inference.

Um, it, I... It may not be that we have this sort of input token versus output token paradigm for, for long, uh, particularly because we have things like, you know, uh, different forms of batching, different forms of caching.

Um, and, like, I don't really know, uh, what that looks like, but I'm very, I'm very curious. I see a lot of opportunity here. If I was an inference provider, um, uh, player, like, that's something I would be trying to offer to people-

**Alessio** [33:24]
Mm-hmm

**Swyx** [33:24]
... as a, as a way to differentiate, because otherwise you're just an API.

**Alessio** [33:27]
Yeah. Yeah, no, it was ki- in a way counterintuitive because most of the struggles with inference as well are just, like, memory bandwidth, you know? So w- we have now models that scale worse at higher batch, you know?

**Swyx** [33:42]
Yeah.

**Alessio** [33:42]
Um, but I'm glad I'm not in that business, I can tell you that as for, uh... There's, there's so much work to be done at, like, so many low levels of the stack. You know, you're try- you're already trying to provide value to the customer on, like, the developer experience a- and all of that, but you also have to get so close to the bare metal to, like, make this model act- Like, like writing a kernel, imagine if you had to write...

You're, like, a CPU cloud provider, and you have to, like, write instruction sets. It's, like, just s- Nobody would get in that business, you know? So I, I salute all of our friends at compute providers doing this work.

And, I mean, Together is doing so much for, like, Tridal and, like, Flash Attention too and, and whatnot, so.

**Swyx** [34:23]
Yeah.

**Alessio** [34:24]
Um.

**Swyx** [34:24]
Yeah, so and, and that's something that I would, I would leave as l- the last part of this sort of war of GPU rich versus poor. Um, so there's... The GPU rich people are the model trainers and the, the infra providers.

They're, they're saying, like, "We have the GPUs. Come use our GPUs, um, you know, and then we, we provide you the best inference," right? Um, and that's, that's what we've, we've been discussing so far. On the other side, on the GPU poor side, um, are, like, all the alternative methods, right?

The modulars, the tinycorps, the QLORAs, um, and, and all, all the other type of stuff. I even put consistency models in there-

**Alessio** [34:56]
Mm-hmm

**Swyx** [34:56]
... because, you know, uh, any, any efficiency or distillation method where you go from, like, you, you reduce your inference or GPU usage by, like, 25 to 40 times is, uh, is a GPU poor-friendly approach.

**Alessio** [35:08]
Right.

**Swyx** [35:11]
Um, so I will also put Apple and MLX in there, and that's also... Like, Apple is finally making moves in, in, uh, inference, and that will be a, a game changer for local models because then you just don't need any cloud inference at all.

You just run it on device, which is fantastic. And then obviously our RWKV and Mamba and-

**Alessio** [35:27]
Mm-hmm

**Swyx** [35:27]
... Stri- Stripe Ty- Hyena from, um, Together. Like, all those, like, emerging models, I don't know... There's something I've, I've been worried about for a latent space. How much attention should we give to the emerging architectures? Because there's a very good chance that, one, these things don't work out.

**Alessio** [35:43]
Mm-hmm.

**Swyx** [35:44]
Two, they take a very long time to work out.

**Alessio** [35:46]
Mm-hmm.

**Swyx** [35:47]
Uh, and then three, um, once they work out, they're, like, for limited domains and, like, not super usable. Um, so I don't know if you have p- opinions on that. I, I, I can g- I can follow up with one conclusion that I've had, but I want to-

**Alessio** [36:00]
Yeah, no, I want to hear and then I'll-

**Swyx** [36:00]
... throw that, throw the question open to you. And so the one conclusion is, um, RWKV and the state-space models, including Mamba, have historically just been pitched as super long context models. And I'm like, "That's not something I need," because I'm okay with 100K context.

I'm okay with RAG and, and, you know, and, uh, recursive summarization, all those, all those techniques to extend your context, like RoPE and YaRN and-

**Alessio** [36:27]
Mm-hmm

**Swyx** [36:27]
... all these, all these things. So I'm like, why do I need million context, uh, models? Why do I need 10 million, 100 million, one billion models? Like, well, why? Um, so the, the more convinced... Like, the, the, the easiest argument is, oh, you can consume, like, very, very high bit rate things like video and DNA strands, and then you can do, like, syn- synbio and all that good stuff.

And I'm like, "Okay, I, I don't know anything about that."

**Alessio** [36:53]
Right.

**Swyx** [36:53]
Like, like what happens if, like, you hallucinate- Like one wrong chain in, in your-

**Alessio** [37:00]
Mm-hmm

**Swyx** [37:00]
... you know, the DNA strand that you're trying to, to synthesize, good luck. Um, I don't think-- I d- I don't know. Um, so all, uh-- like that's why I've been historically underweighting intentionally, uh, our coverage of state space models and, and the non-transformer alternatives, um, until Mamba.

Mamba really changed things where basically for the same amount of compute, um, you can get a lot more mileage or a lot more performance for the same size of model. And then it's a different-- N-Now it's an efficiency story.

Now it's a GPU port story. Uh, it is no longer a long context story. It is, it is just straight up we are strictly more efficient than transformers. I'm like, "Oh, okay, I can get that."

**Alessio** [37:38]
Mm-hmm.

**Swyx** [37:39]
Uh, does that change anything?

**Alessio** [37:41]
Yeah.

**Swyx** [37:41]
I don't know.

**Alessio** [37:43]
Yeah, yeah. No, that makes sense. I think the-- uh, people look at the slope, right? Which is like, oh, you can get the context higher and higher. But in reality it's like if you kept the context smaller instead look at the anti-slope, so to speak.

It's like same context. It's like a lot less compute, you know?

**Swyx** [37:57]
Yes.

**Alessio** [37:58]
Um, so-

**Swyx** [37:59]
Yeah. So, so that was not clear to me until Mamba. And, uh, so I think that's interesting. Um, I do think that, um, there's, there's something I've-- there's a concept I've been trying to call, uh, the, the sour lesson.

You know- ... the bitter lesson is, uh, stop trying to do domain-specific adjustments. Just scale things up and it's, it's gonna work.

**Alessio** [38:17]
Mm-hmm.

**Swyx** [38:17]
That's, that's general intelligence. General intelligence is, uh, uh, dislikes any attempt to, uh, i-imbue with- in-inside of it special intelligence. Like if you have like any if switch case or if statements or like if finance do this, if something do that, don't bother.

Just, just scale things up and it, it's gonna do all of them simultaneously all better at once. That's the bitter lesson. Um, the sour lesson is the, is a, is a parallel, is a corollary, which is, um, stop trying to model artificial intelligence like human intelligence, right?

Like we, we, we-- the neuron was inspired by the brain, but doesn't work exactly like the brain.

**Alessio** [38:56]
Mm-hmm.

**Swyx** [38:56]
Machine learning uses backpropagation. The brain does not use backpropagation. Um, and like so why should-- uh, we keep trying to, to create alternatives to transformers that are, that look like RNNs because we think that humans act like RNNs.

We have a hidden state, uh, and then we, we process new data and we, we update that state. Uh, but maybe ar-artificial intelligence or machine intelligence doesn't work like that. Uh, and maybe, like maybe we just fail every time we try.

**Alessio** [39:25]
Mm-hmm.

**Swyx** [39:27]
So that's the, the, the sour lesson. Every time we try to model things. A-and, uh, my favorite analogy, I actually got this from I think a, a, an old quote from Sam Altman, who was like, you know, like we made the plane, the airplane, uh, it was inspired by birds, but it doesn't work anything like birds, right?

It just is-- and it works very efficiently. Like it's probably the safest mode of transportation that we have. Uh, and it does-- works nothing like a bird. So why should artificial intelligence work like human intelligence? And, uh, th-that is the philosophical debate underlying my, uh, continued, um, cautiousness around state space models.

**Alessio** [40:01]
Mm-hmm. Right.

**Swyx** [40:02]
Which I don't know if it's, um, an, uh, the-- I, I feel very vulnerable saying this because I, I don't think there's any justification once you look at, um, the r- the empirical results or like the, the mathematical justifications for these things.

But there is some grounding in philosophy that you should have when you think about does an idea make sense? Does it-- is it worth exploring?

**Alessio** [40:24]
Yeah. Well, I, I think now there's a lot of work be-being put into it, right? And I think transformers have shown enough success that people are interested in finding the next thing, you know?

**Swyx** [40:37]
Yeah.

**Alessio** [40:38]
So before it wasn't clear if transformers were really gonna work, so people were kind of working on them. Uh, but yeah. Um, okay, maybe in the 2025, uh, recap we're gonna have-

**Swyx** [40:50]
Yeah. I mean-

**Alessio** [40:50]
... more

**Swyx** [40:50]
... we'll, we'll try to do one before that. Uh, so we actually have a, a link. I don't know if you know this, uh, Shreya Rajpal from Guardrails.

**Alessio** [40:56]
Mm-hmm.

**Swyx** [40:57]
She's married to Karan from, um-

**Alessio** [40:58]
From Hazy.

**Swyx** [41:00]
F- sorry?

**Alessio** [41:00]
He was at Hazy, right?

**Swyx** [41:02]
Yeah, from Hazy. Yeah.

**Alessio** [41:03]
Mm-hmm.

**Swyx** [41:03]
Uh, and so now he's one-- started one of the other state space model companies. I forget the name of it.

**Alessio** [41:07]
Mm.

**Swyx** [41:07]
Sorry, it's with a C. Um, and, uh, I, I'm sure, I'm sure like this will be an emerging topic this year as well. So we, we'll, we don't have to wait till next year.

**Alessio** [41:14]
Yeah, yeah. No, I, I think we're gonna have maybe the, the sour lesson- ... um, you know, overview. Uh, then-

**Swyx** [41:21]
Well, I mentioned this in a Luther Discord, and then they were like, "Okay, so what is the, the spicy lesson and what is the-"

**Alessio** [41:26]
The sweet lesson?

**Swyx** [41:27]
"... the salty, salty lesson and what is the sweet lesson?" Yeah, yeah.

**Alessio** [41:29]
I, I want the sweet lesson. Sounds better. Um, cool. Talking about GPU port, let's do multimodality, where I feel that Stable Diffusion was like the, the first GPU port model, you know?

### Multimodality

**Swyx** [41:42]
Yes.

**Alessio** [41:42]
Everybody was running it at home.

**Swyx** [41:43]
Yes, absolutely. I should-- I don't know if I mentioned it. I just didn't mention it. Stability, I think in 2023, um, you know, they, they shipped in incremental things. They, I think-- I don't know if Stable Diffusion 2 was out there.

Um, but e-everyone's talking about SDXL Turbo, which is-

**Alessio** [41:56]
Mm-hmm

**Swyx** [41:56]
... uh, a form which is an alternative to a consistency model, but looks like a consistency model. Uh, they, they shipped Video Diffusion. They shipped a whole bunch of stuff. Uh, but just wasn't as big as 2022 when they, when they, you know, made a huge impact with Stable Diffusion.

**Alessio** [42:09]
Yeah. Yeah, I mean, the-- it's hard to, it's hard to outdo-

**Swyx** [42:11]
It's hard to top that.

**Alessio** [42:12]
... Stable Diffusion. Um, but yeah, Midjourney has been doing great, obviously. I actually finally signed up for a paid account, uh, last month. So-

**Swyx** [42:20]
Uh, Midjourney? Yeah, yeah

**Alessio** [42:21]
... yeah, the, uh, I'm part of the $200 million a year-

**Swyx** [42:24]
You have to-

**Alessio** [42:25]
... that they're getting.

**Swyx** [42:25]
Uh, yeah. So n-now they, uh, it's, it's, uh, what's confirmed is I think like a Business Week article or Economist or Information article that, uh, yeah, this, uh, this team has now reached, uh, at least 200 million ARR, completely bootstrapped.

Um, I think the-- their employee count is somewhere between like 15 and 30 people. I, I don't know if you know-

**Alessio** [42:45]
Mm-hmm

**Swyx** [42:45]
... uh, exact numbers. Um, I have heard rumors, uh, that their revenue is actually higher than that, than was, than was, was reported. Um, but it's between the 200 million to 300 million range, which is crazy.

**Alessio** [42:57]
Yeah, yeah.

**Swyx** [42:59]
Especially if it's like primarily B2C.

**Alessio** [43:01]
Mm-hmm. Yeah.

**Swyx** [43:03]
Which it looks like it is.

**Alessio** [43:04]
Yeah, yeah. It's like B to Fiverr to B. I think, I think there's like a ton of B-

**Swyx** [43:10]
Oh, you mean just people on Fiverr-

**Alessio** [43:11]
You can see the-

**Swyx** [43:12]
Yeah, yeah, yeah. Midjourney specialists

**Alessio** [43:13]
... yeah, yeah. You can like get in Discord and see what people are generating, you know?

**Swyx** [43:17]
Yeah.

**Alessio** [43:17]
And you can see a lot of it is like product placement ads and a, a lot of stuff like that.

**Swyx** [43:22]
Yeah. And, and DALL-E 3 doesn't seem to have any impact on Midjourney.

**Alessio** [43:25]
DALL-E 3 got so much worse after the GPT-4-

**Swyx** [43:28]
Really?

**Alessio** [43:28]
... um, com- like the all-in-one. Well, first of all, before you could generate four images, and they had like very good vibes. Now the vibes are like boomer vibes.

**Swyx** [43:37]
Oh, no.

**Alessio** [43:37]
Every, every time it generates something-

**Swyx** [43:39]
The images I have here of DALL-E 3.

**Alessio** [43:41]
Every, every time it generates something on DALL-E, it looks some like some dusty, old... Yeah, like-

**Swyx** [43:49]
I think it's a-

**Alessio** [43:50]
... mid 2000s

**Swyx** [43:50]
... I think it's a skill issue. I think you've gone too wrong.

**Alessio** [43:52]
No, but that wa- that, that, that was the great thing about DALL-E 3, right? It's like it made the prompt better for you.

**Swyx** [43:58]
Yeah, yeah, yeah.

**Alessio** [43:58]
Like before, like literally like when it first came out, I'm like, "Hey, make, uh, a Colosseum with like llamas." And it was like this beautiful thing. I feel like now it's not. I don't know. Again, it's m- it's a model, right?

**Swyx** [44:10]
Good. So maybe it's-

**Alessio** [44:11]
So it's like maybe I just got unlucky.

**Swyx** [44:12]
Yeah.

**Alessio** [44:12]
I- I'm in the wrong latent space.

**Swyx** [44:14]
Yeah. Exactly, exactly, exactly.

**Alessio** [44:15]
You know? Um, but yeah. Um-

**Swyx** [44:17]
Yeah. There's, there's a lot of players in this. Um, I don't, I don't even think I put like some of the players I- I was really excited about. Like, you know, the Imagen team split out to be- to create Ideogram.

**Alessio** [44:26]
Mm-hmm.

**Swyx** [44:26]
You know, that was, that was a few months ago, and yet I didn't even put it here, uh, 'cause I forgot.

**Alessio** [44:30]
It's too much, right? You can't, you can't, can't keep track of all of it.

**Swyx** [44:34]
Yeah, yeah. Uh, okay, so I, I will just basically say that I, I do think that, um, I used to-- at the end of 2022, start of 2023, I was not as excited about multimodality. Uh, obviously I'm more excited about it now.

Um, I used to think that text-to-image was more like hobbyist kind of, you know, work, but $300 million a year is not a hobbyist.

**Alessio** [44:54]
Right.

**Swyx** [44:57]
It is not like, uh, not, you know, not just like not safe for work because Midjourney doesn't do not safe for work. So it's real. It's a new form of art. It's citizen art. It's, uh, uh, exciting. It's, it's un- unusual and interesting and, uh, a new ki- uh, uh, like and you can't, and you can't even model this as an investor.

You can't even model this on an existing market because like there's just a market of people who would typically not pay for art and now they pay a little bit for art, which is digital, not as good as a human, but it's good enough.

**Alessio** [45:29]
Mm-hmm.

**Swyx** [45:29]
I guess that's what I'm saying.

**Alessio** [45:30]
Yeah. I'm surprised I haven't seen a, a return of, uh, the digital frames that were very popular during the NFTs, uh, boom. People were like, "Oh."

**Swyx** [45:39]
Yeah. Yeah. So, so, uh, this is, um, the very, very first Late in Space post was on the difference between crypto and AI in this respect. So I called this, uh, multiverse versus metaverse.

**Alessio** [45:52]
Mm-hmm.

**Swyx** [45:52]
Crypto is very much about metaverse. Let us create digital scarcity and let us create, uh, tokens that are worth, uh, limited edition, that were something, um, and then you display it proudly in your, in your PFP as your, as your representation of yourself.

Um, and what AI represents is multiverse, which is a, is a very positive sum ins- instead of, instead of zero sum, where like if, if you like a thing, okay, I'll choose a different seed and I'll make a-

**Alessio** [46:17]
Mm-hmm

**Swyx** [46:17]
... a completely equivalent second thing, and that's mine. And, uh, that, that means very different things for like what value is and where value accrues. Uh, so like yeah, I, I, I mean, I still cling to that insight even though I don't know how to make money from it.

I think that, that-- I mean, obviously Midjourney figured it out. Uh, I think Mid- Midjourney like made the right approach there. Um, the other one I, I think I'll highlight is ElevenLabs. I think they were another big winner of last year.

Um, I don't know, did they announce their fundraise? I, I think s- I think so.

**Alessio** [46:46]
Uh, I don't, um-

**Swyx** [46:47]
Rumor is they're now-

**Alessio** [46:48]
Yeah, rumor is-

**Swyx** [46:48]
Rumor is-- I, I can say it. You don't have to say it because I, I only heard it from my friends. Uh, rumor is they're now a unicorn. Uh, and uh, and they just focus on voice synthesis, which again, not s- I did not care about it at the start of 2023.

Now we have used it for parts of Late in Space. Um, I listen, uh, almost every day to an ElevenLabs-generated podcast, the, the Hacker News Daily Recap podcast.

**Alessio** [47:11]
Mm-hmm.

**Swyx** [47:12]
Um, I don't know how-- I don't know what the, the s- the, the room for this to grow is because I always think like it's, it's so inefficient to talk to, to an AI, right? The, the, the bit rate of, of a voice-created thing is so low.

Um, it's only for asyn- asynchronous use cases.

**Alessio** [47:28]
Mm-hmm.

**Swyx** [47:28]
It's only for hands-free, eyes-free use cases. So like why, why would you invest in like, you know, voice, voice generation? I don't know. But like it seems like they're making money.

**Alessio** [47:37]
Right. Yeah, yeah. Yeah, no, I mean, Sarah, my, my wife, yeah, she uses it while she drives-

**Swyx** [47:42]
For-

**Alessio** [47:42]
... um, to talk to ChatGPT.

**Swyx** [47:43]
I see.

**Alessio** [47:44]
Just like-

**Swyx** [47:45]
Yeah. So, so, uh-

**Alessio** [47:45]
... she has questions

**Swyx** [47:46]
... ChatGPT uses their own TTS, and it's not-

**Alessio** [47:48]
Yeah. Yeah, yeah, yeah

**Swyx** [47:49]
... it's not OpenAI's. Okay.

**Alessio** [47:49]
Uh, but you, you can see the modality.

**Swyx** [47:51]
What does, uh-- we should, we should bring Sarah in at some point, but, uh, what does-

**Alessio** [47:55]
Customer interview

**Swyx** [47:56]
... what, what does-

**Alessio** [47:56]
But it's like we're doing a bunch of like, um-

**Swyx** [47:58]
... ChatGPT voice for? Yeah

**Alessio** [47:58]
... we're doing a bunch of like home renovation.

**Swyx** [48:00]
Yeah.

**Alessio** [48:00]
So maybe she's like driving to Home Depot-

**Swyx** [48:03]
Yeah

**Alessio** [48:03]
... and it's like, "Hey, um, what am I supposed to get to replace the sink?" You know, or-

**Swyx** [48:09]
Okay

**Alessio** [48:10]
... all these sort of things that-

**Swyx** [48:11]
Yeah

**Alessio** [48:11]
... maybe were like Google searches before.

**Swyx** [48:13]
Yeah.

**Alessio** [48:13]
Now you can easily do, um, eyes-free-

**Swyx** [48:17]
Yeah

**Alessio** [48:17]
... and hands-free.

**Swyx** [48:17]
Uh, yeah, a lot of people have told me about that, and I just, I-- it's-- I-- when I listen, when I, when I'm by myself, I always listen to podcasts. So I don't have time for ChatGPT, and ChatGPT, you know, the, the-- probably the number one thing they can do for me is give me like a speed, uh, adjustment so I can listen to it.

**Alessio** [48:34]
Yeah, yeah, yeah. Well...

**Swyx** [48:36]
Yeah.

**Alessio** [48:36]
That's funny.

**Swyx** [48:36]
Yeah. Anyway, so like, uh, I'm curious about your, your thoughts on like how an-- as an investor, I think this is the weirdest AI battlefront for, for investing 'cause you don't know the temp.

**Alessio** [48:48]
It's funny because the- there was, um, I'm trying to remember. There was a bunch of companies doing synthetic voices a while ago, and I think the problem, a lot of them got to like good ARR numbers, but the problem was like a repeatability, uh, use case.

**Swyx** [49:03]
Oh.

**Alessio** [49:03]
So people were doing all sort of random stuff, you know? And the problem is not-- it's kinda like Midjourney. The problem is not that there's not maybe a market of interest. It's like how do you build a venture-backed company with like a scalable go-to-market that like can go after a customer segment and like do repeatedly?

I think that's been the challenge. I don't know how ElevenLabs is doing it, but- You could do so many things with voice to-- text to voice that it's like, how do you sell it? You know? The-- who do you call?

Like, uh- Th-th-that's like the hardest thing, right? If you're raising like a series A, a series B, it's like how are you gonna invest this money in sales and marketing to get revenue back?

**Swyx** [49:40]
Yeah.

**Alessio** [49:40]
It's kinda like the basic of it, and-

**Swyx** [49:43]
Okay

**Alessio** [49:43]
... it can be challenging. That's why sometimes investors are like, "You're making money, and that's great for you," but like, how-

**Swyx** [49:50]
There's no industry.

**Alessio** [49:51]
Yeah. It, it's hard. It's hard to like just tie it together.

**Swyx** [49:55]
Okay.

**Alessio** [49:56]
You know?

**Swyx** [49:56]
Um, I would be interested in-- because I feel like there's a category of companies in the early 2010s that did this, meaning they offered an API with no idea how you were gonna use it. Uh, I'm thinking Twilio.

**Alessio** [50:10]
Mm-hmm.

**Swyx** [50:11]
Uh, a-and T-Twilio has a cohort of like sort of API-first companies that are all, all like sort of Twilio, uh, inspired. Um, but yeah. I th-I, the, the, I, I think there's a category or a time in the market when it makes sense to just offer APIs and just let your customers figure it out, and it's a-actually okay.

**Alessio** [50:32]
Yeah.

**Swyx** [50:32]
And then there's sometimes when it's not okay, and I think the default investor mentality right now is that it's not okay if you don't know what your customer is doing.

**Alessio** [50:39]
I think- Well, Twilio, Twilio is a funny example because I think in the middle 2010s, Uber was like fifteen percent-

**Swyx** [50:46]
Yeah, yeah, yeah

**Alessio** [50:46]
... of Twilio's revenue.

**Swyx** [50:47]
Yeah.

**Alessio** [50:48]
Um, I, I-

**Swyx** [50:48]
But like I'm just-- I'm talking like move yourself back as to like Twilio seed investor, Twilio series A investor. They had no idea.

**Alessio** [50:55]
Yeah.

**Swyx** [50:55]
Uber wasn't even around.

**Alessio** [50:56]
But, uh, but I think the, the thing now it's like, um, text to voice is not new.

**Swyx** [51:01]
Yeah.

**Alessio** [51:02]
You know? Like, that's really the thing.

**Swyx** [51:03]
Yeah.

**Alessio** [51:04]
It's like what's new now is that you can generate very good text to then feed into the model.

**Swyx** [51:09]
Yeah.

**Alessio** [51:09]
So that changes why the market is interesting, you know? But if you really think about it, the models today are a little better. They're maybe like fifty percent better than they were three years ago. But the transformer models under the feed, what to say, they're like a billion times better.

So imagine if you have like a, a lot of people use it for like automated, you know, customer support, things like that. Before you had like scripts they were reading. Now you have-- you can have a transformer model converse with the customer, so it makes it a lot more useful in cases.

Um, but we'll see how-

**Swyx** [51:44]
Yeah, we'll see

**Alessio** [51:45]
... that changes.

**Swyx** [51:46]
Um, okay. The, the last thing I'll, I'll mention here, why is this a, a war? Which is, um, OpenAI and Gemini, uh, and Google are working on everything models, uh, versus each of these individual startups all working on their selected modality.

And so, um, uh, this is a question of like are, are the big tech companies going to actually win because they can transfer learning across multiple domains, um, as opposed to each of these things being point solutions in, in their specific things.

The simple answer is obviously everyone will win.

**Alessio** [52:18]
Right.

**Swyx** [52:19]
Because the AI market is so huge. You know, there's a market for the, the b- the, the Amazon basics of like everyth-- you know, one model-

**Alessio** [52:26]
Mm-hmm

**Swyx** [52:26]
... has everything, and then there's a market for no, like, the basics are not good enough. I, I need the special thing. Uh, do you have an opinion on when does one market win o- win over the other, or is it just like everything, everything's gonna win?

**Alessio** [52:39]
Yeah. The-- It's interesting. I think, like, it works when people wouldn't have used the product without the Amazon basics, you know?

**Swyx** [52:46]
Okay.

**Alessio** [52:46]
So like may-maybe an example is like, uh, computer vision, you know?

**Swyx** [52:50]
Yeah.

**Alessio** [52:50]
Like, I mean, we have-

**Swyx** [52:51]
Yeah, vision is so here now

**Alessio** [52:52]
... RuboFlow and more. Yeah. It's like, you know, before people were like, "Why am I bothering trying out to set up a computer vision pipeline and all of that?" Now they can just go on GPT-4 and put an image, and it's like, "Oh, this is good.

I could use this for this," and then they build out something. And maybe they don't use OpenA- GPT-4V. They use RuboFlow or whatever else.

**Swyx** [53:11]
Yeah.

**Alessio** [53:11]
Um, that's kinda how I think about it. It's like what's the thing that enables people to try it, you know? So in a way, the, the God model can do everything fairly okay. It's like DALL-E and Midjourney, you know, all these different things.

Who's like the, uh-- and maybe like Mixtral, the Mixtral inference wars are like another example. It's like I would have never put something in my mo- in my app at like two dollars per million tokens, but I did it at twenty c- twenty-seven cents per million token.

You know? And now it's like, oh, no, I should really do this. It's a lot better.

**Swyx** [53:44]
Yeah.

**Alessio** [53:45]
So that's how I think about how the God model kinda helps the smaller people-

**Swyx** [53:50]
Yeah

**Alessio** [53:50]
... then build more business.

**Swyx** [53:51]
Yeah. Cool. Yeah. Creates a category.

**Alessio** [53:54]
Um.

**Swyx** [53:56]
Yeah, Dragonus.

**Alessio** [53:57]
Last, yeah, last but not least.

### RAG & Ops

**Swyx** [53:58]
Oof.

**Alessio** [53:59]
Where to begin? Um, we had a-all of-- almost all of these people on the podcast too. So.

**Swyx** [54:05]
They're, they're honestly the easiest to talk to because they're, they're, they, they, they look like dev tools.

**Alessio** [54:09]
Mm-hmm.

**Swyx** [54:10]
Um, and you are a dev tools investor. I, I worked in dev tools. Um, and, uh, they all-- I, I think they're also more mature, right? Uh, as, as, as, as businesses. Um, there, there-- there's more of a playbook that is well understood by the, by the customer.

**Alessio** [54:24]
Mm-hmm.

**Swyx** [54:24]
Like, like, yes, I need a new stack here. Uh, maybe not. Uh, so I think, um, the reason-- Okay, so my, my biggest problem with putting databases versus frameworks versus, uh, ops tooling in the same war is that they're not really a war.

They, they work cohesively together.

**Alessio** [54:39]
Mm-hmm.

**Swyx** [54:40]
Except when one thing starts to in-intrude on another thing. And that, that's why I put, uh, the very, very-- I very consciously put together this sequence, which is databases on the left, frameworks in the middle, ops companies on the right.

What does-- what's the first product of LangChain? LangSmith, which is an ops thing.

**Alessio** [54:57]
Mm-hmm.

**Swyx** [54:57]
Right? So, so now suddenly the framework companies are not so friendly with the ops companies, 'cause they are co- trying to compete with the ops companies.

**Alessio** [55:03]
Yep.

**Swyx** [55:03]
And what are ops companies trying to do? The ops companies are trying to produce SDKs that compete with frameworks. Okay, then what are the database companies trying to do? First of all, they're fighting with between each other, right?

There's the non-databases all adding vector features. We had some, uh, people approach us, and we, we had to say no to them 'cause there's just too many. And then there's the vector databases coming up and getting two hundred and thirty-five million dollars to, to, to, to, to build vector databases.

Uh, maybe... Okay, I'll just, uh... You know, obviously you're an active investor in some of these things, so you cannot say everything. But, uh, what are-- Uh, just on databases alone, one of the biggest debates of 2023, uh, where do you stand on the, the whole thing?

**Alessio** [55:43]
That's, uh, that's the million-dollar question. I think it's really-- Well, one, in the start, everything, um... There's kinda like a lot of hype, you know? So like when LangChain came out and LlamaIndex came out, then people were like, "Oh, I need a vector database."

It's like vector... They search vector database, and it's like Chroma, uh, Pinecone, whatever. But then it's like, oh, you can actually just have pgvector in Postgres, and you already have Postgres. Uh, did you know it could do that?

And people are like, "No, I didn't," because nobody really cared.

**Swyx** [56:13]
Yeah.

**Alessio** [56:13]
So like there's not a lot of documentation. Same with, uh, yeah, MongoDB vector, Cassandra, all these things.

**Swyx** [56:19]
Redis, Elasticsearch.

**Alessio** [56:21]
It-- You can actually put vectors and embeddings in, in everything.

**Swyx** [56:25]
It's a different kind of index. Yeah.

**Alessio** [56:26]
You know?

**Swyx** [56:26]
Yeah.

**Alessio** [56:26]
And I think like, I mean, like, Geoff and Anton also, what they always talked about, even early on, is like this is like a active learning platform. This is not just like a vector database. It's like, what do you do with the vectors?

It's like, what's most helpful? It's not where do you store them. So that's kinda the, the change in mentality.

**Swyx** [56:45]
But I think that was old Chroma, by the way. I don't, I don't know if that's the new mes- the current messaging.

**Alessio** [56:49]
Well, uh, but, but I think- ... I, I, I'm just saying like, to them, it's never about this is the best way to put a vector somewhere. It's like, this is the best way to operate on the vectors, and the store is like part of it.

But there's like the pipeline to get things out and everything. You have to build out a lot more. So I think 2023 was like create the data store. I think 2024 is gonna be like, how do I make the data store useful?

Because the vector storage has come out of taste, so there needs to be something else on top of it. Um-

**Swyx** [57:20]
Mm.

**Alessio** [57:21]
Yeah.

**Swyx** [57:22]
Un-unless they can come out with some kind of, like, new distance function or something. I, I keep waiting for Chroma to... Uh, they, they teased a little bit of what they're working on at the AI Engineer Summit, um, which, um, yeah, density and, uh, whatever other fancy formulas Anton is cooking up.

Um, but yeah. The, the-- My, um... So I, I, I think I tweeted about this maybe, like two, three months ago, and I think I pissed off Chroma a little bit. But I, I... The, the best framing of what Anton is, is-- would respond here is the, is what people are embedding within vectors is a very different kind of data from what is already within Postgres and MongoDB and all the others.

Um, so in, in, in some sense, it's net new data.

**Alessio** [57:59]
Hmm.

**Swyx** [58:00]
Um, and that actually struck a chord with me because that's how I started to understand structured versus unstructured data. That's how I started to understand, um... You know, one of my kind of heroes is, um, Mark, who's, who's CTO of MongoDB.

He-- This guy was the former, uh, GM of AWS RDS. Is it-- And for those who don't know, GM is, is like you're the C- you're the mini CEO of that business. And when you work at AWS RDS, you run a-

**Alessio** [58:27]
Mm-hmm

**Swyx** [58:28]
... one, $2 billion a year business. Um, and now-- And then he quits being Mr. Postgres of AWS to join MongoDB, the, the enemy. And, and, a-and he-- and his, uh... When he did-- gave that speech of, like, why he did this, um, he was like, "Actually, if you look at the, the kind of workloads that there's-- that's happening," Postgres is doing well, obviously.

Structured data al-always gonna be there. But unstructured data and, uh, document-type data is just rising exponential rate even faster. And like for him to say that, uh, means different things. Anybody could have said that. Anybody could have pointed-- made a chart that, that showed what he did.

Uh, anybody could have said that. But for him to have said that, I think it was a very big deal.

**Alessio** [59:09]
Mm-hmm.

**Swyx** [59:09]
'Cause he, he's rich. He doesn't have to work, but he like believed in this so much that he was like, "Okay, I'll, I'll just join MongoDB."

**Alessio** [59:15]
Yeah.

**Swyx** [59:16]
So I'm like, okay. There's, there's a real category shift between structured data, unstructured data. I believe it. I don't, I don't think it's just that you can put JSONB inside of Postgres and be done. That's not a, that's not a NoSQL database.

Okay, fine. Um, so what is this new thing o-of vectors? A-and, and, and how do you think about that as, as a new kind of data? And I, I think if there's a third category of, uh, something beyond unstructured data, I don't know what it is.

I, I, I, I, I... Like context or memory or whatever you call it.

**Alessio** [59:46]
Mm-hmm.

**Swyx** [59:46]
Whatever you call this kind of new data, that might belong in a new, new category of database, and that might create the new MongoDB of this era, and it could be any one of these guys. It, um... Right now, um, you know, Pinecone has, has the lead.

Uh, I think they're a $750 million company.

**Alessio** [1:00:03]
Mm-hmm. Valuation, yeah.

**Swyx** [1:00:04]
Yeah. Um, and then all the others are, are much smaller. So like, okay, if there's a, if there's a room for like... If this is, if this is really a new data category and there's room for, for a key player, then it's g- probably gonna be one of these guys.

By the way, I, I left out Weaviate, and I put Q-Qdrant in there. Do you know why?

**Alessio** [1:00:20]
No.

**Swyx** [1:00:21]
Uh, Anthopic and, uh, OpenAI both use, both use, use Qdrant-

**Alessio** [1:00:27]
Hmm

**Swyx** [1:00:28]
... for their internal RAG solutions, which means that for whatever reason, I, I... We should probably interview Qdrant. They passed the evals when Weaviate and, and-

**Alessio** [1:00:36]
Uh-huh

**Swyx** [1:00:36]
... you know, Milvus and all the others didn't, uh, which is interesting.

**Alessio** [1:00:39]
Yeah. Yeah, yeah, yeah.

**Swyx** [1:00:40]
There's, there's a lot that we don't know.

**Alessio** [1:00:42]
Yeah. Interesting. Yeah, I, I think like, I mean, going back to your point of like LangChain building LangSmith, at some point, some of the vector databases are gonna be like, "Why am I letting my customers use LlamaIndex?" You know?

It's like, I should be the RAG interface since I'm owning the data. I think that's-

**Swyx** [1:00:59]
Yes, yes. That's, uh, that's why I put them next to each other.

**Alessio** [1:01:01]
That, that's the fight.

**Swyx** [1:01:01]
Right now they're friends.

**Alessio** [1:01:03]
Yeah, right now. But-

**Swyx** [1:01:04]
Yes

**Alessio** [1:01:04]
... I, I mean, i-if we think about, uh, the Jamstack era- ... you know, you had Vercel started as Zeit, which was just a CDN, and then you had Netlify, you had all these companies, and then Vercel built Next.js, and so they moved down from the CDN to the framework.

**Swyx** [1:01:23]
Mm.

**Alessio** [1:01:24]
You know? And it's like now they use the framework to then enable more cloud and platform products.

**Swyx** [1:01:30]
Mm.

**Alessio** [1:01:32]
Which way is it gonna give this way? I, I think what we learned from before is that you'd rather own the framework and then have the cloud to support it than like just have Netlify and not have your own framework.

Given, just given the way-

**Swyx** [1:01:45]
Ah. Interesting

**Alessio** [1:01:45]
... the two companies are doing now.

**Swyx** [1:01:46]
So, so those who don't know, I worked at Netlify, and I was very, very intimately involved in this. So, so it's, uh-

**Alessio** [1:01:52]
So we don't have to say any private-

**Swyx** [1:01:53]
No, no, no. It's, it's fine. It's fine.

**Alessio** [1:01:54]
Yeah, yeah.

**Swyx** [1:01:54]
It's, it's well known that Vercel won, uh, and, and Netlify has, has pivoted away to a different market. Um- Uh, but is it overle- overlearning from an n of one example-

**Alessio** [1:02:06]
Well-

**Swyx** [1:02:06]
That you always want to own the framework?

**Alessio** [1:02:08]
No, no, no, no. No, because then the counterexample is the same, which is Gatsby.

**Swyx** [1:02:12]
Yes.

**Alessio** [1:02:12]
Where you own the framework, you don't own the cloud-

**Swyx** [1:02:14]
Yeah

**Alessio** [1:02:14]
... and then you don't make money either.

**Swyx** [1:02:15]
Yeah.

**Alessio** [1:02:15]
So the-- It, it's kinda like-- I, I think there-- We still gotta figure out where, like, the gravity is in this market, you know?

**Swyx** [1:02:23]
Mm-hmm.

**Alessio** [1:02:23]
I think a lot of people will say the gravity's in the model. A lot of people will say the gravity's in the embeddings and the data that you put into it. A lot of people don't know what they're talking about.

So I, I think 2024 is supposed to be the year of AI in production. I think we're gonna learn soon who bleeds into where, you know?

**Swyx** [1:02:42]
I, I think that's like a-- that's-- that statement is like a, you know, this is the year of Linux on the desktop thing. Like, it's just always gonna be true.

**Alessio** [1:02:50]
Right.

**Swyx** [1:02:50]
People are always gonna be saying it. How can a-- We're gonna be here one year later, and then it's like, "Yeah, this year is the year of AI production." Uh, and it's always gonna be incrementally more true. But, like, you know, what is the catalyst?

What is the big event that says-- that you, that you will point to and say, "Aha, now it's in production"? I don't know.

**Alessio** [1:03:07]
I think actually being the-- It's not in production, you know? Like, a lot of companies-

**Swyx** [1:03:12]
Sure, sure, sure

**Alessio** [1:03:12]
... a lot-- It, it's funny. Like, one, they're just like a inerent timeline that large companies work within. GPT-4 came out in, like, April. That's like eight months. And it's like most companies don't buy things within eight months-

**Swyx** [1:03:26]
Yeah

**Alessio** [1:03:26]
... or, like, implement them.

**Swyx** [1:03:27]
Yeah.

**Alessio** [1:03:27]
So I, I think, like, part of it is just, like, a physics time limit, that, like, even people that have been really interested, you just cannot go through the whole process-

**Swyx** [1:03:37]
Yeah

**Alessio** [1:03:37]
... of getting them live to all of your customers. So I think we'll see more of that in good and bad, right? It's gonna be a lot of failures and a lot of successes, hopefully. Um, yeah.

**Swyx** [1:03:47]
Yeah. Any other commentary on, like, tooling, RAG, ops, anything like that?

**Alessio** [1:03:52]
Um-

**Swyx** [1:03:52]
I mean, for, for what it's worth, like, I always, I always say, I always tell people, like, as much as I'm interested in fine-tuning, I think RAG is here to stay. Like, don't, don't even f-doubt it. Like, this is a necessary part that every a- every AI engineer should know.

**Alessio** [1:04:04]
Yeah.

**Swyx** [1:04:04]
So-

**Alessio** [1:04:04]
Well, I, I think, uh, yeah, it's tied to the infinite context thing, right?

**Swyx** [1:04:08]
Yeah.

**Alessio** [1:04:08]
I think the, the leftover question is, like, do you wanna have infinite context and hope that the model is good enough at parsing which parts matter to your query? Or do you wanna use RAG and wrap very specific context injection?

**Swyx** [1:04:23]
Yeah.

**Alessio** [1:04:23]
Uh, I think so far most people will say, "I'd rather do context injection with just what I care about than put a whole document in there and hope the model gets it," but maybe that changes.

**Swyx** [1:04:34]
I don't. I, like-

**Alessio** [1:04:35]
Yeah, no, I mean-

**Swyx** [1:04:36]
It cannot-- There's, there's no way it changes.

**Alessio** [1:04:39]
Hey, you know, that's great for LlamaIndex.

**Swyx** [1:04:41]
Yeah, yeah, no, it's great.

**Alessio** [1:04:41]
You know? Like, great. Like, it's gonna make a lot of money, I guess.

**Swyx** [1:04:44]
Yeah, yeah. No, it's not clear that they're gonna make a lot of money, right? Because they're just an open source project. I don't think they've launched a commercial thing yet.

**Alessio** [1:04:50]
I, I don't, I don't think so. Uh, because, uh, yeah, Jerry was talking-

**Swyx** [1:04:55]
I, I thought-

**Alessio** [1:04:55]
... about it on the podcast, but it wasn't-

**Swyx** [1:04:58]
Yeah, yeah

**Alessio** [1:04:58]
... um, yeah.

**Swyx** [1:04:58]
Yeah. So, I mean, we'll see, we'll see what they launch this year. Um, I, I, yeah, I do have-

**Alessio** [1:05:01]
The year of AI in production.

**Swyx** [1:05:02]
Yes.

**Alessio** [1:05:03]
The year of LlamaIndex in production.

**Swyx** [1:05:05]
Yeah. Uh, okay, so that's the four wars. Uh, we also covered a bunch of other, um, non-wars that we, that we skipped over. Um, I did, I did remember that you actually just, um, published a piece on the semantic, uh, versus-

**Alessio** [1:05:18]
The syntax, the semantics

### Code & Agents

**Swyx** [1:05:19]
... syntax things. Do you wanna cover that as a-

**Alessio** [1:05:21]
Yeah

**Swyx** [1:05:22]
... evolution?

**Alessio** [1:05:22]
I, I think, like, I, I kinda mentioned this a couple times on the, on the podcast, but basically the idea of, like, um, code has always been the gateway to programming machines, and we spend a lot of time making it easier.

So you go from, um, punch cards to, like, COBOL to C to Python, um, just to make it easier for the person to read and write the code. And through it, we started adding kinda like these semantic functionalities in it.

So in Python, you can do array.sort. You don't need to know bubble sort. You don't need to know any algorithm that you l-learn in school to do it. And I think the models are kind of like a hundred and X-ing this, which is like, now all you need to do is, like, create a, a sign-up form, you know?

Where people put a name, email, and send it to this endpoint. Um, so it's gonna be a lot easier for people that know the semantics of the business, which is, you know, your product managers or your business people, the layer that goes from customer requirements to implementation, basically, um, and have them intervene in the code.

So, you know, uh, how many times as an engineer you have to, like, go change some button color or, like, some button size, like, these small things that, like, you really shouldn't be doing. Uh, and now you can have people with natural language intervene in the code and write code that can actually be merged and put in production.

Uh, I also wrote the bear case for it, which is like, we already have so much trouble getting engineering teams to collaborate and get all their changes together without conflicts and all of these things that, uh, maybe also having non-technical people try and do things will be hard.

And models are-- They just think about solving the task at hand. They don't think about... I, I've always told my engineers, it's like, "You need to leave the code base better than you found it." You know? If you're, like, writing something, it's like-

**Swyx** [1:07:12]
Ugh

**Alessio** [1:07:13]
... just we cannot always keep adding, like, uh, quick hacks, you know? And I think models are great at quick hacks, but sometimes it's like, oh, this is like the 16th button that you've changed the, uh, style for.

You should make a class for it. And that's like the dumbest example. Um, so I, I think that's-- If that happens, then I think I'll be a lot more bullish on, like, coding agents, you know? Um, but I think now that's kind of the-- Until you can have non-technical people manually query models and look at results and then say, "This is ready to go," it's gonna be hard to have autonomous agents do it, so.

**Swyx** [1:07:52]
Yeah. So, um, I actually had a tweet about it today, um, because, uh, Itamar from Codium actually published, um, a prompt engineer-- a flow engineering as, as his next evolution of prompt engineering.

**Alessio** [1:08:04]
Mm-hmm.

**Swyx** [1:08:04]
Um, and they've been working on, um, you know, i-in IDE agents. Uh, they, they, they call it agents. You can debate about the definition of an agent, uh, you know, uh, at the end of the day. So I, I, uh, my split of it is inner loop versus outer loop.

Um, which I think you, you-

**Alessio** [1:08:18]
Mm-hmm

**Swyx** [1:08:18]
... understand that. Maybe I have to explain it to the audience, because every time I talk about it to developers, they don't, they've never heard of it.

**Alessio** [1:08:23]
Hmm.

**Swyx** [1:08:23]
Um, so inner loop is everything that happens between a Git commit. Uh, outer loop is everything happens after the, the, the commit is, is, is, uh, is committed and it's pushed up for, for PR. Um, so maybe that's too reductive, but that's something like that, right?

Like inner loop happens within your IDE, outer loop happens in GitHub.

**Alessio** [1:08:42]
Mm-hmm.

**Swyx** [1:08:42]
Something like that. Um, okay. So, uh, uh, I think your conception of an agent is outer loopy, uh, especially if it's non-technical, right? Like the, the dream, like you mentioned Sweep.dev in your writeup, um, and there, there's also CodeGen, there's also maybe Morph, uh, depends what Morph is doing.

**Alessio** [1:08:59]
Mm-hmm.

**Swyx** [1:09:00]
Uh, and, and, and there's a bunch of other people all doing this, this stuff. Um, even small developer was also like-

**Alessio** [1:09:05]
Right

**Swyx** [1:09:05]
... you know, write, uh, write it, write in English and then create a code base. And, uh, I think it's just not ready for that. Um, ou- outer loop is, is a, is a mirage, is, uh, that, you know, it's like going to forever be five years away, and the people working on inner loop companies have, have been the, the right bet, and you can work on inner loop agents.

Uh, I think actually Code Interpreter is an inner loop agent-

**Alessio** [1:09:26]
Mm-hmm

**Swyx** [1:09:27]
... uh, in the sense of like, uh, it's like limited self-driving, right? It, it's, it's, it's kind of like you, you have to, you have to, uh, have your attention on it. You have to watch it. It can only drive a small distance, but it is somewhat self-driving.

And so I think if you have this like gradations in your outlook on autonomous agents and you don't expect the w- the, everything to jump to level five at once, but if you have a idea of what level one, two, three, four, five looks like for you, I haven't really defined it apart from this concept of inner loop versus outer loop.

But once you've defined it, then you can be like, oh, like we're making real progress on, on this stage and, and you know, this other stage too early for now, but at some point somebody will do it.

**Alessio** [1:10:04]
Mm-hmm. Yeah. Yeah. I, I think like, yeah, maybe level one is like, I think of it more as a, just the auto-completion in the IDE, you know?

**Swyx** [1:10:13]
Yeah.

**Alessio** [1:10:14]
Level two is like asking Cursor, "Hey, how, how can I make this change?" You know? But then level three should be like, to me it's like we need to separate the inner loop from the IDE, you know? Like I need to make a code change.

Sometimes I don't, I shouldn't go in the IDE. Sometimes I should be in the UI of the product and say, "Hey, that needs to be changed." Kind of like the, uh, all the preview environments companies want you to put comments.

**Swyx** [1:10:42]
Oh, yeah.

**Alessio** [1:10:42]
The PMs put comments.

**Swyx** [1:10:43]
Yeah.

**Alessio** [1:10:44]
Like how do you go from that to code changes? There should be enough there to make the code changes happen-

**Swyx** [1:10:51]
Yeah

**Alessio** [1:10:51]
... you know, through a supervised interface.

**Swyx** [1:10:53]
Yeah, that's outer loop.

**Alessio** [1:10:55]
Uh, yeah. That, but, but that's kind of like I, I think what, uh, these models are doing is like change where the loops start and end, you know? Because now you can create code in the outer loop.

**Swyx** [1:11:05]
Yeah.

**Alessio** [1:11:05]
You know?

**Swyx** [1:11:06]
Yeah.

**Alessio** [1:11:06]
Before you couldn't do it. Um...

**Swyx** [1:11:08]
That's the dream. That's the dream.

**Alessio** [1:11:10]
Yeah.

**Swyx** [1:11:10]
Um, yeah. I, I, I have, uh... Yeah, I anyway, my, my focus right now I, I'll say if anyone cares is like, you know, I think the only thing that's working is inner loop, and you should just use inner loop things aggressively, build inner loop things-

**Alessio** [1:11:22]
Mm-hmm

**Swyx** [1:11:22]
... aggressively, invest in them, uh, and then keep an eye on the outer loop stuff.

**Alessio** [1:11:27]
Yeah.

**Swyx** [1:11:27]
'Cause it's, uh, still very early. Uh, I did invest in CodeGen.

**Alessio** [1:11:30]
Mm-hmm.

**Swyx** [1:11:31]
Uh, this JHacks, uh, thing, which we, which we, uh, uh, mentioned briefly at the Sourcegraph episode. Uh, do we have other things that we wanna mention or do you wanna sort of keep it to-

**Alessio** [1:11:40]
Um-

**Swyx** [1:11:40]
... the four wars?

**Alessio** [1:11:41]
I think that's great. I, I thought it was gonna be much shorter, but we're-

**Swyx** [1:11:44]
Yeah, yeah. We went-

**Alessio** [1:11:44]
... one hours 15 minutes.

**Swyx** [1:11:46]
Yeah, yeah. We probably have some stuff-

**Alessio** [1:11:47]
I thought we were gonna run through everything.

**Swyx** [1:11:48]
Yeah.

**Alessio** [1:11:48]
Yeah.

**Swyx** [1:11:49]
I mean, are there like, okay, maybe like top two things from December that you have commentary on.

**Alessio** [1:11:56]
Um, I, I think the, the needle in a haystack thing.

### Needle & Hardware

**Swyx** [1:12:00]
Okay.

**Alessio** [1:12:00]
So-

**Swyx** [1:12:00]
Maybe you wanna explain that first.

**Alessio** [1:12:01]
Yeah. Basically like, um, Anthropic, uh, should... There was like one example floating around about clothes context window and you basically gave it this like super long context on, I think like things to do in San Francisco or something like that.

And then it was like, what is the most fun thing to do in SF? And it always, it didn't... They made this nice chart of like, okay, based on where it is in the context, it gave a better, worse response.

And then Anthropic responded and they were like, "Oh, you just need to add, here's the most relevant sentence in the context as part of the assistant prompt."

**Swyx** [1:12:36]
Mm-hmm.

**Alessio** [1:12:37]
And then the chart turns all green all of a sudden. And I'm like, we cannot still be here, right? Like it cannot, it cannot... This is like some-

**Swyx** [1:12:48]
And, and you have Anthropic like telling people, "Oh yeah, it's just like just add this magic string and it works."

**Alessio** [1:12:52]
Yeah. It's some like Riley Goodside wizardry. It's like I don't wanna do that anymore. I thought Riley, I, I thought like, you know, in the early days of GPTs, like Riley Goodside was doing so much great work on like prompt engineering and, and whatnot.

We shouldn't be there anymore. There shouldn't be somebody telling me, or like the, the GPT-4 like, uh, I'll give you a $200 tip-

**Swyx** [1:13:14]
Oh, yeah, yeah

**Alessio** [1:13:15]
... if you do this right and like-

**Swyx** [1:13:16]
So I collected a whole bunch of like, uh, state-of-the-art prompting techniques. Um, yeah, yeah. So, so if you tip the model, it, it'll give you better results if you, if you, um, uh, promise that it'll... So, so, okay.

Uh, here's the current state-of-the-art for GPT prompting. It's Monday in October, the most productive day of the year. You have to take a deep breath and you have to think step by step. You have to, uh, return to full script.

You are an expert on everything. I will pay you $20, just do anything I ask you to do. I will tip you $200 every request you answer correctly. Um, and your competitor models said you couldn't do it, but you can do it.

Or I think there's another one that, that I didn't put in here, but it's like, you know, my grandmother's dying. Um, this is an emergency. Please help me do it.

**Alessio** [1:13:53]
Yeah. That, that, that's actually my, um, I think my most viewed tweet ever. Uh, at Open AI Dev Day, I tweeted, uh, "No more return JSON or my grandma's gonna die."

**Swyx** [1:14:04]
Oh, okay, okay, okay.

**Alessio** [1:14:04]
When they announced JSON mode and people, people love the, people love the gay grandma stuff.

**Swyx** [1:14:09]
I haven't heard as much, uh, uptake on JSON mode. Uh, I think it's still-

**Alessio** [1:14:14]
That's, that, that's the thing with all this AI stuff, right? It's like, I mean... And sometimes we're like part of it. If I think about our, uh, um- ChatGPT plugins episode. I think in the moment people are just like, "Oh, this is gonna be such a big deal."

**Swyx** [1:14:26]
Mm-hmm.

**Alessio** [1:14:26]
And then it takes varied amount of times-

**Swyx** [1:14:29]
Yeah

**Alessio** [1:14:29]
... to like really pick up.

**Swyx** [1:14:30]
Yeah.

**Alessio** [1:14:30]
You know?

**Swyx** [1:14:31]
Do you think that'll happen to GPTs?

**Alessio** [1:14:34]
I, I think like most people that I see using GPTs right now are trying to get around some sort of weird limitation of the base model, you know, or just trying to have a better system prompt.

**Swyx** [1:14:47]
Yeah.

**Alessio** [1:14:47]
But like at some point there's limited value to get out of it. So the question is like what's gonna incentivize people to build more on it versus just building their own thing, huh, out of it? I don't know.

**Swyx** [1:15:00]
Yeah. Um, okay, so I guess my pick for highlight of last month, there's two. Uh, one, we finally got Gemini.

**Alessio** [1:15:08]
Right.

**Swyx** [1:15:08]
Uh, I think the marketing was dishonest. I think-

**Alessio** [1:15:11]
Yeah-

**Swyx** [1:15:11]
I think the-

**Alessio** [1:15:11]
... we need the soundboard. Wham, wham, wham.

**Swyx** [1:15:14]
But, but still it's, it is, it is a, it is a sort of model, um, it, it, it is a credible, very, very credible alternative to OpenAI and we should, we should be happy for that because otherwise we live in a, a OpenAI only world.

Um, and, and Gemini is pre- basically the only other sort of leading contender, uh, until Llama 3 drops whenever, whenever Llama 3 comes out.

**Alessio** [1:15:33]
It's kinda-- I mean, Sox said today they're training it, so.

**Swyx** [1:15:36]
Yeah, it sounds like today they're training it.

**Alessio** [1:15:38]
Yeah.

**Swyx** [1:15:38]
Um, for me, I guess, uh, I'm still very interested in like the, the, the hardware meta game. Uh, this is a s- much smaller stakes, but very personal. I think, uh, uh, recently especially, you know, we're recording this mid-January, um, so after CES, after Rabbit R1 launched, um, I think there's a lot of interest in, in hardware.

Um, I don't know how you feel about it as an enterprise software investor but, uh, I think that hardware is hard, but also it, it captures context and is, is-- makes AI usable in ways that, um, you, you cannot currently cons- uh, think about.

Um, and if you-- and, and, you know, everyone dreams of building an assistant like Her-

**Alessio** [1:16:18]
Mm-hmm

**Swyx** [1:16:18]
... in the, in the movie Her, that is a hardware piece. That is actually not only software. And probably the hard part is the en- the engineering for the hardware and, and then the, the sort of AI engineering for the, the assistant within, within the hardware.

Um, so yeah, I mean, yeah, I'm an investor in Tab. Um, I, uh, and I, I see, I see a lot of like, um, you know, interest this month, um, but it started last month-

**Alessio** [1:16:39]
Mm-hmm

**Swyx** [1:16:39]
... uh, with the, the launch of Humane as well.

**Alessio** [1:16:41]
Yeah.

**Swyx** [1:16:41]
I don't know if you have, uh, you have thoughts on any of those things.

**Alessio** [1:16:43]
Well, I think this year we also get the Apple Vision Pro thing, so I think this is-- there's gonna be a ton of experimentation.

**Swyx** [1:16:50]
Yeah.

**Alessio** [1:16:50]
I think Rabbit got, uh, the right nostalgia factor, you know?

**Swyx** [1:16:55]
Yeah.

**Alessio** [1:16:55]
It kind of looks like a-

**Swyx** [1:16:56]
The toy-

**Alessio** [1:16:57]
... Togepi Rex.

**Swyx** [1:16:57]
The toy-looking, yeah.

**Alessio** [1:16:57]
Yeah. It looks like a-

**Swyx** [1:16:58]
Togepi type thing, yeah

**Alessio** [1:16:59]
... Game Boy Advance.

**Swyx** [1:17:00]
Yeah.

**Alessio** [1:17:00]
Something like that. Um, I'm curious to see what you get beyond that. I think like, yeah, I mean, obvious like right where we have the studio building-

**Swyx** [1:17:09]
Yeah

**Alessio** [1:17:09]
... um, Tab and I, I think that's another interesting form factor. And I think if you ask them-- I, I think in our circles a lot of people are like, "Well, what about privacy and all these things?" But he will tell you that we're kinda like a special group, that most people value convenience over privacy-

**Swyx** [1:17:26]
Yes

**Alessio** [1:17:27]
... as you learn from the social medias uh, o- of the last few years. So, um, yeah.

**Swyx** [1:17:31]
Yeah.

**Alessio** [1:17:32]
I'm really curious to-

**Swyx** [1:17:33]
Yeah

**Alessio** [1:17:33]
... to see how it develops.

**Swyx** [1:17:34]
I really like technology where it's, uh, you're slightly uncomfortable with it on a social level. Um, a- and, and so, you know, for, for Uber it was like this regulation around taxis.

**Alessio** [1:17:44]
Mm-hmm.

**Swyx** [1:17:44]
For Airbnb it was, um, you know, staying in strangers' homes. And now it turns out for OpenAI it was, uh, training on people's content.

**Alessio** [1:17:53]
Right.

**Swyx** [1:17:54]
Right? And now it's becoming a, a, a matter of regulation, and OpenAI's data partnerships are, you know, a form of, you know, private regulatory capture, which is a playbook that is fantastic. Like if you-- if-- uh, I hope it was on purpose because whoever did that was- is a, is a genius.

Um, so I'm like, okay, like, you know, I, I, I do think that every great new company, especially on the consumer side, is provocative in that sense. Like they're doing something that is not yet kosher.

**Alessio** [1:18:18]
Mm-hmm.

**Swyx** [1:18:18]
Uh, and, and so I think like the Humanes, the Tabs, um, anything that is, um, working on that front where it's like, yeah, I'm not, I'm not sure I'm comfortable with this and then but maybe it could change.

Uh, that, that, that is, that is a, that is a really interesting shift. Um, so y- I'm excited from, from that point of view, but at the same time most hardware companies fail very, very quickly.

**Alessio** [1:18:38]
Right. Yeah.

**Swyx** [1:18:38]
They, they have a very hot start, and then, you know, everyone puts it in their drawer and then, then never looks at it again. Uh, so I'm very, very aware of that. Uh, but I think it's-- I, I mean, it's something interesting, and I, I do think-- so, uh, here's the, the core thing of it, right?

Avi doesn't think it's a hardware company. Avi-- like most of the, the cost of the, the $600, uh, for, for Tab is going towards GPT costs-

**Alessio** [1:19:00]
Mm-hmm

**Swyx** [1:19:00]
... because it's actually processing context. And the, the whole idea is that context is all you need. Like in this world of like, you know, AI applications, like whoever has the most unique context wins, right? A unique context could be the quality data war, right?

Like a unique context is like, you know, I have Reddit info, I have Stack Overflow info, I have New York Times info. Um, if I have info on everything you say, you say and do at all times, that is something that, uh, no one else has.

A- and, and if, if, if he, if he becomes a good store of that, then like-

**Alessio** [1:19:30]
Mm-hmm

**Swyx** [1:19:30]
... what can you, what can you build with that?

**Alessio** [1:19:31]
Yeah.

**Swyx** [1:19:31]
So I'm most excited for him to, to expose the developer API 'cause then I can come in and do all my software stuff. But, uh, he has to build the hardware layer and get acceptance for that first.

**Alessio** [1:19:40]
Right. Um, yeah, no, I'm, I'm excited to see. I, I'm sure we're gonna see a lot of people walk around with them, so I'm, uh- ... I'm excited to see.

**Swyx** [1:19:48]
I, I actually-- so I, I think he doesn't like me because, uh, I asked for a off button. Like I said, I wanna be able to guarantee you if I-- if we're having a conversation-

**Alessio** [1:19:57]
Right

**Swyx** [1:19:57]
... I want to show you, you see it's off, right? It's, it's kind of like, oh, yeah, my phone is on silent mode, right?

**Alessio** [1:20:02]
Mm-hmm.

**Swyx** [1:20:02]
There's a physical silent mode button. But now-- but he, he just wants it to be always on.

**Alessio** [1:20:07]
That's a whole new market, uh, like a soundproof- ... like soundproof storage for your AI pendant so that you can guarantee the person-

**Swyx** [1:20:16]
Yeah, yeah

**Alessio** [1:20:16]
... cannot hear you. Um, awesome. No, this was, this was fun. Um, please, if you're still listening after one hour 21 minutes, let us know, uh, what we did right, what we did wrong, what you would like to see differently.

Uh, it's the first time we tried this out. Um, but yeah.

**Swyx** [1:20:33]
Awesome. Thanks for doing this.

**Alessio** [1:20:34]
Cool.

---

This library is powered by PodHood (https://podhood.com), the podcast website platform.
