# [State of AI Startups] Memory/Learning, RL Envs & DBT-Fivetran — Sarah Catanzaro, Amplify

Latent Space · 2025-12-30

<https://addtry.com/4a76e39c-694d-4f7d-8598-fd7ab731fb6d>

Sarah Catanzaro of Amplify Partners argues that the DBT-Fivetran merger was not the death of the modern data stack but a path to IPO targeting $600M+ combined revenue, that data catalogs failed because they were built for humans rather than machines, and that RL environments are a fad. She notes that frontier labs use dbt and Fivetran for training data curation and agent analytics, and criticizes $100M+ seed rounds raised without near-term roadmaps. For 2026, she identifies personalization via memory and continual learning as the key retention unlock, observing that AI founders are unfamiliar with growth concepts like k-factor. She prefers real-world logs over synthetic RL environments, citing Cursor's use of user activity, and says the most exciting startups combine hard research problems (RAG, rule-following, continual learning) with applications that were previously impossible.

## Questions this episode answers

### Are simulated RL environments a viable long-term approach for training AI agents?

Sarah Catanzaro calls RL environments a fad. She argues the best environment is the real world—using actual logs and user traces from apps like DoorDash—rather than synthetic clones. While they may create short-term value, she believes building app clones isn't useful long-term, and companies like Cursor successfully use real user activity to improve their agents.

[0:24](https://addtry.com/4a76e39c-694d-4f7d-8598-fd7ab731fb6d?t=23500)

### Why are AI startups raising $100 million+ in seed rounds without a clear product roadmap?

Sarah Catanzaro observes that many founders raise over $100M at seed with no near-term milestones, often to inflate valuation for hiring leverage. She warns this is alarming because investors have only days to commit and can't assess execution ability. Moreover, employees may end up with worthless equity if the company exits below the capital raised, since valuations are fabricated until a liquidity event.

[0:11](https://addtry.com/4a76e39c-694d-4f7d-8598-fd7ab731fb6d?t=10667)

### What does the dbt-Fivetran merger reveal about the state of the modern data stack?

Sarah Catanzaro states the dbt-Fivetran merger is not the end of the modern data stack but a strategic move toward IPO. Both companies were growing and beating revenue targets, but the public market now expects over $600M in revenue. The combined entity nears that threshold, accelerating liquidity. She also notes frontier AI labs are adopting both tools for training data management and agent analytics, sustaining demand.

[0:01](https://addtry.com/4a76e39c-694d-4f7d-8598-fd7ab731fb6d?t=1150)

## Key moments

- **[0:00] Intro**
- **[1:06] Merger**
  - [1:09] Sarah Catanzaro: 'a lot of people look at the dbt Fivetran merger and talk about the end of the modern data stack, and I think that is a fundamentally wrong take.'
  - [1:50] IPO bar now above $600M revenue; dbt and Fivetran combined targeting $600M.
  - [2:28] Frontier AI labs like Thinking Machines used dbt within weeks of formation for training data management.
- **[3:56] Training Data**
- **[5:23] Data Catalogs**
  - [7:44] Data catalogs failed because they targeted human discoverability, not machine governance; opportunity in metadata services.
- **[8:16] Frontier Stacks**
- **[10:13] Seed Craze**
  - [10:32] $100M+ seed rounds with no near-term roadmap forcing 7-day decisions are alarming.
  - [13:43] Antithesis raised $100M seed led by Jane Street for deterministic simulation testing, used by Palantir.
  - [16:04] Sarah Catanzaro: 'valuation, until a company exits, it is an entirely made-up number.'
- **[17:06] World Models**
- **[18:56] Personalization**
  - [18:56] Memory and continual learning will be a key 2026 theme as AI apps need personalization to improve low retention.
- **[23:25] RL Environments**
  - [23:42] RL environments are a fad; real-world logs and traces from apps like DoorDash are better than clones.
- **[25:37] Startup Bets**
  - [25:47] Sarah Catanzaro's favorite startups marry hard research (RAG, rule following) with killer apps that were previously impossible.
- **[27:57] Outro**

## Speakers

- **Sarah Catanzaro** (guest)

## Topics

Startups, Data Tools, Reinforcement Learning

## Mentioned

Antithesis (company), Atlan (company), Cognition (company), Data World (company), Fivetran (company), Habia (company), Harvey (company), Metaphor (company), OpenAI (company), Runway (company), Sierra (company), Snowflake (company), Spiral (company), Cursor (product), Hex (product), Looker (product), RocksDB (product), Tableau (product), Vortex (product), dbt (product)

## Transcript

### Intro

**Host** [0:03]
Lighten space like 20 to five. Breakups. Lighten space like- Okay. We're here with Sarah Catanzaro from Amplify. Welcome.

**Sarah Catanzaro** [0:15]
Thank you.

**Host** [0:15]
First time on the pod.

**Sarah Catanzaro** [0:16]
Great to be here. I-

**Host** [0:17]
Took too long

**Sarah Catanzaro** [0:17]
... I know, I know. We've, we've known each other for so long. And yeah, never made an appearance.

**Host** [0:22]
Uh, it also made the transition from data to AI, I guess. I, I don't know if, if... Um, I did. I don't know if you were always, like, s- as deep on, on, on AI. Um, but obviously there's a lot of simpatico.

**Sarah Catanzaro** [0:33]
Yeah. I've always actually kind of oscillated between data and AI.

**Host** [0:37]
Sure.

**Sarah Catanzaro** [0:38]
Um, like arguably, I started my career in, quote-unquote, "AI." It was just more like symbolic systems back then. But as you said, I think, like, they're, they're so symbiotic, like, it- it's almost hard to divorce them.

**Host** [0:50]
Yeah.

**Sarah Catanzaro** [0:50]
That's actually what brought me into data. I was like, "I want to better understand what happens when I write a SQL query," so...

**Host** [0:56]
Yeah. Um, let's briefly touch on data because I, I think obviously that's, that's a lot of where you and I first met. Uh, dbt Fivetran, that was so cool. I mean, or-

**Sarah Catanzaro** [1:05]
Yeah

**Host** [1:06]
... how do you, how do you, how do you think about the end of the modern data stack?

### Merger

**Sarah Catanzaro** [1:09]
Okay. So, so, like, a lot of people look at the, like, dbt Fivetran, uh, merger and, like, talk about the end of the modern data stack, and I think that is, like, a fundamentally wrong take. Both of these companies were growing, you know, very healthily.

Both of these com-

**Host** [1:28]
And you f- you funded dbt?

**Sarah Catanzaro** [1:29]
We funded dbt. So, so like both of the companies were actually, like, beating their revenue targets. I think what you're more seeing is a, you know, IPO environment wherein companies are expected to have far more than, you know, like 100 million revenue.

And so-

**Host** [1:47]
What would you say the bar is now? 300?

**Sarah Catanzaro** [1:48]
No. Like above-

**Host** [1:50]
Five

**Sarah Catanzaro** [1:50]
... 600.

**Host** [1:51]
600.

**Sarah Catanzaro** [1:51]
Yeah, yeah.

**Host** [1:52]
And the combined company is 400?

**Sarah Catanzaro** [1:54]
Uh, I believe that they'll actually be close to 600.

**Host** [1:57]
Okay.

**Sarah Catanzaro** [1:57]
I don't have the exact number.

**Host** [1:59]
But they clearly just getting ready-

**Sarah Catanzaro** [2:01]
Yes

**Host** [2:01]
... for IPO.

**Sarah Catanzaro** [2:01]
So, so, so, you know, basically, like, the merger was a way to accelerate that path to liquidity. You know, as you might remember-

**Host** [2:09]
A- and they were the presumptive winners in their categories anyway, so...

**Sarah Catanzaro** [2:11]
Exactly, exactly. Um, you know, I think one of the things that has actually, uh, pleasantly surprised me, um, and this speaks to, again, the symbiotic relationship between, you know, data and AI, many of the big frontier labs are actually using both dbt and Fivetran.

I recall talking to folks at, um, Thinking Machines, like, within weeks of the company's formation, and dbt was already an important part of their stack. Certainly, like, training data sets need to be managed. We need insight into what users are doing on these platforms and, in fact, like the way in which you would analyze interactions with an agent or analyze interactions with an LLM is even more complicated.

And so while I think perhaps, like, uh, the demand for analytics engineers, the demand for data scientists, uh, didn't explode in the way that some people thought, like analytics engineers are not one-third of personnel- ... uh, that doesn't actually mean that the demand for the tools, uh-

**Host** [3:13]
Yeah

**Sarah Catanzaro** [3:13]
... is not still, like, very prevalent.

**Host** [3:15]
Well, you got what you wanted. You wanted to democra- democratize things. You, you got it.

**Sarah Catanzaro** [3:18]
Yeah, yeah. I mean, I guess we democra- we, we, we, we democratized things by, uh, uh, perhaps reducing the need for the people. I don't know whether or not that is a good thing, but honestly, I do think that, like, the fact that, uh, it is easier than ever, uh, from a tooling standpoint for people to make data-driven decisions is probably a step in the right direction.

Um, and I've become actually convinced that, like, while every company does need analytics engineers and does need data scientists, they probably don't need armies of them. Um, and probably having like a moderately sized data and analytics team is a good thing.

**Host** [3:56]
Yeah. So you touched on an, an interesting thing I wasn't planning to ask, but this is interesting. So I come from the data field. Data was sy- synonymous of analytics.

### Training Data

**Sarah Catanzaro** [4:04]
Yeah.

**Host** [4:05]
But you're now saying that the dbt Fivetran are being used for training data. Is there any notable differences in the workloads or the requirements?

**Sarah Catanzaro** [4:14]
Undoubtedly there will be. I mean, I think one of the things that we saw with, uh, analytics that, you know, was surprising to some of the people in the data infrastructure space was that, like, the workloads were actually quite predictable.

Uh, they were quite predictable because, like, many of them were actually not being generated, uh, by humans, but rather by deterministic systems. So, like, a lot of it, uh, was, you know, like BI dashboards that are, you know, Tableau that is actually hitting your database.

Or maybe not Tableau, but like, uh, Looker, um, or you know, Hacks or something like that. I think with, uh, uh, like analyzing, curating, preparing data sets, it's a bit more ad hoc. Um, and so undoubtedly it will be, uh, less predictable.

I don't know if that really changes the way that we approach developing data infrastructure. You know, I talked... Uh, like some people are quite interested still in like things like learned indexes, learned optimizers, and it's a bit easier to build a learned optimizer, uh, if you have more predictable workloads.

And so, uh, it could change the way that we approach things like that.

**Host** [5:23]
Yeah. Data catalogs, do they become more important? Are they transferred?

### Data Catalogs

**Sarah Catanzaro** [5:26]
Oh, man, like straight to the gut. So, so that was something I got wrong. I-

**Host** [5:30]
I'm sorry, I don't know the background. What, what, what did you-

**Sarah Catanzaro** [5:33]
I, I, I just, I really believed that data catalogs were going to become an important part of, you know, the modern data stack. Um-

**Host** [5:42]
And the, the players are Atlan, uh, I, I, the, those, those-

**Sarah Catanzaro** [5:46]
Atlan

**Host** [5:46]
... she's, she's Singaporean, so I-

**Sarah Catanzaro** [5:48]
Yeah, yeah. Uh, the, there, there was, uh-

**Host** [5:51]
Data World

**Sarah Catanzaro** [5:52]
... Girl, Data World, Metaphor within our portfolio.

**Host** [5:55]
They've all struggled as a category?

**Sarah Catanzaro** [5:56]
They all have struggled a bit as a category. Uh, many of them have been, you know, acquired subsequently, uh, which suggests that like this was not, you know, perhaps a, a standalone category. As a data scientist, like I spent so much time working on data catalogs.

Um, and so, you know, I kind of felt like This was, like, this was the thing I wanted.

**Host** [6:19]
Yeah.

**Sarah Catanzaro** [6:19]
Like, I didn't wanna have to, like-

**Host** [6:21]
And, and-

**Sarah Catanzaro** [6:21]
... build the... Yeah

**Host** [6:21]
... more to the point also, like, pre-training data, you have a lot more he- heterogeneous data all over the place.

**Sarah Catanzaro** [6:26]
Yeah.

**Host** [6:27]
And, like, you, you need to keep on top of it, and, and you need to make it discoverable, accessible, and all that. So why didn't it work?

**Sarah Catanzaro** [6:34]
So, so I think there were a couple of things. I think we have seen some consolidation in, uh, the, you know, modern data stack, uh, particularly around, you know, some of the key components, whether it was, you know, Fivetran or DBT or, uh, you know, Hex or, you know, Snowflake.

Um, many of these products offered kind of, like, data cataloging, uh, capabilities as a feature, and I think for humans, that was good enough. Like, the, the data catalog that you had available in Snowflake was good enough. Uh, the data ca- logging capabilities available in DBT, like, tho- those were good enough.

Like, da- um, DBT-

**Host** [7:11]
DBT, like, obviously as they built the cloud, they were going to build it.

**Sarah Catanzaro** [7:14]
Yeah, yeah.

**Host** [7:14]
Like, what else do you do?

**Sarah Catanzaro** [7:15]
I mean, it's actually funny. In fact, uh, my colleague Bar at Amplify was, uh, the, like, products lead on the- these kind of-

**Host** [7:22]
DBT, yeah

**Sarah Catanzaro** [7:22]
... like, metadata services. Um, I think, ah, it's still not obvious to me, but I think one opportunity that might have existed and/or could've been realized was the opportunity to build data catalogs not for, uh, humans, but, you know, for, you know, machines.

This would look a little bit more like, you know, m- metadata services.

**Host** [7:48]
Mm.

**Sarah Catanzaro** [7:48]
Um, I don't just mean for agents, although I think, you know, that opportunity is arising more, but even, uh, uh, like microservices and things like that. Um-

**Host** [7:58]
Okay.

**Sarah Catanzaro** [7:59]
Yeah. So, so I do wonder at times, like, if we built data catalogs for the wrong people, uh, and potentially even, you know, for the wrong use cases. Like, I think a lot of, uh, data cataloging companies ended up focusing on, like, discoverability when perhaps, like, the real market opportunity was in governance.

### Frontier Stacks

**Sarah Catanzaro** [8:16]
Uh-

**Host** [8:16]
Governance, very important. Uh, any other comments just about what you know so far about the data stacks of the large labs? You know, I guess obviously a lot of data people might, who might be listening would want to sell into them.

**Sarah Catanzaro** [8:30]
Yeah. I mean, a couple of observations. One is that, you know, they are actually paying careful attention to their data stacks. I think they're thinking about, you know, problems ranging from, you know, data discoverability to, uh, data preparation, to even things like the efficiency of data loading.

Like, if you're unable to load data to a GPU efficiently, then the GPU is going to sit idle-

**Host** [8:54]
Yeah

**Sarah Catanzaro** [8:54]
... and that's going to be a kind of like a, a cost-

**Host** [8:57]
That is-

**Sarah Catanzaro** [8:57]
Yeah.

**Host** [8:57]
Yeah.

**Sarah Catanzaro** [8:58]
Yeah. Exactly. Um, so-

**Host** [9:00]
But what, what solution is, um, handles that? I don't actually-

**Sarah Catanzaro** [9:03]
I mean, I get to, to talk about-

**Host** [9:05]
Plug.

**Sarah Catanzaro** [9:05]
Yes, exactly. Plug my portfolio company is, uh... We have a portfolio company called Spiral that has developed a file format called Vortex. Um, uh, and, uh, they make data loading, like, super efficient. Um-

**Host** [9:18]
Specifically to GPUs?

**Sarah Catanzaro** [9:20]
Specifically to GPUs.

**Host** [9:21]
Okay.

**Sarah Catanzaro** [9:22]
Yeah, yeah.

**Host** [9:23]
Good to know.

**Sarah Catanzaro** [9:24]
One of the things that has surprised me, though, is actually that, like, so much data infrastructure has actually scaled quite elegantly to meet the AI use case. Uh, I-

**Host** [9:32]
You would hope.

**Sarah Catanzaro** [9:34]
You, you would, but like, like, uh, the scale of these AI companies, it's incredible. Um, and so like-

**Host** [9:41]
Uh, kind of like it's not as big as ads.

**Sarah Catanzaro** [9:43]
Maybe. Maybe. Yeah. I, I, I think that could change, you know, like a- uh, as agents actually become kind of, like, more prevalent or, and are interfacing with each other, and therefore, like, perhaps, like, the number of transactions explodes.

I have a friend who works on transactional databases at OpenAI, and I was like, "So you must be, like, building databases." Like, this is, like, a paradigm shift in terms of, like, the scale that, like, databases are, like, going to need to handle.

And he's like, "No, we use RocksDB." Like it, it, it's fine.

**Host** [10:11]
That's the one they acquired, right?

**Sarah Catanzaro** [10:12]
It scales. Yes, exactly.

**Host** [10:13]
Yeah, yeah. Um, very cool. Okay, let's just talk about funding a round minute 'cause o- obviously-

### Seed Craze

**Sarah Catanzaro** [10:17]
Yeah

**Host** [10:17]
... that's, that's, like, a, a big theme this year. What comes to mind in terms of looking back at 2025, uh, what stands out?

**Sarah Catanzaro** [10:25]
It was crazy. Um, yeah.

**Host** [10:28]
You can give anonymized examples of, like, what, what, what does crazy look like?

**Sarah Catanzaro** [10:32]
Yeah, I mean, I think crazy looks like, uh, raising upwards of $100 million.

**Host** [10:40]
Seed?

**Sarah Catanzaro** [10:41]
S- like, upwards of $100 million in a seed round where you have a long-term vision but not a near-term roadmap.

**Host** [10:49]
Yeah.

**Sarah Catanzaro** [10:49]
Uh, this is something that I'm seeing happening, uh, not just occasionally, but quite frequently.

**Host** [10:56]
Yes.

**Sarah Catanzaro** [10:56]
Um, and it definitely makes me anxious because firstly, like, uh, when founders are asking me, you know, "How much should I raise?" I'm typically saying, like-

**Host** [11:08]
Three, like, five.

**Sarah Catanzaro** [11:09]
Well, like, like, what do you need to do?

**Host** [11:11]
Yeah, yeah.

**Sarah Catanzaro** [11:11]
Like, what are your milestones for the next-

**Host** [11:14]
Your use of funds, you know

**Sarah Catanzaro** [11:14]
... let's call it, like, 12 to 24 months. Uh- ... what resources do you need in terms of, you know, headcount, compute, uh, equipment to, uh, unlock those milestones? And then, like, maybe add, like, a 20% buffer or something like that.

**Host** [11:28]
Yes.

**Sarah Catanzaro** [11:29]
Um, but doing that analysis requires you to, like, understand what you're going to build in the next zero to, let's call it, like, 24 months. But I've talked to some companies and they're like, "We're building a frontier lab for X."

And like, okay, cool, like, I get the long-term vision. There is an opportunity to, uh, you know, make AI more secure, make AI more humane, uh, make AI more data efficient, whatever it might be. Um, so, so, like, I'm bought into the long-term vision, and that, that, you know, for me as an investor is super important.

Like, so let's talk about, like, what your team's going to work on in the next six months. They're like, "Uh, maybe we might build a consumer app. Like, uh, you know, we're debating-"

**Host** [12:09]
I feel like I know exactly the company you're talking about.

**Sarah Catanzaro** [12:11]
But, but, but like, I wish I was talking about, like, one specific company. I'm actually talking about, like, several companies.

**Host** [12:18]
Oh, no.

**Sarah Catanzaro** [12:19]
Um, a- and look, like, I'd be a hypocrite to say that, like, I've never done investments like that. Uh, but I've done investments like that when, like, I really know the people and I'm like- They're gonna figure it out.

What is frightening about this funding environment is that you meet a founder, they're like, "I'm raising, you know, $100 million. I'm raising, like, a billion dollars," maybe at times. Um, and you need to make a decision in seven days, and I can't tell you what I'm gonna do for the next six months.

**Host** [12:46]
Yep.

**Sarah Catanzaro** [12:46]
And so, like, you have no way of even gaining conviction that they're going to figure it out-

**Host** [12:51]
Yeah

**Sarah Catanzaro** [12:51]
... because you only have, like, seven days to get to know them. I think what some of the founders are missing is, like, you only have seven days to get to know me. If you haven't figured it out, like, you probably want a partner who's going to be working closely with you to help you figure it out.

**Host** [13:06]
Oh, I mean, they're absolutely viewing it as transactional, right?

**Sarah Catanzaro** [13:09]
Yeah.

**Host** [13:09]
Like, you know, they don't care.

**Sarah Catanzaro** [13:10]
No, they care about, you know, the most money at the highest valuation. I mean, the crazy thing is that they don't even seem to care about d- dilution. It's just, like, the most money at the highest valuation.

**Host** [13:20]
Yeah, and, and, but, you know, it does send a signal that helps.

**Sarah Catanzaro** [13:23]
So, so, ah, I mean, I th- yes, I think it does right now send a signal.

**Host** [13:28]
Okay, I'll, I'll tell you how, how it, it affects me, and I hate it. I hate it, all right? Antithesis came out of stealth this week, right? And the, the, it's like the o- the only thing I know about them is they do something, something in AI testing, and Jane Street led a seed round of $100 million.

**Sarah Catanzaro** [13:43]
We invested in, in it too. I can tell you what they do, but-

**Host** [13:46]
Okay

**Sarah Catanzaro** [13:46]
... they, they do-

**Host** [13:47]
Well, like-

**Sarah Catanzaro** [13:47]
... like a deterministic simulation thing

**Host** [13:48]
... gate testing-

**Sarah Catanzaro** [13:49]
Yeah

**Host** [13:49]
... the, the thing that is the, the lead is the money.

**Sarah Catanzaro** [13:51]
Yeah.

**Host** [13:52]
And then like, okay, well, who else uses it other than Jane Street? Like, what do you do that's innovative?

**Sarah Catanzaro** [13:57]
Palantir.

**Host** [13:58]
Oh, okay. All right.

**Sarah Catanzaro** [13:58]
Warp Stream.

**Host** [13:59]
So, yeah. Okay, so-

**Sarah Catanzaro** [14:00]
But, but, but yeah. Okay

**Host** [14:00]
... anyway, so, so, um, maybe, maybe, um, Antithesis is a bad example 'cause they're actually legit, but like, uh, you know, there's, there's a lot of similar examples where, uh, they just lead with the money, and, like, there's no, not much substantiation behind it.

Maybe it's just bad storytelling, and that's why I as a podcaster get to talk to the... I just talked to, uh, General, General Intuition, and, like, once you spend some time with them, then you're like, "Oh, okay, this is why they raised $100 million."

But, like, without that context, it's, like, really hard to understand anything.

**Sarah Catanzaro** [14:27]
Well, and, and, and like I think there are some companies that are raising, you know, uh, $100 million or more because they need it. Like, a good example might be, like, Periodic. In addition to, you know, need-

**Host** [14:38]
Wet lab, yeah

**Sarah Catanzaro** [14:38]
... yeah, they need to build out a wet lab, and, like, designing a wet lab that can support high-throughput biology, which is absolutely critical to, you know, their goals, uh, that's costly. So, so, so like I understand why they need that, that funding, but again, there are others where, like, they don't have these near-term milestones.

I think the thing that is a little bit, you know, perturbing to me, many of them are doing it because it makes it easier for them to hire 'cause, you know, there are all of these candidates who, like, want to be...

want to work at a company that is, like, a unicorn or a near unicorn. They're pitching-

**Host** [15:13]
'Cause the alternative is work at a big lab where, you know, it's, yeah, the prestige and the money is there.

**Sarah Catanzaro** [15:18]
Yeah. Well, or the alternative is, like, work at, like, an early-stage startup. Uh, but, but, but, but, but like there's something about, like, the big valuation that becomes enticing.

**Host** [15:28]
Yeah.

**Sarah Catanzaro** [15:28]
They're also kind of, uh, pitching candidates. They, they have a compelling equity pitch where they're like, "Okay, maybe you're getting, you know, less than, uh, zero point, like, uh, 1% of the company, but, like, given the valuation, uh, the value of, uh, your equity is already, you know, like $10 million," or something like that.

**Host** [15:49]
And, and they also, uh, guarantee the dollar value-

**Sarah Catanzaro** [15:52]
Uh

**Host** [15:52]
... on the equity.

**Sarah Catanzaro** [15:52]
Y- y- you mean that, like, they'll offer them a loan to, to pay?

**Host** [15:56]
Uh, a buyback, uh, if, if, if it goes, um-

**Sarah Catanzaro** [16:00]
Yeah

**Host** [16:01]
... if you wanna sell it.

**Sarah Catanzaro** [16:02]
Yeah, but, but, but, but, but like the-

**Host** [16:03]
Because they have so much cash.

**Sarah Catanzaro** [16:04]
Like, but the thing though is that, like, the valuation is a made-up number. Like, valuation, until a company exits, it is an entirely made-up number. So, like, I could just be like, "Uh, you know what? The latent space pod, that is worth $5 billion," and we could agree.

Like, uh, we, like, I as an investor could say, like, "That is the price," and now, now the company is worth $5 billion. Like, do you think that, like, if you were to, to-

**Host** [16:28]
Yeah, it's not, it's not real. It's not, it's not actual-

**Sarah Catanzaro** [16:30]
It's not real

**Host** [16:30]
... uh, it's exacted in any volume.

**Sarah Catanzaro** [16:31]
Um, and given the, the funding amounts that they're raising too, like, if they spent that and they, you know, get acquired for less than, uh, that amount-

**Host** [16:40]
Yeah

**Sarah Catanzaro** [16:40]
... then, like, their teams are getting nothing. I wish people were kind of, like, more sensitive to this dynamic and thinking more about, like, what is the upside associated with the company, and, you know, more fundamentally, like, do I deeply believe in this vision?

'Cause I think, like, uh, joining companies because, like, they have a billion-dollar valuation, it's just, it's not the right way to choose a job.

**Host** [17:03]
I hear you. Okay, so there obviously we can go about that forever.

**Sarah Catanzaro** [17:05]
Oh, yeah.

**Host** [17:06]
There's, uh, and there's, there's a lot of... there's also some stuff with, like, cyclical, uh, funding and all that stuff, but, um, I, I, I do wanna be more relevant to engineers and researchers.

### World Models

**Sarah Catanzaro** [17:16]
Yeah.

**Host** [17:17]
Uh, what is... what are the, the themes that are, that are really strong, right? So one, one thing I'll point out is w- world, uh, world models-

**Sarah Catanzaro** [17:24]
Oh, yeah. Yeah

**Host** [17:24]
... just in general are a really strong bet. I would say, uh, so I have a... every NeurIPS, I go to this, like, group of researchers, and we take a vote on the top themes of the year. Everyone's extremely skeptical about world models.

I think it's a trailing indicator because LLMs have been so enormously successful. You're like, "I don't need anything else." I don't know if you have a take on world models or any other top theme of the year.

**Sarah Catanzaro** [17:46]
My, like, take on world models is that, like, we have not yet defined, like, what a world model is.

**Host** [17:51]
Oh, yeah, there's like three definitions right now.

**Sarah Catanzaro** [17:53]
Yeah. I think there's a lot of confusion about, like, what a world model is and therefore, you know, what it should be used for. Uh, we're already seeing, you know, plenty of, like, market potential for video models, including for things as, like, perhaps, like, banal as, like, uh, video editing.

I think, you know, we're already seeing some applications of world models to things like autonomous driving and potentially even coding, but again, it really hinges upon, like, how are you defining world models? And I think one challenge that people have seen is that, like, world models perhaps designed for, uh, one specific use case might not generalize to others.

So as an-

**Host** [18:31]
Yeah

**Sarah Catanzaro** [18:31]
... example of this, like, world models for, uh, like, video game generation might not, like, generalize to, like, factory settings or, or-

**Host** [18:40]
Yeah

**Sarah Catanzaro** [18:40]
... robotics. Um, I use the word might, like, strategically because I think, like, it is potentially a research problem-

**Host** [18:46]
And the other argument is that they work

**Sarah Catanzaro** [18:47]
... that might be figured out.

**Host** [18:47]
Yeah, so-

**Sarah Catanzaro** [18:48]
Yeah

**Host** [18:48]
... that's part of the General Intuition podcast that we did-

**Sarah Catanzaro** [18:51]
Yeah

**Host** [18:51]
... is that they have some evidence.

**Sarah Catanzaro** [18:52]
Yeah, yeah. I think, like, it is possible. It's just we're not there yet today.

**Host** [18:56]
Yeah.

### Personalization

**Sarah Catanzaro** [18:56]
A theme that I've been spending a lot of time thinking about is, uh, memory management and continual learning. I work with a lot of-

**Host** [19:05]
Same startup that I was thinking about.

**Sarah Catanzaro** [19:08]
Okay. The, the... I, I think I know what startup you're, you're thinking-

**Host** [19:11]
Yeah

**Sarah Catanzaro** [19:11]
... about as, as, as well. But I actually, like, I see, I see, like, a lot of market potential for, um, uh, memory management and continual learning. Uh, my interest in this is actually more driven by conversations with, uh, practitioners.

Personalization is so important right now. I think what we're seeing is that, like, a lot of AI application companies, they're growing really quickly, but they suffer from, you know, relatively low retention, relatively high churn. So, you know, if you're developing an app like Cursor, how do you ensure that your users don't, you know, switch over to, uh, uh, you know, Windsurf?

**Host** [19:49]
Windsurf.

**Sarah Catanzaro** [19:49]
Yes. Uh, or, you know, Cloud Code, or Cognition, or, or, uh, you know, whatever else, um, when they release new features.

**Host** [19:58]
Yeah. Cursor Rules isn't enough, right? Like, it's, it's, like, the shittiest form of memory.

**Sarah Catanzaro** [20:01]
Yeah, yeah.

**Host** [20:02]
But, and, and, and, you know, and it's great. Um, but yeah, I, I agree with that, but also it's like, as a... I, I've, I've publicly mused about this before where, like, uh, memorization... memory is very, uh, poorly implemented today in a lot of surfaces.

Like, even ChatGPT, I wouldn't say, like, people are particularly excited about. Okay. All right.

**Sarah Catanzaro** [20:21]
Yeah, yeah.

**Host** [20:22]
You feel stronger about it than I do.

**Sarah Catanzaro** [20:23]
Yeah, yeah. I mean, I, I, I wish ChatGPT had, you know, much better-

**Host** [20:27]
Yeah

**Sarah Catanzaro** [20:28]
... uh, memory completion

**Host** [20:28]
... this has both been the leading one.

**Sarah Catanzaro** [20:29]
Mm-hmm.

**Host** [20:30]
I don't know. Um, so a- and then I think, like, just in general it makes product management harder because what is the product? It's a combination of you plus memory, and like, when you have a bug, is it the memory or is it something core?

Um, a- and, and that's, as a user, especially if it's consumer, it's... there's gonna be zero patience for any of this.

**Sarah Catanzaro** [20:53]
I agree, but that said, like, consumers seem to be, like, tolerating products with, like, no implementation of memory today.

**Host** [21:01]
Yeah.

**Sarah Catanzaro** [21:01]
So I think, like-

**Host** [21:01]
Yeah, early adopter

**Sarah Catanzaro** [21:02]
... better is still-

**Host** [21:03]
Yeah

**Sarah Catanzaro** [21:03]
... probably better than, like, what, what, what exists now. Better is better than nothing, I guess.

**Host** [21:08]
Would you agree with the statement that basically, let's say, a key theme of 2026 is this personalization? I would call it kind of like the consumerization of, uh, AI, uh, in the same way that consumerization of, of enterprise was-

**Sarah Catanzaro** [21:20]
Yeah

**Host** [21:20]
... a trend, like, 10 years ago.

**Sarah Catanzaro** [21:21]
Yeah, I mean, I think that is a good way to putting it too. Like, I don't for, for, for what it's worth think, like, this is just a, like, consumer or prosumer phenomena. If you are an enterprise that is adopting, again, like, a Devon or Augment or something like that-

**Host** [21:35]
Yeah

**Sarah Catanzaro** [21:35]
... you probably also want your models to kind of, like, learn the, like, kind of-

**Host** [21:38]
So I, I'm not limiting-

**Sarah Catanzaro** [21:39]
Learn. Yeah.

**Host** [21:39]
Yeah.

**Sarah Catanzaro** [21:40]
Yeah.

**Host** [21:40]
Like, you start to... Uh, like k-factor, I had to explain what that is to so many founders and, you know, like, this, this... Like, if you're in normal SaaS, this is what you obsess over, and to AI founders they're like, "What do you mean growth just doesn't just show up?"

Like

**Sarah Catanzaro** [21:56]
Yeah, yeah. I mean, it has though. But, but, but, but I think, like, it has because, uh, for a while, uh, you know, AI has just felt magical.

**Host** [22:07]
Yeah.

**Sarah Catanzaro** [22:07]
But, like, now we're getting more accustomed to the magic, and it's no longer enough, and I think, you know, we need to, uh, revert to some of the, like, old tips and tricks for retaining people and, you know, bringing, bringing them in.

Personalization is one of them. Um, uh, I always kind of intermingle, like, memory and continual learning because I think, like, one interesting element of personalization is not just learning, you know, facts about your or your preferences, but, like, actually learning new skills from interactions with you and, you know, learning as the world changes.

Like, there are new versions of, uh, languages and frameworks and, you know, other repos that are coming out all the time. The world is changing all the time. Human intelligence is incredibly dynamic, and yet, like, uh, artificial intelligence is just so static today.

**Host** [22:57]
Yeah, yeah.

**Sarah Catanzaro** [22:57]
But, like-

**Host** [22:58]
So it must update weights-

**Sarah Catanzaro** [22:59]
Yeah

**Host** [23:00]
... for you.

**Sarah Catanzaro** [23:00]
But, but, but that also means that, like, it's an interesting kind of, like, systems problem because, like, if you must update weights then, like, you know, weights become stateful, and today, like, inference is not stateful. So, so, you know, I think, I think there's going to be, like, a lot of kind of fun, gnarly problems to figure out as we figure out things like-

**Host** [23:17]
Yeah

**Sarah Catanzaro** [23:17]
... personalization and continual learning.

**Host** [23:19]
And, and that, that's also a fascinating infrastructure problem because you have to load and unload and, uh, you know, cache and all the, all the, all the good stuff.

**Sarah Catanzaro** [23:25]
Yeah, exactly.

### RL Environments

**Host** [23:25]
Um, one more thing. I, I think we have time for one more take-

**Sarah Catanzaro** [23:28]
Yeah

**Host** [23:28]
... uh, on RL environments.

**Sarah Catanzaro** [23:30]
Yeah.

**Host** [23:30]
Huge topic. Uh, is it just a Docker container with some custom software loaded and logging stuff out? What are the good ones like, and what, what are the average ones like?

**Sarah Catanzaro** [23:42]
So I know I'm going on record on this, and, like, I'm actually okay to be wrong, but I think RL environments is just a fad. Um-

**Host** [23:50]
Oh, God. Oh, no.

**Sarah Catanzaro** [23:52]
Uh-

**Host** [23:52]
They're all, they're all fake? I mean, like, what... I mean, people are... Like, okay, the, the thing that I... makes me take it seriously, the labs I know are paying seven, eight figures for RL environments for other... Like, and they could build it in-house.

They're not, and I don't understand why.

**Sarah Catanzaro** [24:08]
I mean, they were paying seven to eight figures for, like, piss-poor data annotation too.

**Host** [24:14]
Yeah.

**Sarah Catanzaro** [24:14]
Uh, so, uh, like, and then data labeling before. Like, the labs have a lot of money. I think perhaps, like, RL environments could create some value in the short term, but I think to, to the point about, like, what makes a good RL environment, what makes a bad RL environment, I think the best RL environment is, is, you know, the, the real world.

Um, why would I, you know, want to, uh, buy a DoorDash clone when, like, I can just use, uh, logs and traces from, you know, DoorDash itself? It doesn't mean that we don't need to spend-

**Host** [24:50]
You can roll out in parallel-

**Sarah Catanzaro** [24:51]
Y- yeah

**Host** [24:52]
... faster.

**Sarah Catanzaro** [24:52]
I mean, I think, like, using the real world, using real apps as, like, RL, RL environment is in fact, like, the best thing, and this is what Cursor does. Like, they actually do use, uh, you know, real user activity on their platform to, you know, significantly, like, improve both their coding agents as well as tab, and I think that's one of the, the approaches that has, like, made the platform so compelling.

It doesn't... Like, you still need to figure out, like, the right rubrics. You still need to figure out, like, the right set of tasks. Uh, so there are some aspects of RL environment design, you know, at least as we're talking about it today, uh, that I think are going to remain incredibly relevant, but, like, just building a clone of an app I think is not that useful.

**Host** [25:36]
Yeah.

**Sarah Catanzaro** [25:36]
Yeah.

### Startup Bets

**Host** [25:37]
Okay. Yeah, that's, that is haunting. Um, we have maybe three minutes for any other stuff that you think, uh, about just the state of star- startups in general, state of funding.

**Sarah Catanzaro** [25:47]
Yeah. I, I, I... So, so maybe I can talk about, like, just the archetype startup that is, like, most exciting to me.

**Host** [25:53]
Yes.

**Sarah Catanzaro** [25:54]
I, I-

**Host** [25:54]
Press for startups.

**Sarah Catanzaro** [25:55]
Yeah, yeah. I love investing in, you know, infra tools, platforms, et cetera. And as we talked about with continual learning, I think, like, there will be opportunities for, like, new tools, platforms, and infra in the future. I've spent a lot of time thinking about, like, applications today.

Um-

**Host** [26:11]
Yeah

**Sarah Catanzaro** [26:12]
... and specifically, like, the relationship between research and applications. An example of this is, like, I think there were a lot of advances in RAG, and the biggest beneficiaries of these advances were the application companies for whom, you know, retrieval was a critical unlock.

So as an example of this, you know, like, uh, Harvey, Habia.

**Host** [26:35]
I knew you were gonna say Harvey.

**Sarah Catanzaro** [26:36]
They... Yeah. I mean, they, they, they have, like, really interesting RAG implementations. They have hired researchers, like really good researchers, uh, to kind of advance the state-of-the-art, and that enables them to build a better product. I feel this way very much about, like, rule following and customer support.

Rule following is, like, a hard research problem, but if you solve rule following, then you unlock, you know, better customer support. And I think, uh, you know, a lot of, uh, Sierra's success can be attributed to, like, their focus on this.

So I've been thinking about, like, uh, e-even for something like continual learning or memory, what is, like, the killer use case where you can either offer a dramatically better experience by having a good memory implementation, or you can do something that was just not possible today?

I think you can also think about this in the inverse. Like, uh, and often the best companies emerge in this way. They're like, "I'm trying to do this thing, but in order to actually do it, I need to solve this hard technical problem."

Uh, that, that, that's kind of like the story of Runway. Uh-

**Host** [27:39]
Mm.

**Sarah Catanzaro** [27:40]
I don't think they would've built models if they didn't have to. Uh, but I love that, that combination of, like, we're delivering something that is, like, better for consumers, better for prosumers, better for users, uh, but we're doing so by solving th-these, like, really gnarly research and engineering problems.

**Host** [27:57]
Yeah. I, I don't wanna, um... Yeah, go ahead. There's, there's so, there's so much that, that I wanna sort of dig into there, but we're short on time. Uh, thank, just thank you in general. Um, I don't know if, I don't know if you have, like, a, a, a general call to startups for, like, a page somewhere that you wanna point people to.

### Outro

**Sarah Catanzaro** [28:13]
Uh, Twitter. Or X, whatever it's called.

**Host** [28:16]
Yep.

**Sarah Catanzaro** [28:17]
Uh, yeah. You can find me-

**Host** [28:19]
Can't be AI without Twitter

**Sarah Catanzaro** [28:19]
... you can find me there.

**Host** [28:19]
Yeah.

**Sarah Catanzaro** [28:19]
Or on South Park.

**Host** [28:20]
Okay.

**Sarah Catanzaro** [28:21]
With the one-eyed dog. I'm easy to spot.

**Host** [28:23]
Oh, okay.

**Sarah Catanzaro** [28:23]
Yeah.

**Host** [28:24]
Well, thank you so much for your time. I know you gotta go, but, uh, appreciate it.

**Sarah Catanzaro** [28:26]
Of course. It was great seeing you, and thanks for having me.

**Host** [28:29]
Yeah. Thanks.

---

This library is powered by PodHood (https://podhood.com), the podcast website platform.
