# The AI Memory Problem: Why Long Context Isn’t Enough — Dan Biderman, Engram Co-founder & CEO

Latent Space · 2026-07-13

<https://addtry.com/92bf3a8a-6ed4-4699-926e-c42dab96fe0b>

Engram co-founder and CEO Dan Biderman tells hosts Allen Park and Sean that long context and RAG aren't enough for AI memory—continual learning and gradient-based weight updates are needed. He argues Engram's approach compresses company knowledge into 'cartridges' that let models reason with far fewer tokens, overcoming 'context rot' and the inefficiency of re-reading massive corpora. Biderman draws on his background in Israeli special forces and computational neuroscience to explain how training creates intuition beyond text retrieval, citing examples like Harvey's legal queries where holistic understanding beats search. He envisions personal AI weights that improve like Tamagotchis, driven by user-specific feedback loops, and stresses that token efficiency and intelligence are inseparable—doing more with less enables harder problems. Engram is hiring infrastructure engineers to deploy millions of continuously updated memories.

## Questions this episode answers

### What are Engram's "knowledge cartridges" and how do they differ from RAG?

Knowledge cartridges are compact model weights created by training on a corpus with gradient descent, compressing knowledge up to 1,000x. Unlike RAG, which retrieves text chunks, cartridges give the model a chef-like intuition, enabling it to reason holistically with fewer tokens, avoid context rot, and tackle queries where the whole is greater than the sum of its parts.

[10:02](https://addtry.com/92bf3a8a-6ed4-4699-926e-c42dab96fe0b?t=602000)

### Why does Engram's CEO believe long context and RAG will break down as enterprise data scales to trillions of tokens?

Dan Biderman argues that with trillion-token corpora, long contexts cause confusion (context rot) and RAG-based retrieval misses holistic insights. For example, answering "which M&A deals haven't we completed this year?" requires reading all client files—a task that can cost thousands of dollars in frontier model tokens for a simple employee query, making the current approach economically unsustainable.

[24:41](https://addtry.com/92bf3a8a-6ed4-4699-926e-c42dab96fe0b?t=1481000)

### What is Engram's vision for personal AI models that improve like Tamagotchis?

Engram aspires to give every user a personalized model or adapter that learns from their data and feedback, akin to nurturing a Tamagotchi. Users control their own weights, and direct feedback tightens the improvement loop. Over time, these models may run locally on devices, becoming experts in the user's knowledge, discerning which feedback to trust, and autonomously managing what to remember.

[27:24](https://addtry.com/92bf3a8a-6ed4-4699-926e-c42dab96fe0b?t=1644000)

## Key moments

- **[0:00] Intro**
  - [0:26] Engram announces its $98M seed round for AI memory technology.
- **[1:45] Military to AI**
  - [1:45] Dan Biderman transitioned from Israeli naval special operations to AI research founder.
  - [4:32] Israeli military culture gives young people responsibility and multiple career shots, creating resilient founders, says Dan Biderman.
- **[7:12] Context Bet**
  - [7:13] Q: Why did Engram focus on AI memory and continual learning?
  - [9:10] Engram creates 'cartridges'—compact knowledge capsules 1000x more compressed than raw text—to load into models.
  - [10:49] Just like a chef needs intuition beyond recipes, Engram trains models to develop true understanding of knowledge corpora.
- **[14:10] Data Scale**
  - [14:43] Within 18 months, enterprises will have trillions of tokens of internal data, causing context rot for LLMs, predicts Dan Biderman.
  - [17:49] Long-context models still suffer from 'context rot'—more data leads to confusion—requiring neural memory like Engram's, argues Dan Biderman.
  - [23:46] Dan Biderman: 'We're destroying prefill' by training knowledge into model weights to skip costly preliminary reads.
- **[24:31] Enterprise & Personal**
  - [25:02] A law firm query like 'which M&A deals haven't been completed?' costs thousands with RAG; Engram's training makes it holistic and cheap.
  - [27:31] Engram's vision: personal AI models that improve with use, like Tamagotchis that get smarter the more you interact.
- **[30:00] Memory Architecture**
  - [30:00] Q: How does Engram decide what knowledge lives in model weights versus in text?
  - [32:25] Engram aims for autonomous memory where models learn from implicit user feedback, getting better for each individual without manual labeling.
  - [35:22] Dan Biderman: 'Getting smarter means exerting less energy to solve increasingly harder problems.'
- **[38:03] Team & Hiring**
  - [38:23] Engram's research team pairs seasoned PhDs with 'cracked' young talent from Stanford and Berkeley, merging product focus with research depth.
  - [43:08] Q: What roles is Engram hiring for to build continual learning infrastructure?
- **[45:25] Efficiency**
  - [45:54] Doing more with less is the next AI paradigm, enabling longer horizon agents, says Dan Biderman.

## Speakers

- **Allen Park** (host)
- **Sean** (host)
- **Dan Biderman** (guest)

## Topics

Language Models

## Mentioned

Databricks (company), Engram (company), Harvey (company), Mosaic (company), ChatGPT (product), Fable (product), Llama (product)

## Transcript

### Intro

**Dan Biderman** [0:00]
I'm too absorbed in the cooking, man.

**Allen Park** [0:02]
Yeah.

**Dan Biderman** [0:02]
Sorry, I'm fried, um, the entire show I pretended to be the cooking expert here.

**Allen Park** [0:08]
Okay.

**Dan Biderman** [0:08]
It has to be pure crazy.

**Allen Park** [0:14]
Hey guys, welcome to the Latent Space Cooking Show, where we invite founders and researchers and let them cook. Today we have a very special guest, co-founder and CEO Dan Biderman of Engram.

**Dan Biderman** [0:25]
Nice to be here.

**Allen Park** [0:26]
Yes, thank you for coming. Um, I heard, or we got connected through Jack, who is saying that you're a very good cook. And again, congrats on the $98M seed round. Do you cook often? How frequent is it? I assume you're constantly researching and building and.

**Dan Biderman** [0:40]
Yeah. So I would say since I started the company, there's a little bit less cooking of this type, uh, cooking other things. But cooking is something I've been doing since I, I was a kid. My parents cook, my grandparents used to cook, and it's, it's a way for me to kind of relax and digest things and think things through.

**Allen Park** [0:58]
Great. Yeah. Well, I'm very excited. Today's a little different, where we're actually going to have Dan lead us through one of the recipes. Do you want to reveal to us what we're making?

**Dan Biderman** [1:07]
Yeah. So I was thinking we can make, um, beef meatballs with some vegetables and have them swimming in some kind of, uh, white sauce with white wine and, and, uh, chicken stock. It's a pretty nice Mediterranean type, uh, meatballs that I've kind of stumbled upon recently when visiting my parents in Tel Aviv and cooking stuff with what they had at home.

So I thought it would be nice to, to redo this with you and see how it feels.

**Allen Park** [1:32]
Great. Well, I'm very excited. Do you want to kick us off?

**Dan Biderman** [1:35]
We're going to fry the onion. I'm going to chop it, and I'm going to chop some garlic, and then we'll, uh, make, make them into balls and, and fry them.

**Allen Park** [1:42]
Fry them. Great. Okay.

**Dan Biderman** [1:43]
Yeah.

**Allen Park** [1:43]
Well, let's get started. We'll both do it in parallel.

### Military to AI

**Dan Biderman** [1:45]
Awesome.

**Allen Park** [1:46]
Um, but yeah, I guess to start, I'm kind of curious about your background. So you wanted to become a professor, if I recall correctly, and you also had experience in the Israeli military or special forces.

**Dan Biderman** [1:59]
Yeah.

**Allen Park** [1:59]
Um, how did that turn you into becoming a founder of a research.

**Dan Biderman** [2:03]
Yeah.

**Allen Park** [2:03]
AI research company?

**Dan Biderman** [2:05]
Yeah. So I guess, uh, it's, it's easy to connect the dots, uh, looking backward, but in real time, it was not. I didn't, uh, grow up planning to be this founder and a chef working on this hot topic in AI.

**Allen Park** [2:20]
Mm-hmm.

**Dan Biderman** [2:20]
Things came to it. Um, so I grew up in Tel Aviv and, uh, and as you said, uh, was an officer and, and working on special operations, specifically naval special operations, which was very entrepreneurial. It was, uh, basically finding and identifying crazy things we could do and, and prove to people that we should scale resources and, and people to actually do them.

So it was a lot of, a lot of researching, a lot of pitching, and then a lot of hard work to make things real.

**Allen Park** [2:51]
Yeah.

**Dan Biderman** [2:52]
Um, and then after I finished this, um, I went to university, um, and studied cognitive, uh, neuroscience in Israel, uh, in an interesting program where, um, actually a lot of Israeli novelists and professors and scientists and.

**Allen Park** [3:09]
Mm-hmm.

**Dan Biderman** [3:09]
And even chefs. You have Tamagotchi; the Israeli chef went to this program many years before me. Um, and so I went there, got really curious about the brain, curious about human behavior.

**Allen Park** [3:22]
Yeah.

**Dan Biderman** [3:22]
Curious about statistics. And then I went to, to New York to do a PhD in computational neuroscience, um, which is a community that's been one of those stubborn communities that cared about neural networks.

**Allen Park** [3:35]
Mm-hmm.

**Dan Biderman** [3:35]
For many decades before it was a hot topic.

**Allen Park** [3:38]
Yeah.

**Dan Biderman** [3:38]
Uh, here in the valley. Um, so I was there and studied with some statisticians and physicists and biologists trying to understand, um, how neural networks can be used to understand the brain.

**Allen Park** [3:50]
Mm-hmm.

**Dan Biderman** [3:51]
To understand animal behavior. And, uh, and then decided that I wanted to go deeper into the LLM world after seeing ChatGPT. So I started working at Mosaic, uh, where I worked on LoRA and met a bunch of great people.

And then from there, I went to work with Chris Ray at Stanford.

**Allen Park** [4:10]
Chris Ray, yeah.

**Dan Biderman** [4:10]
And Scott Linderman, who both are co-founders of my current company.

**Allen Park** [4:14]
Yeah.

**Dan Biderman** [4:14]
And where I met my co-founder Sabri and my other co-founders Jack and Jesse, uh, who worked on very similar topics from their labs in Cornell and Berkeley.

**Allen Park** [4:25]
Yeah. No, it seems like a very, very research-focused and also serendipitous. Going back a little on your background.

**Dan Biderman** [4:32]
Yeah.

**Allen Park** [4:32]
Is there something about the Israeli special forces? Because I think it was the Wiz co-founders,right, who were also ex-special forces.

**Dan Biderman** [4:39]
Yeah. Uh, we worked pretty closely with Assaf, uh, Rapa Poort, the Wiz CEO.

**Allen Park** [4:44]
Mm-hmm.

**Dan Biderman** [4:44]
Who's, uh, uh, a mentor to me. Um, and, um, yeah, there's interesting things going on. Uh, there, I would say the, the most important thing is actually, uh, when you're 18, it's joining the intelligence or taking on responsibility.

Uh, the interesting bit about it is that you're not just in a student mode. You're not just heads down doing exams and projects. Rather, you interact with, with grownups, and you have resources, and you argue your, your for your, uh, budgets and things like that.

**Allen Park** [5:18]
Interesting. Oh, okay.

**Dan Biderman** [5:19]
Um, and it forces you to kind of build a little bit of social skills, a little bit of maturity.

**Allen Park** [5:25]
Mm-hmm.

**Dan Biderman** [5:25]
Um, which, which I benefited from.

**Allen Park** [5:28]
Yeah.

**Dan Biderman** [5:28]
Um, and, um, you know, there's a reason many countries in the world choose not to have everyone go to the army. It's a long time, and it's, and it's, uh, uh, a package deal. And I was, uh, in the Israeli army in different years, not the ones that are the recent ones.

**Allen Park** [5:46]
Gotcha.

**Dan Biderman** [5:46]
Um, so it's not like I recommend everyone goes to the army, but I would say that it forces you to kind of think about other s other, other aspects of your personality, not just the intellectual ones. And also generally, I would say Israel as a culture is a place where, um, you can basically you get multiple shots at goal.

**Allen Park** [6:06]
Mm-hmm.

**Dan Biderman** [6:06]
If you're not the best in high school, um, you still might have a, a good position in the military. And if you're really good, more doors open up for you for university. And even if you didn't do anything super interesting in, in military and got to university and found out you have good and deep inclinations in science and engineering, you can be the most successful person.

And we have some of those Israelis who are, um, leaders, uh, in various, uh, academic and technological places.

**Allen Park** [6:34]
Oh, that's great.

**Dan Biderman** [6:34]
Um, so I would say, yeah, it's mostly like you can basically.

**Allen Park** [6:38]
The culture.

**Dan Biderman** [6:39]
The cards are dealt multiple times.

**Allen Park** [6:40]
Yeah.

**Dan Biderman** [6:41]
More emphasis on social aspects, more emphasis on being mature, on being a team player. Uh, but again, it comes at a price that you get, uh, you, you're older when you get to things. Um, so I started college when I was 24, which is, uh, a, a middle PhD in, in an American, uh, university.

But yeah.

**Allen Park** [7:05]
That makes sense. I think it's great how it gives you a lot of shots on goal and is a very beautiful culture. I know fast forwarding to today, back to Engram.

### Context Bet

**Dan Biderman** [7:13]
Yeah.

**Allen Park** [7:13]
What kind of prompted you to focus on this issue with context? Were there specific breakthroughs? Was it, you know, the performance of open source models or even just the issue of Context Rot, long horizon agents, or was it just that you're very passionate about this area as a whole?

**Dan Biderman** [7:30]
Yeah. I'll answer that in a second. Can we.

**Allen Park** [7:34]
Yeah.

**Dan Biderman** [7:34]
Can we heat up some.

**Allen Park** [7:35]
Yes. Let's start getting going.

**Dan Biderman** [7:37]
Oil here to fry up the.

**Allen Park** [7:39]
Great. So we can turn on the Impulse stove.

**Dan Biderman** [7:42]
So yeah, I, my, my PhD was focusing on, on, on a field that's not super in vogue today, but I think will become super crucial again, which is semi-supervised learning.

**Allen Park** [7:56]
Hmm. Okay.

**Dan Biderman** [7:57]
Which is you have very little data. You don't have infinite trillion tokens of data. You have very infinite, you have a very small set of examples. And from them, you want to extrapolate and learn efficiently for, uh, and, and achieve things more, more generally.

**Allen Park** [8:13]
Mm-hmm.

**Dan Biderman** [8:14]
So how to be data efficient. So my whole academic work was basically all of my papers had the same plot. We're on the X-axis. I had.

**Allen Park** [8:24]
Mm-hmm.

**Dan Biderman** [8:24]
Some, some cost, some resource, and on the Y-axis, I had, um, I had, uh, accuracy, sort of Pareto curve. Uh, and that's been my, my upbringing.

**Allen Park** [8:35]
Mm-hmm.

**Dan Biderman** [8:36]
Efficiency, how to do more with less, um, salt.

**Allen Park** [8:39]
Yeah.

**Dan Biderman** [8:40]
Yeah. You can put some salt. Um, and I went to, to Chris Ray's lab, um, and.

**Allen Park** [8:45]
At Stanford,right?

**Dan Biderman** [8:46]
At Stanford. And there I worked on agents, on ideas called minions, uh, and, and asked kind of the question of how can we have agents interface with very large, uh, corpora of knowledge in a way that's economical, and when we introduce some patterns that involve sub-agents, uh, a bit before it was, uh, popular.

**Allen Park** [9:08]
Oh, sorry. We can also put this side of your glass.

**Dan Biderman** [9:09]
And from that perspective of efficiency and of cost, um, we started asking ourselves, actually.

**Allen Park** [9:17]
Yeah.

**Dan Biderman** [9:17]
Um, is there a more efficient way to interface with data that's not just agentic orchestration and things like this? Can we actually use the, the magic of training to, um, have models that are efficient, faster, and, and, and can operate with fewer tokens?

And some of these ideas kind of were developed in parallel by some of us, uh, Jesse, Sabri, and myself.

**Allen Park** [9:40]
Yeah.

**Dan Biderman** [9:40]
Uh, and the common theme was that if you have a very large corpus of knowledge, um, and you allow the model, you give the model time to study it in advance to basically ask itself questions, give itself quizzes, try to solve problems, and then train it using gradient descent in the same way you would pre-train a model.

**Allen Park** [10:01]
Mm-hmm.

**Dan Biderman** [10:02]
Um, then you can create very compact representations of your context that can later on, you can load them. We call those cartridges.

**Allen Park** [10:10]
Mm-hmm.

**Dan Biderman** [10:10]
These are like capsules of knowledge you can load in and out of the model that are like a brain state that describes the model's world and the corpus in a way that's going to be 1,000x more compressed. And when you use them, you can basically, uh, use far fewer tokens, be less confused, more accurate, and things like this.

So we got into.

**Allen Park** [10:30]
And these are task-specific cartridges?

**Dan Biderman** [10:32]
So they can be corpus-specific, like documents inside a company. They can be task-specific, like you want to learn a certain skill. But the idea is that in many cases, it makes sense to go beyond textual representations.

**Allen Park** [10:45]
Yeah.

**Dan Biderman** [10:45]
So when we're cooking now and we have all these, like, great textbooks.

**Allen Park** [10:49]
Mm-hmm.

**Dan Biderman** [10:50]
Our cookbooks here, Samina Swat and, you know, a book from French Laundry. These are amazing books that I.

**Allen Park** [10:55]
Yeah.

**Dan Biderman** [10:55]
I respect. But to say that because we have the books and we can read them, to say that we deeply know how to cook to the same level of Samina Swat or, or, or Thomas Keller.

**Allen Park** [11:05]
Yeah.

**Dan Biderman** [11:06]
It's incorrect,right?

**Allen Park** [11:07]
Yeah.

**Dan Biderman** [11:07]
We can read those things and we can repeat all their steps in a very robotic way. And we can come in, current LLMs are like coming into the kitchen first time every time, reading the textbook, uh, cooking the dish, measuring everything.

Uh, but they don't have the intuition of, uh, of a chef that's pinching salt and kneading, kneading dough and things like this. So the kind of thing we're after with this kind of, uh, training and creating those cartridges is this kind of intuition in the models that goes beyond notes and recipes to the kind of intelligence that allows then, allows you to come, come up with the next recipe, thing that hasn't been explored before to do the next, uh, move, the next extrapolation.

**Allen Park** [11:49]
Yeah. I guess double-clicking on building this intuition, could you kind of contextualize it more on how it would be different from extracting, let's say, using the cooking example? If you get all the useful notes and useful sections from a cookbook, that will help you actually understand all the complexities of making a dish.

Just providing it, I guess that would be similar to RAG, getting just those chunks.

**Dan Biderman** [12:13]
Yeah.

**Allen Park** [12:13]
And then understanding that, I guess, in context and then providing an output.

**Dan Biderman** [12:18]
Yeah. So this is, this is an excellent method. Um, we do not take the bet that I'm putting two eggs.

**Allen Park** [12:24]
Yeah.

**Dan Biderman** [12:24]
In here, if that's okay. Um.

**Allen Park** [12:26]
And is this enough grated carrot and zucchini?

**Dan Biderman** [12:28]
Yeah. Excellent. You can put it in here. Um.

**Allen Park** [12:31]
Yeah.

**Dan Biderman** [12:33]
So we're not taking the bet that no ne no notes need to be taken,right?

**Allen Park** [12:39]
All of it.

**Dan Biderman** [12:39]
Yeah. Put all of it inside.

**Allen Park** [12:40]
Mm-hmm.

**Dan Biderman** [12:41]
Um, all the greatest chefs in the world, they have notes and they have books and they.

**Allen Park** [12:46]
Yeah.

**Dan Biderman** [12:46]
They have diaries and they document what worked and what didn't work with their experiments.

**Allen Park** [12:50]
Mm-hmm.

**Dan Biderman** [12:50]
But they also have brains that they also have hands and, and, uh, and, uh, a tongue.

**Allen Park** [12:58]
Yeah.

**Dan Biderman** [12:58]
That can remind them how something tastes and what worked and what failed and what was easy and what was hard. So, uh, in all of our work, we never say that textual representations are useless.

**Allen Park** [13:11]
Mm-hmm.

**Dan Biderman** [13:11]
We basically use them all the time and we construct those wikis and knowledge bases and things like this.

**Allen Park** [13:17]
Yeah.

**Dan Biderman** [13:18]
Um, at the same time, what we say is that the layer above those, the learn the layer of intuition and learning, uh, that is in the form of numbers of parameters, um, is, is the, the full experience of the human chef.

The best chef in the world is all the notes combined with a nervous system that reads those notes.

**Allen Park** [13:38]
Gotcha.

**Dan Biderman** [13:38]
And can implement them and innovate them and maybe don't re-return to the same notes over and over again. There are some dishes where they don't need to, to reread them.

**Allen Park** [13:45]
Mm-hmm.

**Dan Biderman** [13:46]
And the current problem in, uh. So I was saying that, um, we want, we want the best of both worlds.

**Allen Park** [13:54]
Mm-hmm.

**Dan Biderman** [13:54]
Every knowledge worker, if they can't write notes and they cannot document the events of the day, they would be, uh, in a, a disadvantage.

**Allen Park** [14:02]
Mm-hmm.

**Dan Biderman** [14:02]
But, uh, if you wipe their brain every evening, they would also be at a severe disadvantage.

**Allen Park** [14:06]
Yeah.

**Dan Biderman** [14:07]
Um, so we want to have the best of both worlds. And, uh, the thing is that textual representations can take you, uh, a, a long way. Uh, but the thing we're thinking about is like, look, the, at the rate at which knowledge is being created.

### Data Scale

**Allen Park** [14:22]
Mm-hmm.

**Dan Biderman** [14:23]
With now agents working on behalf of knowledge workers, creating artifacts, code, documents, uh, presentations, I think that people don't fully comprehend the size of the.

**Allen Park** [14:37]
Yeah.

**Dan Biderman** [14:37]
Of the knowledge workspaces they will deal with in 18 months.

**Allen Park** [14:42]
Yeah.

**Dan Biderman** [14:43]
Um, in 18 months, many companies would have, uh, maybe trillions of tokens, which.

**Allen Park** [14:50]
Of internal company.

**Dan Biderman** [14:52]
Of internal company data, proprietary data. I'm talking about like maybe trillions. It sounds exaggerated, but I don't think it's an impossibility if they're really AI native.

**Allen Park** [14:59]
Yeah.

**Dan Biderman** [15:00]
And I think, and, and what are trillions of tokens? Uh, like when I was at Mosaic two, three years ago, like we call this internet scale data, pre-training data.

**Allen Park** [15:10]
Mm-hmm.

**Dan Biderman** [15:10]
So imagine every company has data that it's basically internet scale data.

**Allen Park** [15:14]
Um, yeah.

**Dan Biderman** [15:17]
And I might need, uh, salt.

**Allen Park** [15:18]
Yeah, we'll have tea.

**Dan Biderman** [15:19]
And pepper and the bread crumbs.

**Allen Park** [15:21]
Yeah.

**Dan Biderman** [15:21]
For this. And we don't have a sink here, so hands will be a tiny bit dirty.

**Allen Park** [15:26]
No worries.

**Dan Biderman** [15:27]
The audience here is forgiving. Um.

**Allen Park** [15:31]
Gotta get your hands dirty.

**Dan Biderman** [15:32]
Yeah.

Awesome. I guess the principle here I once heard from a chef is that a ratio of one to one of meat with everything else is usually a healthy way to make meatballs. So I think we're, we're there.

**Allen Park** [15:48]
Yeah, that's fair. We are on salt and pepper.

**Dan Biderman** [15:50]
Um, nice meatballs. They're a bit warm with those fried, um.

**Allen Park** [15:55]
Oh. We're back with the fix.

**Dan Biderman** [15:58]
We're back.

**Allen Park** [15:59]
We're back with the fixed meatballs.

**Dan Biderman** [16:01]
With the fixed meatballs. It's, it's team building activity here.

**Allen Park** [16:05]
Yes.

**Dan Biderman** [16:06]
Oh, yeah. Let's put those spices that you brought.

**Allen Park** [16:08]
Yeah.

**Dan Biderman** [16:09]
I trust you on, on these ones. Cumin, let's put a little bit, not a ton.

**Allen Park** [16:14]
Yeah. We won't dump it like last time.

**Dan Biderman** [16:16]
Cumin is my wife's favorite spice.

**Allen Park** [16:18]
Oh, yeah. Cumin's great.

**Dan Biderman** [16:20]
I have mixed feelings about it, but.

**Allen Park** [16:21]
Oh, okay.

**Dan Biderman** [16:22]
Yeah. It's good. Yeah.

**Allen Park** [16:23]
Good.

**Dan Biderman** [16:24]
Good. So we have the meatballs. Now is the time to make the balls.

**Allen Park** [16:28]
So we were talking about what were we talking about?

**Dan Biderman** [16:34]
I'm.

**Allen Park** [16:34]
I think I'm just.

**Dan Biderman** [16:35]
I'm too absorbed.

**Allen Park** [16:36]
Oh.

**Dan Biderman** [16:36]
I'm too absorbed in the cooking, man.

**Allen Park** [16:38]
Yeah.

**Dan Biderman** [16:38]
I'm sorry. I'm.

**Allen Park** [16:39]
No, it's good. You were.

**Dan Biderman** [16:39]
You were pitching things here.

**Allen Park** [16:41]
You're talking about how companies will have a corpus of data that will be at the scale of the internet?

**Dan Biderman** [16:47]
Yeah. Yeah. And maybe today.

**Allen Park** [16:49]
Maybe today.

**Dan Biderman** [16:49]
Some they don't have it. And maybe today you can go relatively far.

**Allen Park** [16:54]
Yeah.

**Dan Biderman** [16:54]
With textual representations. They're also interpretable. They're very good.

**Allen Park** [16:58]
Mm-hmm.

**Dan Biderman** [16:59]
Um, but I just think at a certain scale, even those textual representations will be hard to make.

**Allen Park** [17:04]
Okay.

**Dan Biderman** [17:04]
If you have trillion tokens, how do you create exactly a wiki or an index of those that you keep updated all the time? How big is this?

**Allen Park** [17:11]
Yeah.

**Dan Biderman** [17:11]
Um, knowledge base that you create, how expensive will it be to process it with frontier models that know nothing about your, your company?

**Allen Park** [17:19]
So is it just that using frontier models will be expensive because that's the start from scratch every time? Is that the main?

**Dan Biderman** [17:25]
So there's the element of expensive because you reread more things, you consume more to tokens. That's one.

**Allen Park** [17:30]
Yeah.

**Dan Biderman** [17:31]
But two is like for the agentic tasks of 18 months from now, inside those major repositories of knowledge and asking the models more and more things and under specified ways.

**Allen Park** [17:42]
Mm-hmm.

**Dan Biderman** [17:42]
I suspect that the accuracy, uh, of the models would, would go down. The, the phenomenon of context rot,right?

**Allen Park** [17:49]
Yeah.

**Dan Biderman** [17:49]
The model has to read more. It will be less accurate. And we know this, and it will remain the same thing at even at the 10 million, uh, context window scale.

**Allen Park** [17:58]
Okay. That makes sense. Okay. Let's go wash our hands real quick.

**Dan Biderman** [18:00]
Yeah.

**Allen Park** [18:00]
And we'll beright back.

**Dan Biderman** [18:02]
Great.

**Allen Park** [18:02]
And we're back. Hands are all clean. So now what's the next thing? Just frying the meatballs?

**Dan Biderman** [18:06]
Um, yeah. Let's fry those meatballs.

**Allen Park** [18:07]
Okay. Do you want to turn this one?

**Dan Biderman** [18:09]
Yeah.

**Allen Park** [18:09]
You just press it and then turn it. Very intuitive.

**Dan Biderman** [18:12]
And we'll put some oil.

**Allen Park** [18:13]
Okay. We could put these in and.

**Dan Biderman** [18:16]
Yeah.

**Allen Park** [18:17]
While we do that, I guess more so on the question of like long horizon agents.

**Dan Biderman** [18:22]
Yeah.

**Allen Park** [18:23]
Um, what's the main issue that Engram is trying to solve? Soright now, if I were to steel man the opposing view, um, maybe company, internal company data isn't actually big enough where I'd need to put it into the weights.

Like what's the issue with just having RAG or, um, having specific models or even cheaper models since open source models are also very performant to handle a lot of tasks?

**Dan Biderman** [18:47]
Yeah. So I would say like, um, a thorny question in the research community is, can you come up with an example where only in weights training would work, where in-context learning will fail?

**Allen Park** [19:01]
Mm-hmm.

**Dan Biderman** [19:02]
And it turns out it's very hard to devise such examples. And for every example you give, someone can ask, well, what happens if the, the next, uh, the next fable has a 10 million context window?

**Allen Park** [19:16]
Mm-hmm.

**Dan Biderman** [19:16]
Um, and the way I see it in kind of my scientific upbringing, I see all of these questions of, of continual learning and memory as questions of long context, uh, in disguise.

**Allen Park** [19:28]
Yeah.

**Dan Biderman** [19:28]
Um, if the models could see a whole, whole company data and in principle would have this infinite context window, what then is the limitation?

**Allen Park** [19:37]
Mm-hmm.

**Dan Biderman** [19:38]
So the limitation is twofold. One is that we know even at very small scales that the more context you feed to the model, the more confused it gets.

**Allen Park** [19:47]
Okay.

**Dan Biderman** [19:47]
It's called the context rot. So you can feed in a certain, uh, number of tokens into the model and not get an error, but it doesn't mean that the model can reason in, in a, in a holistic way about them.

That's one thing. And.

**Allen Park** [19:59]
And what's the problem with compaction? Is compaction not as useful, would you say?

**Dan Biderman** [20:03]
So I think compaction is, is improving by the day.

**Allen Park** [20:07]
Yeah.

**Dan Biderman** [20:07]
And compaction for those in the audience who don't know what it is, it's like models actually managing their own context, uh, evicting certain tokens and, and keeping others. Uh, compaction works. Uh, that too, when you go into a longer horizon, compaction by definition is lossy.

**Allen Park** [20:24]
Mm-hmm.

**Dan Biderman** [20:24]
You discard some and keep other things. And I think it's, it's a, it's a correct way to, to go. Uh.

**Allen Park** [20:31]
Yeah.

**Dan Biderman** [20:31]
But it's, it's also like very deterministic.

**Allen Park** [20:34]
Yeah.

**Dan Biderman** [20:34]
Either you're in or you're out. And the current versions of compaction, uh, show these kinds of, uh, uh, issues when, when very deep into the session you can get, um, confused and you can get, um, forgetful.

**Allen Park** [20:49]
Mm-hmm.

**Dan Biderman** [20:50]
So we think compaction will be part of the story.

**Allen Park** [20:52]
Okay.

**Dan Biderman** [20:52]
We think another part of the story is some sort of neural memory trace, which too is a lossy thing. It evicts some and, and keeps some, but not in the text, uh, representation and the weights representation.

**Allen Park** [21:05]
Okay.

**Dan Biderman** [21:05]
But I would say the main thing we're trying to solve or the main thing for which continual learning is needed, one is token efficiency and cost, which is a major issue that wasn't actually an issue when we started the company late last year and became moreurgent now.

And problem number two is if you can do, uh, if you can do the same thing with, with fewer resources, when you scale up to very large resources, suddenly you can take on tasks that were previously not possible.

**Allen Park** [21:35]
Yeah.

**Dan Biderman** [21:36]
Way more long horizon.

**Allen Park** [21:37]
Okay.

**Dan Biderman** [21:37]
Way more, uh, adaptive. And we're not quite there yet. Uh, we're focusingright now on, on the first component, which is getting these models to reason on large context with fewer tokens and do it in a way that's, that's more, uh, uh, le-less confused.

**Allen Park** [21:57]
Yeah.

**Dan Biderman** [21:57]
But we think that eventually, uh, part of the.

**Allen Park** [22:01]
Yeah. I'll take pictures. Thank you.

**Dan Biderman** [22:03]
Part of the, part of the solution for very hard tasks in, in science and engineering and defense and all that stuff will involve some form of gradient-based updates.

**Allen Park** [22:15]
Mm-hmm.

**Dan Biderman** [22:16]
Um, during, uh, during, uh, doing these long horizon tasks.

**Allen Park** [22:20]
Gotcha.

**Dan Biderman** [22:20]
Uh, some people call this, uh, test time compute or test time training. Um, these are just different names for the same thing. I think we have a, a proof of existence.

**Allen Park** [22:31]
Mm-hmm.

**Dan Biderman** [22:31]
From pre-training that, um, you can pack a lot of information in very few numbers.

**Allen Park** [22:38]
Yeah.

**Dan Biderman** [22:38]
Very efficiently. Um.

**Allen Park** [22:39]
Mm-hmm.

**Dan Biderman** [22:40]
So the examples we like to give is that, uh, if you take a Llama 70B model and you load one article from Wikipedia, which is a few tens of kilobytes, and you have the model read this, uh, the brain state of the model when reading this few tens of kilobytes is like 80 gigabytes.

**Allen Park** [22:57]
Mm-hmm.

**Dan Biderman** [22:57]
80 gigabytes on, on the HBM of the GPU.

**Allen Park** [23:00]
Yeah.

**Dan Biderman** [23:01]
It's an insane amount. And the entire set of parameters of this model would be like 140 or, or, or so gigabytes.

**Allen Park** [23:07]
Mm-hmm.

**Dan Biderman** [23:07]
FP16. So those 100 or so gigabytes with some distortion represent the entire internet. And this one article about Taylor Swift is like same order of magnitude memory consumption on the GPU. So it's highly memory inefficient.

**Allen Park** [23:24]
Yeah.

**Dan Biderman** [23:24]
That's a systems problem. That's the, the KV cache monstrosity that the smartest people in the world are trying to solve from the chip side of things and from the, from the software and kernel side of things. Um, but yeah.

So I would say another way to look at what we're doing from that angle, from the systems angle, is basically, uh, if we, if we can get a bit more technical.

**Allen Park** [23:46]
Yeah.

**Dan Biderman** [23:46]
Instead of doing those, uh, prefills where the model's just reading and reading and reading a corpus, uh, we are kind of like the we're destroying prefill. We're scaling training compute in some other time so we can load the thing into the model and can just immediately start decoding.

**Allen Park** [24:03]
Okay.

**Dan Biderman** [24:03]
Or prefill a little bit. And this kind of goes hand in hand with trends in how data centers are built out, this aggregating prefill and decode and doing this on different specialized, uh, cards.

**Allen Park** [24:15]
Mm-hmm.

**Dan Biderman** [24:15]
And, uh, and this is part of our.

**Allen Park** [24:17]
Yeah.

**Dan Biderman** [24:17]
Part of our initial interest in the thing.

**Allen Park** [24:20]
Mm-hmm.

**Dan Biderman** [24:20]
Okay. So it seems like we're mostly ready in here. I think it's a good time for our.

**Allen Park** [24:27]
Have the wine.

**Dan Biderman** [24:27]
White wine.

**Allen Park** [24:28]
Okay.

**Dan Biderman** [24:28]
On both of them.

**Allen Park** [24:29]
Great.

**Dan Biderman** [24:29]
Uh, maybe.

**Allen Park** [24:30]
Do you wanna do that?

**Dan Biderman** [24:30]
Yeah.

**Allen Park** [24:31]
Yeah.

**Dan Biderman** [24:31]
Yeah.

### Enterprise & Personal

**Allen Park** [24:32]
And do you have any tangible examples on this when you're talking about.

**Dan Biderman** [24:35]
Yeah.

**Allen Park** [24:36]
Your approach.

**Dan Biderman** [24:36]
So if you think about, um,

let's think, for example, about like an enterprise, uh, a firm.

**Allen Park** [24:44]
Mm-hmm.

**Dan Biderman** [24:45]
Or, um, investment banking, a law firm or investment banking, they can have many client matters, uh, many.

**Allen Park** [24:53]
Like a lot of knowledge work.

**Dan Biderman** [24:55]
A lot. Yeah. So many client matters. They have various clients, clients do, uh, financing, mergers and acquisitions and things like this.

**Allen Park** [25:02]
Yeah.

**Dan Biderman** [25:02]
And take loans and do deals.

**Allen Park** [25:03]
You wanna put this on or not yet?

**Dan Biderman** [25:05]
Um, no, not yet.

**Allen Park** [25:05]
Okay.

**Dan Biderman** [25:06]
Will you wanna let it, uh, let it.

**Allen Park** [25:09]
Let it reduce a little? Okay.

**Dan Biderman** [25:10]
Yeah. And so.

**Allen Park** [25:11]
Yeah.

**Dan Biderman** [25:12]
For example, these are the kinds of things we, we work with Harvey.

**Allen Park** [25:14]
Mm-hmm.

**Dan Biderman** [25:15]
Very, very large file systems. Uh, and there's many queries that agents might run into. Either humans ask them or the agents have to solve them.

**Allen Park** [25:23]
Mm-hmm.

**Dan Biderman** [25:23]
Which are these kinds of like ambient hard questions that are not easily searchable with RAG. For example, if you wanna ask like, which M&A deals haven't we completed this year? So to actually solve this problem, you have to go client matter by client matter.

**Allen Park** [25:40]
Yeah.

**Dan Biderman** [25:40]
Read all the files. You can't read in any place that it was not completed. You have to take the gist.

**Allen Park** [25:45]
You have to understand.

**Dan Biderman** [25:46]
The, the thing hasn't the loop hasn't been closed.

**Allen Park** [25:48]
Yeah.

**Dan Biderman** [25:48]
Thing hasn't been completed. And now you can solve these tasks with, with frontier models and compaction.

**Allen Park** [25:53]
Mm-hmm.

**Dan Biderman** [25:53]
And when you ask them to do so, they will consume thousands of dollars for queries that we think are harmless.

**Allen Park** [25:58]
Mm-hmm.

**Dan Biderman** [25:58]
That every employee in the company would be able to answer.

**Allen Park** [26:01]
Mm-hmm.

**Dan Biderman** [26:01]
So this is just like one example, but these kinds of holistic things that are, you know, you can get the gist if you read everything, but you can't find a thing in one thing. The kind of queries where the whole is greater than the sum of its parts.

So that's where, uh, this kind of magic of training, uh, comes in. And that's the, the magic that, that, you know, Ilya and others.

**Allen Park** [26:22]
Yeah.

**Dan Biderman** [26:23]
Have shown us with pre-training,right?

**Allen Park** [26:24]
Yeah.

**Dan Biderman** [26:24]
You learn the entire web in these sets of weights, and suddenly the model can infer things, can generalize, can interpolate and extrapolate to new things. And that's the kind of knowledge we wanna do.

**Allen Park** [26:35]
Mm-hmm.

**Dan Biderman** [26:35]
And it, and it's not a, not a coincidence that we're training on the entire web and we're not just doing RAG over the entire web.

**Allen Park** [26:43]
Yeah.

**Dan Biderman** [26:44]
Or putting it in the catalog and renewing it in time. Because we do think that learning from a lot of knowledge somehow creates these associations in the model.

**Allen Park** [26:51]
Gotcha.

**Dan Biderman** [26:51]
So some of it is concrete problems of now. Other parts of it are bets that in 18 months from now, the scale of the data will require the methods that we know from pre-training work.

**Allen Park** [27:03]
Gotcha. And so are you doing parameter efficient fine-tuning for.

**Dan Biderman** [27:07]
Yeah.

**Allen Park** [27:07]
Specific companies like LoRA based off of the corpus?

**Dan Biderman** [27:10]
Yeah. Yeah. So our, our ambition in the long term.

**Allen Park** [27:12]
But also start to write?

**Dan Biderman** [27:13]
Yeah.

**Allen Park** [27:14]
Yeah.

**Dan Biderman** [27:14]
So our ambition in the long term, uh, is that every person has a, has a model or, or a part of the model or a set of weights.

**Allen Park** [27:24]
Mm-hmm.

**Dan Biderman** [27:25]
Uh, that.

**Allen Park** [27:27]
Yes, sir.

**Dan Biderman** [27:28]
That represents their knowledge, their expertise.

**Allen Park** [27:31]
Yeah.

**Dan Biderman** [27:31]
Learns from them. That the more time they spend with the model, the better it gets for them. The more data they give the model, the better the model is for them.

**Allen Park** [27:39]
Yeah.

**Dan Biderman** [27:39]
They control those sets of weights. It's theirs, um.

**Allen Park** [27:43]
And the rights?

**Dan Biderman** [27:44]
Uh, and they're incentivized to let, let, let's let it boil over, you know?

**Allen Park** [27:48]
Okay. Yeah.

**Dan Biderman** [27:48]
Um, so they're incentivized to we can, we can put the spices up.

**Allen Park** [27:53]
Mm-hmm.

**Dan Biderman** [27:53]
So they're incentivized to, to make it better in the same way that you would work with, with a Tamagotchi. The more you nurture it, the happier it is. And that's the kind of thing we would like to create with the model.

So that's the ultimate continual learning. And we think on the extreme scale, on the single human scale.

**Allen Park** [28:08]
Mm-hmm.

**Dan Biderman** [28:08]
Um, turns out this is not just a research problem. That's a major infra problem. And in the long, long term, I do think these things will actually run on people's devices.

**Allen Park** [28:17]
Yeah.

**Dan Biderman** [28:17]
And, uh, we're seeingright now the, the new hardware on personal computers is already, uh, you know, soon approaching the ability to run inference on close to trillion parameter models.

**Allen Park** [28:29]
Mm-hmm.

**Dan Biderman** [28:30]
Uh, which will be very interesting for that kind of personalization. But in the shorter term, we do think that, uh, the kind of large corpora of knowledge, very dense, with a lot of expertise, uh, can be found in the enterprise.

**Allen Park** [28:43]
Mm-hmm.

**Dan Biderman** [28:43]
Um, that's where people spend most of their time. That's where AI is, is actually being used the most.

**Allen Park** [28:49]
Yeah. And enterprises.

**Dan Biderman** [28:50]
Um, so we, we, we go there and that's our bet. And there we're, we're our bet is that parameter efficient fine-tuning methods like LoRA, like cartridges, like memory layers, different things we, we contributed to as, as a research team, um, can actually represent knowledge and can be combined with other methods that are like context management, traceable methods that, that people can, can use and understand and, and audit.

**Allen Park** [29:15]
Yeah.

**Dan Biderman** [29:16]
So it's a combination.

**Allen Park** [29:17]
Gotcha.

**Dan Biderman** [29:18]
So again, it's like, it's like the chef that has the cookbook and it has the recipes, has the diary, but also has the nervous system that learns from every session. Uh.

**Allen Park** [29:28]
Gotcha.

**Dan Biderman** [29:28]
So it's boiling here. It's all.

**Allen Park** [29:30]
Yeah.

**Dan Biderman** [29:30]
Pretty efficient here.

**Allen Park** [29:32]
Yeah. Very powerful stove.

**Dan Biderman** [29:34]
Yeah.

**Allen Park** [29:36]
Very powerful. And then medium, low, then you wanna put this on?

**Dan Biderman** [29:40]
Yeah.

**Allen Park** [29:40]
Salt?

**Dan Biderman** [29:41]
Let's mix it up. Yeah.

**Allen Park** [29:42]
Okay.

**Dan Biderman** [29:42]
Did you put some salt?

**Allen Park** [29:43]
Here, I'll put some salt in.

**Dan Biderman** [29:44]
Um, we're gonna have yellow rice courtesy of Allen.

**Allen Park** [29:50]
Yes.

**Dan Biderman** [29:51]
Um, great.

**Allen Park** [29:52]
Well, we're nearly there.

**Dan Biderman** [29:54]
Yeah.

We'll let the these guys cook for a little bit.

### Memory Architecture

**Allen Park** [30:00]
With what you guys are attacking, how do you determine what, you know, should live inside of the weights? What should still kind of be handled with RAG? Um, where things should be orchestrated?

**Dan Biderman** [30:11]
Yeah. So this is a major open question, I would say, for us and for everyone else. And it's not just an open question in, in, uh, AI and, and a startup. It's an open question in the study of human memory from its, its, uh, inception.

It's like what kind of knowledge should be internalized and what kind of knowledge should be externalized. Does it make sense for you to remember everything you've seen as a person? Uh, people who have that, uh, often are not enjoying that capability.

**Allen Park** [30:42]
Yeah.

**Dan Biderman** [30:42]
And it's, it can be very distracting. It can be, uh, at some times very scary. Um, so a certain amount of forgetting is healthy. Certain things you wanna remember in the weights and certain things you wanna put in, in text.

So I would say it's an open research problem. The way to work on it is to train models both to train models to manage it themselves. And that's an active area for us. Have the model know, like, without any explicit supervision signal to determine this kind of stuff I can pull from my brain.

**Allen Park** [31:10]
Mm-hmm.

**Dan Biderman** [31:10]
And that kind of stuff I rather keep in notes. And you can imagine these kinds of things, uh, relate to like the saliency of certain parts in the data, how, how often, the frequency at which they repeat, and the affordances of like what can you do when you know this fact from your brain.

**Allen Park** [31:28]
Yeah.

**Dan Biderman** [31:28]
Uh, you can imagine that, uh, remembering, I don't know, remembering, uh, the, the, your room number in a hotel for tonight is less important than remembering your partner's phone number or remembering your address,right?

**Allen Park** [31:44]
Yeah.

**Dan Biderman** [31:44]
So.

**Allen Park** [31:45]
So it's higher signal data to.

**Dan Biderman** [31:46]
Higher signal data. And now the thing is, if you start manually, heuristically saying this is in, this is out, then it becomes a whack-a-mole.

**Allen Park** [31:54]
Okay.

**Dan Biderman** [31:54]
Every, every person in every enterprise has different data, and you can't really very easily pick and choose what goes in, what goes out. So the holy grail is have the model learn for itself, have it operate with a notebook where it can take notes, have it operate with a, with a brain, associative parameter efficient thing that it can read from, and have it decide when to go to each and do this with training in an unconstrained way.

And we're working on it, and more breakthroughs are needed.

**Allen Park** [32:20]
Yeah.

**Dan Biderman** [32:20]
But I think that, that is the dream, that the model learns what, what comes in and what comes out, and we get out of the way.

**Allen Park** [32:26]
Gotcha. Okay. And so that's the angle you wanna achieve where it's completely autonomous, so there's no human in the loop to kind of tell it what to fetch. Um.

**Dan Biderman** [32:35]
The human in the loop can be the user.

**Allen Park** [32:36]
Mm-hmm.

**Dan Biderman** [32:37]
And if the user chooses to say keep this in or keep this out, we would like the model to listen to the user.

**Allen Park** [32:42]
Yeah.

**Dan Biderman** [32:43]
But we would lo not like to depend on the user,right?

**Allen Park** [32:46]
Okay.

**Dan Biderman** [32:46]
So user feedback is something we can learn from, and we can learn from implicitly, but we don't want a person supervising every step because that's not the way, uh, people enjoy using, uh, language models. But I would say the kind of models that we're building, uh, unlike other models where you, you do thumbs up and thumbs down, and you're basically like helping the provider maybe in the next version give you something that's more workable.

**Allen Park** [33:10]
Yeah.

**Dan Biderman** [33:10]
Here, if you give a thumbs up or thumbs down or you say something, you know that someone's gonna scale compute on what you said, and someone's gonna go and practice, uh, to get better at what you said. And this is kind of the thing we wanna get to, like building trust with the user that, uh, they're listened to, and they're making a model that's better.

And it's not generally better for everyone. It's better for them.

**Allen Park** [33:31]
Gotcha. So the benefit is it being a tighter loop. Um, compared to the gen, like you said, a general model provider, the UI having a thumbs up, thumbs down, you don't know if that's actually gonna contribute to the feedback.

**Dan Biderman** [33:40]
Yeah. There's tighter loop, and we use a different machinery.

**Allen Park** [33:42]
Okay.

**Dan Biderman** [33:42]
If, if we allow ourselves to use the machinery of training, we know we can, we can hammer that in. There's no uncertainty about it. And, and another important thing to say is not everything that a user tells you is ground truth,right?

**Allen Park** [33:54]
Mm-hmm.

**Dan Biderman** [33:54]
Not all of us, including myself, are Einsteins, and we can say things to the model where we think we'reright and the model is wrong. And increasingly, the models will get better, and increasingly, they'll know more things than we do.

So the model in some way has to learn and understand and kind of like discern what, which feedback is valuable and which feedback should be ignored.

**Allen Park** [34:13]
Mm-hmm.

**Dan Biderman** [34:13]
But there too, I think the holy grail is to get out of the way and have the model learn it if you define theright objectives for training.

**Allen Park** [34:20]
Gotcha. Okay. So get out of the model's way. Have there been any moments that have really kind of shown you what is still needed, like, or what's still possible? 'Cause it seems very ambitious to be in an end state.

**Dan Biderman** [34:34]
Yeah.

**Allen Park** [34:34]
Where a model autonomously can handle this all. Um, and I think like multi-agent setups and evenright now too, just constrained to coding still has problems.

**Dan Biderman** [34:43]
Yeah.

**Allen Park** [34:44]
Um, and so I guess when you get to knowledge work and to specific enterprises, it seems a lot higher stakes.

**Dan Biderman** [34:48]
Yeah.

**Allen Park** [34:49]
Um, and you can't make as many mistakes. And so have there been moments or even research topics that you're kind of seeing great progress inright now that can get you there?

**Dan Biderman** [34:57]
Yeah. Yeah. I would say like we're we, we will share more of our results.

**Allen Park** [35:00]
Mm-hmm.

**Dan Biderman** [35:01]
Like in the coming weeks and months.

**Allen Park** [35:02]
Okay.

**Dan Biderman** [35:02]
And we're kind of keeping, keeping the logs of it.

**Allen Park** [35:05]
Still confidential.

**Dan Biderman** [35:05]
Yeah.

**Allen Park** [35:05]
Yeah.

**Dan Biderman** [35:05]
But I would say the, the, the theme is and the kind of thing you're seeing, and we're not the only ones seeing it, but I would say like we're looking very closely into it is, is behaviors around token efficiency, the ability of the models to go where they need to go and solve things faster.

Um, and, and basically it's, it's this kind of view of intelligence where getting smarter means, um, exerting less energy, uh, to, to solve increasingly harder problems.

**Allen Park** [35:35]
Yeah.

**Dan Biderman** [35:36]
And these are the kind of things we're going for. Um, and, and a lot of it, it comes into, into life in the form of like, uh, efficiency and speed, but we will share more of that soon.

**Allen Park** [35:48]
Okay.

**Dan Biderman** [35:48]
Yeah.

**Allen Park** [35:48]
I guess on the token efficiency point, do you also heavily consider like model routing? 'Cause I assume there may be some cheaper models that may do the same job for much less, but there may be some tasks that it may take, uh, much cheaper open source model, like $100 worth of tokens, when a much smarter model may be able to do it in one shot, like very quickly for much cheaper.

**Dan Biderman** [36:10]
Yeah. I would say that, um, routing is an interesting thing.

**Allen Park** [36:14]
Mm-hmm.

**Dan Biderman** [36:14]
Uh, and, and it there's a reason so many, uh, enterprises and computer scientists are looking into it because, uh, we have models that are overkill for many things.

**Allen Park** [36:26]
Yeah.

**Dan Biderman** [36:26]
You, you don't need Fable to tell you how much like, um.

**Allen Park** [36:30]
Water to put to.

**Dan Biderman** [36:31]
Soft on to put in.

**Allen Park** [36:32]
Yeah.

**Dan Biderman** [36:32]
Rice and stuff. Um, at the same time, you do, you do need to know when to go to Fable when you are trying to crack something that's, that's, uh.

**Allen Park** [36:39]
Mm-hmm.

**Dan Biderman** [36:39]
That's above your pay grade. Um, so I think routing is a great direction. Um, routing too, it's not that easy as someone who's worked on it.

**Allen Park** [36:49]
Yeah.

**Dan Biderman** [36:49]
Um, you can train, uh.

**Allen Park** [36:50]
Or what are the main challenges having the experience of seeing the difficulties of routing?

**Dan Biderman** [36:56]
So I think routing will be part of the solution there for sure.

**Allen Park** [36:58]
Yeah.

**Dan Biderman** [36:58]
And I think many people, not just myself, say the solution is multimodal. It's not Engram taking over.

**Allen Park** [37:05]
Yeah.

**Dan Biderman** [37:05]
It is one model, and you teach it things.

**Allen Park** [37:07]
Yeah.

**Dan Biderman** [37:07]
And, and, uh, you can close Stargate. Um, that's not our approach here.

**Allen Park** [37:11]
Yeah.

**Dan Biderman** [37:11]
Um, the solution will involve some form of routing. Um, and in that part of routing, our, our thing would be like, you know, your best friend or your employee, you cannot fire someone who's been looking over your shoulder the whole time and can tell you where to look, where to focus, and can even go out and ask Fable things in a more targeted way with more context that Fable can then go and work on it for days and solve very high-stakes tasks for you.

Um, so I think routing is great. Uh, it's just, uh, it's just, uh, like any other research area, uh, it's an unsolved thing. There's been progress on it, but a lot of more work is required to actually route things to the, theright model, theright time, uh, theright cost.

Um, and models are continually updated. Uh, versions are coming out. So yeah.

**Allen Park** [37:58]
Yeah. Okay. That makes sense. It seems like there's still a lot of open questions and even some of the work you're doing is confidential, which makes sense.

### Team & Hiring

**Dan Biderman** [38:04]
Yeah.

**Allen Park** [38:04]
And has to release more. I guess now shifting towards a team.

**Dan Biderman** [38:07]
Yeah.

**Allen Park** [38:08]
Um, yeah. How is it like working with such a mixed group of folks, like some who had professors, um, while you were doing studying.

**Dan Biderman** [38:15]
Yeah.

**Allen Park** [38:16]
And some that you also met and have left their PhDs or finished their PhDs as well.

**Dan Biderman** [38:20]
Yeah.

**Allen Park** [38:20]
Like Jack.

**Dan Biderman** [38:20]
Yeah.

**Allen Park** [38:21]
Um, and yeah, having such like a diverse group.

**Dan Biderman** [38:23]
So I would say, uh, there's many strengths to our group. Um, I'm not sure diversity is one of them. So if you're watching this.

**Allen Park** [38:30]
Yeah.

**Dan Biderman** [38:30]
We have a lot of researchers.

**Allen Park** [38:32]
Yeah.

**Dan Biderman** [38:32]
Uh, we could, we could diversify.

**Allen Park** [38:34]
Research focused. Yeah.

**Dan Biderman** [38:34]
We could diversify a little bit. Um, I would say, um, for this niche that we're trying to solve, which is memory and continual learning, taking knowledge, shoving it into weights, I think our team is the most specialized team in that, in that kind of thing.

Uh, and our team, our team is fun. You know, Jack.

**Allen Park** [38:54]
Yeah.

**Dan Biderman** [38:54]
Jack is a, is a fun guy.

**Allen Park** [38:56]
Jack is a very fun guy.

**Dan Biderman** [38:57]
Uh, and, and Jack is teaching us a lot of things about how to, how to think clearly, communicate ourselves clearly internally and to the external world, just see in the same way. Very complementary kind of like approaches. Uh, she thought a lot about how humans and AI work together, uh, in, in this closed loop.

**Allen Park** [39:18]
Mm-hmm.

**Dan Biderman** [39:18]
Of a machine and, and human. Um.

**Allen Park** [39:21]
Yeah.

**Dan Biderman** [39:22]
Sabrina and I were tiny bit more on the, on the systems, uh, side. And I was, Scott and I worked on statistics.

**Allen Park** [39:29]
Yeah.

**Dan Biderman** [39:29]
So it's all researchers and not diverse in this kind of way, but it is, as you said, diverse in our inclinations.

**Allen Park** [39:35]
Yeah.

**Dan Biderman** [39:35]
Some of us are more mathy, some of us are more systemsy, some of us are more like AI leaders. And so it's been interesting. The way we try to do this is to take those PhDs with a lot of experience and pair them up with like those kind of, uh, you know, up-and-coming, rising, uh, cracked types that.

**Allen Park** [39:55]
Mm-hmm.

**Dan Biderman** [39:55]
That have joined our company. Uh, for example, Shizza or Dhruv from Stanford and Berkeley respectively.

**Allen Park** [40:02]
Yeah.

**Dan Biderman** [40:02]
They're, they have research backgrounds and have written papers, um, but they're kind of, um, you know, entering this field with a lot of momentum. And we like to pair them up with someone who's been, who's been around the field for a few years, has some intuitions, can warn them from the rabbit holes.

And this, I think, powerful combination of different levels of expertise, uh, different levels of freshness of thought, uh, is making our place an interesting one. And I would say.

**Allen Park** [40:30]
Yeah.

**Dan Biderman** [40:30]
Culturally, all of us from day one, uh, in the, in the fall, winter of '25, which was a crazy time in terms of frontier lab.

**Allen Park** [40:40]
Mm-hmm. Yeah.

**Dan Biderman** [40:41]
Uh, recruiting and, and, and AGI anxiety. I would say all of us entered into the startup world in a very sober way, knowing that it's not just a research club in here.

**Allen Park** [40:53]
Mm-hmm.

**Dan Biderman** [40:53]
It's not just a journal club in here. The thing that's missing is products.

**Allen Park** [40:56]
Yeah.

**Dan Biderman** [40:56]
And products need to be distributed.

**Allen Park** [40:58]
Gotcha.

**Dan Biderman** [40:59]
You have to earn theright to play, um, by, by selling things that people love. So we're working very hard on that as well.

**Allen Park** [41:04]
That makes sense. I guess more of a fun question.

**Dan Biderman** [41:06]
Yeah.

**Allen Park** [41:07]
Out of founders, who do you think has the best taste in food? I assume you guys sometimes like order food or even like go out. Like, are there some that stand out to you? 'Cause I'm the.

**Dan Biderman** [41:15]
Yeah. I would say, uh, my co-co-founder Sabrina. He's, he's consistently, uh, he loves life.

**Allen Park** [41:24]
Mm-hmm.

**Dan Biderman** [41:24]
He loves, uh, good food. He's very well known in the company as someone who's the surf and turf guy.

**Allen Park** [41:30]
Gotcha.

**Dan Biderman** [41:31]
Uh, he will have it for lunch and dinner. And, um, yeah, being on DoorDash with him has been an inspiration for me. Um, and I would say, I would say, uh, Jack and Jesse have, have good taste as well.

**Allen Park** [41:43]
Good taste as well.

**Dan Biderman** [41:44]
Uh, yeah.

**Allen Park** [41:45]
Great.

**Dan Biderman** [41:45]
It's, it's fun to be with them. And we're now in like a small office, like an apartment. So we have all the lunches and then often dinners together.

**Allen Park** [41:52]
Yeah.

**Dan Biderman** [41:52]
And it's, it's, uh, a bit like a family, you know? We at home, you see your parents and siblings. Sometimes it's too much, but it's like, you know, you're, you're not never gonna forget.

**Allen Park** [42:02]
Yeah.

**Dan Biderman** [42:02]
That period.

**Allen Park** [42:03]
No, it does sound very fun. Have you cooked for them yet?

**Dan Biderman** [42:05]
Um, have I? Um, maybe not enough.

**Allen Park** [42:10]
Okay.

**Dan Biderman** [42:10]
Maybe I should.

**Allen Park** [42:11]
Maybe this week.

**Dan Biderman** [42:11]
Maybe this weekend I'll cook some.

**Allen Park** [42:13]
Okay.

**Dan Biderman** [42:13]
Um, they're cooking other things. They're cooking stuff on the science, but yeah.

**Allen Park** [42:17]
That's true.

**Dan Biderman** [42:17]
Um.

**Allen Park** [42:18]
They're cooking on the research.

**Dan Biderman** [42:19]
Yeah. On the research and there.

**Allen Park** [42:22]
Yeah. That's great. Okay. The rice should be done.

**Dan Biderman** [42:24]
Should be done.

**Allen Park** [42:25]
Do we wanna give it a try? Let's see. But we'll see. First bite.

**Dan Biderman** [42:31]
It's good. It's nice.

**Allen Park** [42:32]
Oh, yeah. It is. Yeah. The bottom's definitely more cooked. Um, you can also probably let that sit.

**Dan Biderman** [42:37]
Yeah. Now, these guys.

**Allen Park** [42:38]
Yeah. Would you say this is ready?

**Dan Biderman** [42:40]
These guys are, ooh, probably ready. Um, we can perhaps open it up a little bit.

**Allen Park** [42:47]
Yeah.

**Dan Biderman** [42:48]
Let it, uh, concentrate a bit. We can let those guys concentrate for a couple minutes here. And then, um, and then.

**Allen Park** [42:59]
Yeah.

**Dan Biderman** [42:59]
And then maybe we can cut those things with our handsright now.

**Allen Park** [43:02]
Yeah. Do you have any call to actions? Are you guys looking for specific folks, hiring for types of researchers?

**Dan Biderman** [43:08]
Yeah. Uh, I would say like, um, what we're trying to build.

**Allen Park** [43:13]
Mm-hmm.

**Dan Biderman** [43:13]
Which is those systems that continually learn. Obviously, there are many open problems on the research side, how you do this without destroying the model, how you do it.

**Allen Park** [43:22]
Yeah.

**Dan Biderman** [43:22]
In a cost-efficient way. Um, what data do you learn from, um, and stuff like that. But it's also an extremely ambitious, um, infrastructure problem.

**Allen Park** [43:33]
Mm-hmm.

**Dan Biderman** [43:33]
If you truly believe in the possibility that there's gonna be trillions of tokens of, of, uh.

**Allen Park** [43:39]
Data within the company.

**Dan Biderman** [43:40]
Enterprise data.

**Allen Park** [43:41]
Yeah.

**Dan Biderman** [43:41]
Or even personal data.

**Allen Park** [43:43]
Yeah.

**Dan Biderman** [43:43]
And if you truly believe that.

**Allen Park** [43:44]
Yeah.

**Dan Biderman** [43:44]
We can get to the level where we have those kinds of parameter-efficient adapters for.

**Allen Park** [43:49]
Yeah.

**Dan Biderman** [43:49]
Every person and team, you suddenly think about deployments that involve millions of different endpoints stored in different places.

**Allen Park** [43:57]
Mm-hmm.

**Dan Biderman** [43:57]
That need to be efficiently read from disk to HBM.

**Allen Park** [44:01]
Yeah.

**Dan Biderman** [44:01]
Uh, and then use that inference time.

**Allen Park** [44:04]
Mm-hmm.

**Dan Biderman** [44:04]
Um, and swapped and updated. It's gonna be if things work out for us, this thing ha will have a massive compute footprint and many new questions on, on systems and, and balancing of AI workloads in new ways. So.

**Allen Park** [44:18]
Mm-hmm.

**Dan Biderman** [44:19]
The kind of people that I think, um, could enjoy them and help us a lot are those, uh, one, uh, LLM kind of, um, performance engineers, research engineers who know how to make things go burr.

**Allen Park** [44:31]
Yeah.

**Dan Biderman** [44:31]
We have some of them.

**Allen Park** [44:32]
Make things go burr.

**Dan Biderman** [44:33]
Um, and we have, uh, Cade Daniel, who, uh, was one of the inference, uh, leads at Databricks.

**Allen Park** [44:41]
Yeah.

**Dan Biderman** [44:41]
And one of the core contributors of the LLM.

**Allen Park** [44:43]
Mm-hmm.

**Dan Biderman** [44:43]
And we are all kind of like, uh, systems inclined, but we think infrastructures engineers, uh, people who know how to work, uh, and, um, and build those large APIs and databases, I think would have a very fun time working on questions that they can't find in other places.

**Allen Park** [45:01]
Mm-hmm.

**Dan Biderman** [45:02]
Um, yeah. And, and generally, like, we always are, are open to, to smart and creative people who think out of the box and, and are committed to, to, to working on interesting problems.

**Allen Park** [45:14]
Yeah.

**Dan Biderman** [45:15]
Um, yeah.

**Allen Park** [45:18]
No, that's, that's great. It seems like a very creative mix of researchers and people who also care about the infrastructure, as you said.

**Dan Biderman** [45:25]
Yeah. Yeah.

### Efficiency

**Allen Park** [45:25]
Um.

**Dan Biderman** [45:25]
And I also wanted to say something, I guess.

**Allen Park** [45:28]
Yeah.

**Dan Biderman** [45:28]
I was just thinking about it while I was talking before that a lot of the use cases for this kind of training and continual learning.

**Allen Park** [45:34]
Yeah.

**Dan Biderman** [45:35]
Uh, involve.

**Allen Park** [45:36]
Mm-hmm.

**Dan Biderman** [45:36]
Um, let's give it a teeny bit more.

**Allen Park** [45:40]
Yeah.

**Dan Biderman** [45:41]
A bit less fluids there.

**Allen Park** [45:42]
Okay.

**Dan Biderman** [45:43]
Uh, the quizzing.

**Allen Park** [45:43]
Yeah.

**Dan Biderman** [45:43]
And so.

**Allen Park** [45:44]
Mm-hmm.

**Dan Biderman** [45:46]
The, the point for me is the principle is any kind of like efficiency and intelligence, they cannot really be de-decoupled.

**Allen Park** [45:53]
Okay.

**Dan Biderman** [45:54]
Sometimes people think if you're building something that's more efficient that can save you dollars, therefore you're not in the premium category. You're in the, uh, you know, you're, you're making the, the cheaper product. And that and intelligence, this is just, um, you know, purely wrong,right?

**Allen Park** [46:11]
Mm-hmm.

**Dan Biderman** [46:11]
So the more you can do with less, the more ambitious tasks you can solve longer term. Um.

**Allen Park** [46:18]
Gotcha.

**Dan Biderman** [46:18]
So I think that, um, the current paradigm of scaling with AI has been doing more with more.

**Allen Park** [46:26]
Mm-hmm.

**Dan Biderman** [46:27]
Uh, and it took us extremely far, and it will keep being a valuable way to, to build intelligence.

**Allen Park** [46:33]
Yeah.

**Dan Biderman** [46:33]
But I think the next paradigm and many leaders of the labs are seeing it, uh, involves a certain element of doing more with less.

**Allen Park** [46:39]
More with less.

**Dan Biderman** [46:40]
Um, to take on longer horizon tasks and harder problems generally. So I think going beyond enterprises and going beyond efficiency, that's where I hope to go.

**Allen Park** [46:51]
Yeah.

**Dan Biderman** [46:51]
Uh, if we solve and, and, and the challenges that we're facingright now.

**Allen Park** [46:55]
Mm-hmm. Okay. That makes sense. I think thinking of both coupled does seem very important, especially as you want to, like you said, do more ambitious tasks.

**Dan Biderman** [47:03]
Yeah.

**Allen Park** [47:04]
Okay. Should we mix in the spinach or just let it steam on top?

**Dan Biderman** [47:07]
Yeah. You can try mix it up.

**Allen Park** [47:08]
Okay.

**Dan Biderman** [47:08]
And just try it. It will steam.

**Allen Park** [47:11]
Let's soften a little bit. And I could cut this lemon.

**Dan Biderman** [47:21]
Yeah. Not too big.

**Allen Park** [47:29]
And then just squeeze the lemon in.

**Dan Biderman** [47:32]
Yeah.

**Allen Park** [47:32]
Okay.

**Dan Biderman** [47:32]
The, the sauce,right?

**Allen Park** [47:34]
Yeah.

**Dan Biderman** [47:36]
Okay. Do you wanna squeeze this one? This one has a little less, but this one's up.

**Allen Park** [47:42]
Amazing. And where can people find you?

**Dan Biderman** [47:45]
Where can people find us?

**Allen Park** [47:46]
Yeah. Are you guys.

**Dan Biderman** [47:47]
They, they can find us at engram.com.

**Allen Park** [47:49]
Engram.com.

**Dan Biderman** [47:50]
Um, they can find us NSF. Um.

**Allen Park** [47:53]
You guys are pretty fun.

**Dan Biderman** [47:53]
They can write to me, uh, at dan@engram.com, um, to talk about things. Um, and yeah, I would say like as we're scaling up different parts of the company that involve engineering, that involve product and business, there's a lot of things to talk about beyond the, the frontier AI type research.

Uh, and there's a lot for us to learn from smart people.

**Allen Park** [48:17]
Great. Awesome. That's exciting. Should we try it?

**Dan Biderman** [48:20]
Yeah.

I, I challenge every, every leader of a NeoLab to come to my house in Noy Valley and cook things with me. And, uh, and I'm sure we can learn a lot from each other.

**Allen Park** [48:39]
Cheers.

**Dan Biderman** [48:44]
Good?

**Allen Park** [48:45]
Mm-hmm.

**Dan Biderman** [48:46]
I think it turned out well now.

**Allen Park** [48:48]
Yeah. Very good. Wow. That's very good.

**Dan Biderman** [48:52]
Uh, maybe a little bit Persian.

**Allen Park** [48:54]
Mm-hmm.

**Dan Biderman** [48:54]
I would say. Um, with the yellow rice. So great.

**Allen Park** [48:58]
I'm a big fan. Okay.

Sean?

**Sean** [49:03]
Rice is amazing. Rice and spinach is amazing on this one. Actually, I'm just gonna finish the whole thing after.

**Allen Park** [49:08]
Yeah.

**Sean** [49:09]
Mm-hmm. Or meatballs are a bit softer than that.

**Allen Park** [49:11]
Yeah.

**Sean** [49:12]
Mm-hmm.

**Dan Biderman** [49:13]
It's way too heavy.

**Sean** [49:14]
But it's so tasty.

**Dan Biderman** [49:15]
The entire show, I pretended to be the cooking expert here.

**Allen Park** [49:17]
Okay.

**Dan Biderman** [49:18]
It has to be pure, pure crazy. Just kidding.

**Allen Park** [49:21]
This is like a 8 out of 10, 9 out of 10. Mm-hmm. Great. But yeah, no, I think that's basically it. It turned out pretty well. I mean, how was it? Was it fun? Did you enjoy it?

**Dan Biderman** [49:30]
It was the funnest, uh, podcast.

**Allen Park** [49:33]
Okay.

**Dan Biderman** [49:33]
I ever had.

**Allen Park** [49:34]
Yeah.

**Dan Biderman** [49:34]
Makes you feel at home.

**Allen Park** [49:35]
That's true.

**Dan Biderman** [49:36]
Um, easier to talk about things.

**Allen Park** [49:38]
Thank you again for coming.

**Dan Biderman** [49:39]
Yeah.

**Allen Park** [49:39]
And hopefully it was a fun time.

**Dan Biderman** [49:40]
It was super fun.

**Sean** [49:41]
My weights have been updated.

**Allen Park** [49:42]
That's good.

**Dan Biderman** [49:43]
That's good.

---

This library is powered by PodHood (https://podhood.com), the podcast website platform.
