LALatent SpaceJul 13, 2026· 49:44

The AI Memory Problem: Why Long Context Isn’t Enough — Dan Biderman, Engram Co-founder & CEO

Engram co-founder and CEO Dan Biderman tells hosts Allen Park and Sean that long context and RAG aren't enough for AI memory—continual learning and gradient-based weight updates are needed. He argues Engram's approach compresses company knowledge into 'cartridges' that let models reason with far fewer tokens, overcoming 'context rot' and the inefficiency of re-reading massive corpora. Biderman draws on his background in Israeli special forces and computational neuroscience to explain how training creates intuition beyond text retrieval, citing examples like Harvey's legal queries where holistic understanding beats search. He envisions personal AI weights that improve like Tamagotchis, driven by user-specific feedback loops, and stresses that token efficiency and intelligence are inseparable—doing more with less enables harder problems. Engram is hiring infrastructure engineers to deploy millions of continuously updated memories.

  1. 0:00Intro
  2. 1:45Military to AI
  3. 7:12Context Bet
  4. 14:10Data Scale
  5. 24:31Enterprise & Personal
  6. 30:00Memory Architecture
  7. 38:03Team & Hiring
  8. 45:25Efficiency

Powered by PodHood

Transcript

Intro0:00

Dan Biderman0:00

I'm too absorbed in the cooking, man.

Allen Park0:02

Yeah.

Dan Biderman0:02

Sorry, I'm fried, um, the entire show I pretended to be the cooking expert here.

Allen Park0:08

Okay.

Dan Biderman0:08

It has to be pure crazy.

Allen Park0:14

Hey guys, welcome to the Latent Space Cooking Show, where we invite founders and researchers and let them cook. Today we have a very special guest, co-founder and CEO Dan Biderman of Engram.

Dan Biderman0:25

Nice to be here.

Allen Park0:26

Yes, thank you for coming. Um, I heard, or we got connected through Jack, who is saying that you're a very good cook. And again, congrats on the $98M seed round. Do you cook often? How frequent is it? I assume you're constantly researching and building and.

Dan Biderman0:40

Yeah. So I would say since I started the company, there's a little bit less cooking of this type, uh, cooking other things. But cooking is something I've been doing since I, I was a kid. My parents cook, my grandparents used to cook, and it's, it's a way for me to kind of relax and digest things and think things through.

Allen Park0:58

Great. Yeah. Well, I'm very excited. Today's a little different, where we're actually going to have Dan lead us through one of the recipes. Do you want to reveal to us what we're making?

Dan Biderman1:07

Yeah. So I was thinking we can make, um, beef meatballs with some vegetables and have them swimming in some kind of, uh, white sauce with white wine and, and, uh, chicken stock. It's a pretty nice Mediterranean type, uh, meatballs that I've kind of stumbled upon recently when visiting my parents in Tel Aviv and cooking stuff with what they had at home.

So I thought it would be nice to, to redo this with you and see how it feels.

Allen Park1:32

Great. Well, I'm very excited. Do you want to kick us off?

Dan Biderman1:35

We're going to fry the onion. I'm going to chop it, and I'm going to chop some garlic, and then we'll, uh, make, make them into balls and, and fry them.

Allen Park1:42

Fry them. Great. Okay.

Dan Biderman1:43

Yeah.

Allen Park1:43

Well, let's get started. We'll both do it in parallel.

Military to AI1:45

Dan Biderman1:45

Awesome.

Allen Park1:46

Um, but yeah, I guess to start, I'm kind of curious about your background. So you wanted to become a professor, if I recall correctly, and you also had experience in the Israeli military or special forces.

Dan Biderman1:59

Yeah.

Allen Park1:59

Um, how did that turn you into becoming a founder of a research.

Dan Biderman2:03

Yeah.

Allen Park2:03

AI research company?

Dan Biderman2:05

Yeah. So I guess, uh, it's, it's easy to connect the dots, uh, looking backward, but in real time, it was not. I didn't, uh, grow up planning to be this founder and a chef working on this hot topic in AI.

Allen Park2:20

Mm-hmm.

Dan Biderman2:20

Things came to it. Um, so I grew up in Tel Aviv and, uh, and as you said, uh, was an officer and, and working on special operations, specifically naval special operations, which was very entrepreneurial. It was, uh, basically finding and identifying crazy things we could do and, and prove to people that we should scale resources and, and people to actually do them.

So it was a lot of, a lot of researching, a lot of pitching, and then a lot of hard work to make things real.

Allen Park2:51

Yeah.

Dan Biderman2:52

Um, and then after I finished this, um, I went to university, um, and studied cognitive, uh, neuroscience in Israel, uh, in an interesting program where, um, actually a lot of Israeli novelists and professors and scientists and.

Allen Park3:09

Mm-hmm.

Dan Biderman3:09

And even chefs. You have Tamagotchi; the Israeli chef went to this program many years before me. Um, and so I went there, got really curious about the brain, curious about human behavior.

Allen Park3:22

Yeah.

Dan Biderman3:22

Curious about statistics. And then I went to, to New York to do a PhD in computational neuroscience, um, which is a community that's been one of those stubborn communities that cared about neural networks.

Allen Park3:35

Mm-hmm.

Dan Biderman3:35

For many decades before it was a hot topic.

Allen Park3:38

Yeah.

Dan Biderman3:38

Uh, here in the valley. Um, so I was there and studied with some statisticians and physicists and biologists trying to understand, um, how neural networks can be used to understand the brain.

Allen Park3:50

Mm-hmm.

Dan Biderman3:51

To understand animal behavior. And, uh, and then decided that I wanted to go deeper into the LLM world after seeing ChatGPT. So I started working at Mosaic, uh, where I worked on LoRA and met a bunch of great people.

And then from there, I went to work with Chris Ray at Stanford.

Allen Park4:10

Chris Ray, yeah.

Dan Biderman4:10

And Scott Linderman, who both are co-founders of my current company.

Allen Park4:14

Yeah.

Dan Biderman4:14

And where I met my co-founder Sabri and my other co-founders Jack and Jesse, uh, who worked on very similar topics from their labs in Cornell and Berkeley.

Allen Park4:25

Yeah. No, it seems like a very, very research-focused and also serendipitous. Going back a little on your background.

Dan Biderman4:32

Yeah.

Allen Park4:32

Is there something about the Israeli special forces? Because I think it was the Wiz co-founders,right, who were also ex-special forces.

Dan Biderman4:39

Yeah. Uh, we worked pretty closely with Assaf, uh, Rapa Poort, the Wiz CEO.

Allen Park4:44

Mm-hmm.

Dan Biderman4:44

Who's, uh, uh, a mentor to me. Um, and, um, yeah, there's interesting things going on. Uh, there, I would say the, the most important thing is actually, uh, when you're 18, it's joining the intelligence or taking on responsibility.

Uh, the interesting bit about it is that you're not just in a student mode. You're not just heads down doing exams and projects. Rather, you interact with, with grownups, and you have resources, and you argue your, your for your, uh, budgets and things like that.

Allen Park5:18

Interesting. Oh, okay.

Dan Biderman5:19

Um, and it forces you to kind of build a little bit of social skills, a little bit of maturity.

Allen Park5:25

Mm-hmm.

Dan Biderman5:25

Um, which, which I benefited from.

Allen Park5:28

Yeah.

Dan Biderman5:28

Um, and, um, you know, there's a reason many countries in the world choose not to have everyone go to the army. It's a long time, and it's, and it's, uh, uh, a package deal. And I was, uh, in the Israeli army in different years, not the ones that are the recent ones.

Allen Park5:46

Gotcha.

Dan Biderman5:46

Um, so it's not like I recommend everyone goes to the army, but I would say that it forces you to kind of think about other s other, other aspects of your personality, not just the intellectual ones. And also generally, I would say Israel as a culture is a place where, um, you can basically you get multiple shots at goal.

Allen Park6:06

Mm-hmm.

Dan Biderman6:06

If you're not the best in high school, um, you still might have a, a good position in the military. And if you're really good, more doors open up for you for university. And even if you didn't do anything super interesting in, in military and got to university and found out you have good and deep inclinations in science and engineering, you can be the most successful person.

And we have some of those Israelis who are, um, leaders, uh, in various, uh, academic and technological places.

Allen Park6:34

Oh, that's great.

Dan Biderman6:34

Um, so I would say, yeah, it's mostly like you can basically.

Allen Park6:38

The culture.

Dan Biderman6:39

The cards are dealt multiple times.

Allen Park6:40

Yeah.

Dan Biderman6:41

More emphasis on social aspects, more emphasis on being mature, on being a team player. Uh, but again, it comes at a price that you get, uh, you, you're older when you get to things. Um, so I started college when I was 24, which is, uh, a, a middle PhD in, in an American, uh, university.

But yeah.

Allen Park7:05

That makes sense. I think it's great how it gives you a lot of shots on goal and is a very beautiful culture. I know fast forwarding to today, back to Engram.

Context Bet7:12

Dan Biderman7:13

Yeah.

Allen Park7:13

What kind of prompted you to focus on this issue with context? Were there specific breakthroughs? Was it, you know, the performance of open source models or even just the issue of Context Rot, long horizon agents, or was it just that you're very passionate about this area as a whole?

Dan Biderman7:30

Yeah. I'll answer that in a second. Can we.

Allen Park7:34

Yeah.

Dan Biderman7:34

Can we heat up some.

Allen Park7:35

Yes. Let's start getting going.

Dan Biderman7:37

Oil here to fry up the.

Allen Park7:39

Great. So we can turn on the Impulse stove.

Dan Biderman7:42

So yeah, I, my, my PhD was focusing on, on, on a field that's not super in vogue today, but I think will become super crucial again, which is semi-supervised learning.

Allen Park7:56

Hmm. Okay.

Dan Biderman7:57

Which is you have very little data. You don't have infinite trillion tokens of data. You have very infinite, you have a very small set of examples. And from them, you want to extrapolate and learn efficiently for, uh, and, and achieve things more, more generally.

Allen Park8:13

Mm-hmm.

Dan Biderman8:14

So how to be data efficient. So my whole academic work was basically all of my papers had the same plot. We're on the X-axis. I had.

Allen Park8:24

Mm-hmm.

Dan Biderman8:24

Some, some cost, some resource, and on the Y-axis, I had, um, I had, uh, accuracy, sort of Pareto curve. Uh, and that's been my, my upbringing.

Allen Park8:35

Mm-hmm.

Dan Biderman8:36

Efficiency, how to do more with less, um, salt.

Allen Park8:39

Yeah.

Dan Biderman8:40

Yeah. You can put some salt. Um, and I went to, to Chris Ray's lab, um, and.

Allen Park8:45

At Stanford,right?

Dan Biderman8:46

At Stanford. And there I worked on agents, on ideas called minions, uh, and, and asked kind of the question of how can we have agents interface with very large, uh, corpora of knowledge in a way that's economical, and when we introduce some patterns that involve sub-agents, uh, a bit before it was, uh, popular.

Allen Park9:08

Oh, sorry. We can also put this side of your glass.

Dan Biderman9:09

And from that perspective of efficiency and of cost, um, we started asking ourselves, actually.

Allen Park9:17

Yeah.

Dan Biderman9:17

Um, is there a more efficient way to interface with data that's not just agentic orchestration and things like this? Can we actually use the, the magic of training to, um, have models that are efficient, faster, and, and, and can operate with fewer tokens?

And some of these ideas kind of were developed in parallel by some of us, uh, Jesse, Sabri, and myself.

Allen Park9:40

Yeah.

Dan Biderman9:40

Uh, and the common theme was that if you have a very large corpus of knowledge, um, and you allow the model, you give the model time to study it in advance to basically ask itself questions, give itself quizzes, try to solve problems, and then train it using gradient descent in the same way you would pre-train a model.

Allen Park10:01

Mm-hmm.

Dan Biderman10:02

Um, then you can create very compact representations of your context that can later on, you can load them. We call those cartridges.

Allen Park10:10

Mm-hmm.

Dan Biderman10:10

These are like capsules of knowledge you can load in and out of the model that are like a brain state that describes the model's world and the corpus in a way that's going to be 1,000x more compressed. And when you use them, you can basically, uh, use far fewer tokens, be less confused, more accurate, and things like this.

So we got into.

Allen Park10:30

And these are task-specific cartridges?

Dan Biderman10:32

So they can be corpus-specific, like documents inside a company. They can be task-specific, like you want to learn a certain skill. But the idea is that in many cases, it makes sense to go beyond textual representations.

Allen Park10:45

Yeah.

Dan Biderman10:45

So when we're cooking now and we have all these, like, great textbooks.

Allen Park10:49

Mm-hmm.

Dan Biderman10:50

Our cookbooks here, Samina Swat and, you know, a book from French Laundry. These are amazing books that I.

Allen Park10:55

Yeah.

Dan Biderman10:55

I respect. But to say that because we have the books and we can read them, to say that we deeply know how to cook to the same level of Samina Swat or, or, or Thomas Keller.

Allen Park11:05

Yeah.

Dan Biderman11:06

It's incorrect,right?

Allen Park11:07

Yeah.

Dan Biderman11:07

We can read those things and we can repeat all their steps in a very robotic way. And we can come in, current LLMs are like coming into the kitchen first time every time, reading the textbook, uh, cooking the dish, measuring everything.

Uh, but they don't have the intuition of, uh, of a chef that's pinching salt and kneading, kneading dough and things like this. So the kind of thing we're after with this kind of, uh, training and creating those cartridges is this kind of intuition in the models that goes beyond notes and recipes to the kind of intelligence that allows then, allows you to come, come up with the next recipe, thing that hasn't been explored before to do the next, uh, move, the next extrapolation.

Allen Park11:49

Yeah. I guess double-clicking on building this intuition, could you kind of contextualize it more on how it would be different from extracting, let's say, using the cooking example? If you get all the useful notes and useful sections from a cookbook, that will help you actually understand all the complexities of making a dish.

Just providing it, I guess that would be similar to RAG, getting just those chunks.

Dan Biderman12:13

Yeah.

Allen Park12:13

And then understanding that, I guess, in context and then providing an output.

Dan Biderman12:18

Yeah. So this is, this is an excellent method. Um, we do not take the bet that I'm putting two eggs.

Allen Park12:24

Yeah.

Dan Biderman12:24

In here, if that's okay. Um.

Allen Park12:26

And is this enough grated carrot and zucchini?

Dan Biderman12:28

Yeah. Excellent. You can put it in here. Um.

Allen Park12:31

Yeah.

Dan Biderman12:33

So we're not taking the bet that no ne no notes need to be taken,right?

Allen Park12:39

All of it.

Dan Biderman12:39

Yeah. Put all of it inside.

Allen Park12:40

Mm-hmm.

Dan Biderman12:41

Um, all the greatest chefs in the world, they have notes and they have books and they.

Allen Park12:46

Yeah.

Dan Biderman12:46

They have diaries and they document what worked and what didn't work with their experiments.

Allen Park12:50

Mm-hmm.

Dan Biderman12:50

But they also have brains that they also have hands and, and, uh, and, uh, a tongue.

Allen Park12:58

Yeah.

Dan Biderman12:58

That can remind them how something tastes and what worked and what failed and what was easy and what was hard. So, uh, in all of our work, we never say that textual representations are useless.

Allen Park13:11

Mm-hmm.

Dan Biderman13:11

We basically use them all the time and we construct those wikis and knowledge bases and things like this.

Allen Park13:17

Yeah.

Dan Biderman13:18

Um, at the same time, what we say is that the layer above those, the learn the layer of intuition and learning, uh, that is in the form of numbers of parameters, um, is, is the, the full experience of the human chef.

The best chef in the world is all the notes combined with a nervous system that reads those notes.

Allen Park13:38

Gotcha.

Dan Biderman13:38

And can implement them and innovate them and maybe don't re-return to the same notes over and over again. There are some dishes where they don't need to, to reread them.

Allen Park13:45

Mm-hmm.

Dan Biderman13:46

And the current problem in, uh. So I was saying that, um, we want, we want the best of both worlds.

Allen Park13:54

Mm-hmm.

Dan Biderman13:54

Every knowledge worker, if they can't write notes and they cannot document the events of the day, they would be, uh, in a, a disadvantage.

Allen Park14:02

Mm-hmm.

Dan Biderman14:02

But, uh, if you wipe their brain every evening, they would also be at a severe disadvantage.

Allen Park14:06

Yeah.

Dan Biderman14:07

Um, so we want to have the best of both worlds. And, uh, the thing is that textual representations can take you, uh, a, a long way. Uh, but the thing we're thinking about is like, look, the, at the rate at which knowledge is being created.

Data Scale14:10

Allen Park14:22

Mm-hmm.

Dan Biderman14:23

With now agents working on behalf of knowledge workers, creating artifacts, code, documents, uh, presentations, I think that people don't fully comprehend the size of the.

Allen Park14:37

Yeah.

Dan Biderman14:37

Of the knowledge workspaces they will deal with in 18 months.

Allen Park14:42

Yeah.

Dan Biderman14:43

Um, in 18 months, many companies would have, uh, maybe trillions of tokens, which.

Allen Park14:50

Of internal company.

Dan Biderman14:52

Of internal company data, proprietary data. I'm talking about like maybe trillions. It sounds exaggerated, but I don't think it's an impossibility if they're really AI native.

Allen Park14:59

Yeah.

Dan Biderman15:00

And I think, and, and what are trillions of tokens? Uh, like when I was at Mosaic two, three years ago, like we call this internet scale data, pre-training data.

Allen Park15:10

Mm-hmm.

Dan Biderman15:10

So imagine every company has data that it's basically internet scale data.

Allen Park15:14

Um, yeah.

Dan Biderman15:17

And I might need, uh, salt.

Allen Park15:18

Yeah, we'll have tea.

Dan Biderman15:19

And pepper and the bread crumbs.

Allen Park15:21

Yeah.

Dan Biderman15:21

For this. And we don't have a sink here, so hands will be a tiny bit dirty.

Allen Park15:26

No worries.

Dan Biderman15:27

The audience here is forgiving. Um.

Allen Park15:31

Gotta get your hands dirty.

Dan Biderman15:32

Yeah.

Awesome. I guess the principle here I once heard from a chef is that a ratio of one to one of meat with everything else is usually a healthy way to make meatballs. So I think we're, we're there.

Allen Park15:48

Yeah, that's fair. We are on salt and pepper.

Dan Biderman15:50

Um, nice meatballs. They're a bit warm with those fried, um.

Allen Park15:55

Oh. We're back with the fix.

Dan Biderman15:58

We're back.

Allen Park15:59

We're back with the fixed meatballs.

Dan Biderman16:01

With the fixed meatballs. It's, it's team building activity here.

Allen Park16:05

Yes.

Dan Biderman16:06

Oh, yeah. Let's put those spices that you brought.

Allen Park16:08

Yeah.

Dan Biderman16:09

I trust you on, on these ones. Cumin, let's put a little bit, not a ton.

Allen Park16:14

Yeah. We won't dump it like last time.

Dan Biderman16:16

Cumin is my wife's favorite spice.

Allen Park16:18

Oh, yeah. Cumin's great.

Dan Biderman16:20

I have mixed feelings about it, but.

Allen Park16:21

Oh, okay.

Dan Biderman16:22

Yeah. It's good. Yeah.

Allen Park16:23

Good.

Dan Biderman16:24

Good. So we have the meatballs. Now is the time to make the balls.

Allen Park16:28

So we were talking about what were we talking about?

Dan Biderman16:34

I'm.

Allen Park16:34

I think I'm just.

Dan Biderman16:35

I'm too absorbed.

Allen Park16:36

Oh.

Dan Biderman16:36

I'm too absorbed in the cooking, man.

Allen Park16:38

Yeah.

Dan Biderman16:38

I'm sorry. I'm.

Allen Park16:39

No, it's good. You were.

Dan Biderman16:39

You were pitching things here.

Allen Park16:41

You're talking about how companies will have a corpus of data that will be at the scale of the internet?

Dan Biderman16:47

Yeah. Yeah. And maybe today.

Allen Park16:49

Maybe today.

Dan Biderman16:49

Some they don't have it. And maybe today you can go relatively far.

Allen Park16:54

Yeah.

Dan Biderman16:54

With textual representations. They're also interpretable. They're very good.

Allen Park16:58

Mm-hmm.

Dan Biderman16:59

Um, but I just think at a certain scale, even those textual representations will be hard to make.

Allen Park17:04

Okay.

Dan Biderman17:04

If you have trillion tokens, how do you create exactly a wiki or an index of those that you keep updated all the time? How big is this?

Allen Park17:11

Yeah.

Dan Biderman17:11

Um, knowledge base that you create, how expensive will it be to process it with frontier models that know nothing about your, your company?

Allen Park17:19

So is it just that using frontier models will be expensive because that's the start from scratch every time? Is that the main?

Dan Biderman17:25

So there's the element of expensive because you reread more things, you consume more to tokens. That's one.

Allen Park17:30

Yeah.

Dan Biderman17:31

But two is like for the agentic tasks of 18 months from now, inside those major repositories of knowledge and asking the models more and more things and under specified ways.

Allen Park17:42

Mm-hmm.

Dan Biderman17:42

I suspect that the accuracy, uh, of the models would, would go down. The, the phenomenon of context rot,right?

Allen Park17:49

Yeah.

Dan Biderman17:49

The model has to read more. It will be less accurate. And we know this, and it will remain the same thing at even at the 10 million, uh, context window scale.

Allen Park17:58

Okay. That makes sense. Okay. Let's go wash our hands real quick.

Dan Biderman18:00

Yeah.

Allen Park18:00

And we'll beright back.

Dan Biderman18:02

Great.

Allen Park18:02

And we're back. Hands are all clean. So now what's the next thing? Just frying the meatballs?

Dan Biderman18:06

Um, yeah. Let's fry those meatballs.

Allen Park18:07

Okay. Do you want to turn this one?

Dan Biderman18:09

Yeah.

Allen Park18:09

You just press it and then turn it. Very intuitive.

Dan Biderman18:12

And we'll put some oil.

Allen Park18:13

Okay. We could put these in and.

Dan Biderman18:16

Yeah.

Allen Park18:17

While we do that, I guess more so on the question of like long horizon agents.

Dan Biderman18:22

Yeah.

Allen Park18:23

Um, what's the main issue that Engram is trying to solve? Soright now, if I were to steel man the opposing view, um, maybe company, internal company data isn't actually big enough where I'd need to put it into the weights.

Like what's the issue with just having RAG or, um, having specific models or even cheaper models since open source models are also very performant to handle a lot of tasks?

Dan Biderman18:47

Yeah. So I would say like, um, a thorny question in the research community is, can you come up with an example where only in weights training would work, where in-context learning will fail?

Allen Park19:01

Mm-hmm.

Dan Biderman19:02

And it turns out it's very hard to devise such examples. And for every example you give, someone can ask, well, what happens if the, the next, uh, the next fable has a 10 million context window?

Allen Park19:16

Mm-hmm.

Dan Biderman19:16

Um, and the way I see it in kind of my scientific upbringing, I see all of these questions of, of continual learning and memory as questions of long context, uh, in disguise.

Allen Park19:28

Yeah.

Dan Biderman19:28

Um, if the models could see a whole, whole company data and in principle would have this infinite context window, what then is the limitation?

Allen Park19:37

Mm-hmm.

Dan Biderman19:38

So the limitation is twofold. One is that we know even at very small scales that the more context you feed to the model, the more confused it gets.

Allen Park19:47

Okay.

Dan Biderman19:47

It's called the context rot. So you can feed in a certain, uh, number of tokens into the model and not get an error, but it doesn't mean that the model can reason in, in a, in a holistic way about them.

That's one thing. And.

Allen Park19:59

And what's the problem with compaction? Is compaction not as useful, would you say?

Dan Biderman20:03

So I think compaction is, is improving by the day.

Allen Park20:07

Yeah.

Dan Biderman20:07

And compaction for those in the audience who don't know what it is, it's like models actually managing their own context, uh, evicting certain tokens and, and keeping others. Uh, compaction works. Uh, that too, when you go into a longer horizon, compaction by definition is lossy.

Allen Park20:24

Mm-hmm.

Dan Biderman20:24

You discard some and keep other things. And I think it's, it's a, it's a correct way to, to go. Uh.

Allen Park20:31

Yeah.

Dan Biderman20:31

But it's, it's also like very deterministic.

Allen Park20:34

Yeah.

Dan Biderman20:34

Either you're in or you're out. And the current versions of compaction, uh, show these kinds of, uh, uh, issues when, when very deep into the session you can get, um, confused and you can get, um, forgetful.

Allen Park20:49

Mm-hmm.

Dan Biderman20:50

So we think compaction will be part of the story.

Allen Park20:52

Okay.

Dan Biderman20:52

We think another part of the story is some sort of neural memory trace, which too is a lossy thing. It evicts some and, and keeps some, but not in the text, uh, representation and the weights representation.

Allen Park21:05

Okay.

Dan Biderman21:05

But I would say the main thing we're trying to solve or the main thing for which continual learning is needed, one is token efficiency and cost, which is a major issue that wasn't actually an issue when we started the company late last year and became moreurgent now.

And problem number two is if you can do, uh, if you can do the same thing with, with fewer resources, when you scale up to very large resources, suddenly you can take on tasks that were previously not possible.

Allen Park21:35

Yeah.

Dan Biderman21:36

Way more long horizon.

Allen Park21:37

Okay.

Dan Biderman21:37

Way more, uh, adaptive. And we're not quite there yet. Uh, we're focusingright now on, on the first component, which is getting these models to reason on large context with fewer tokens and do it in a way that's, that's more, uh, uh, le-less confused.

Allen Park21:57

Yeah.

Dan Biderman21:57

But we think that eventually, uh, part of the.

Allen Park22:01

Yeah. I'll take pictures. Thank you.

Dan Biderman22:03

Part of the, part of the solution for very hard tasks in, in science and engineering and defense and all that stuff will involve some form of gradient-based updates.

Allen Park22:15

Mm-hmm.

Dan Biderman22:16

Um, during, uh, during, uh, doing these long horizon tasks.

Allen Park22:20

Gotcha.

Dan Biderman22:20

Uh, some people call this, uh, test time compute or test time training. Um, these are just different names for the same thing. I think we have a, a proof of existence.

Allen Park22:31

Mm-hmm.

Dan Biderman22:31

From pre-training that, um, you can pack a lot of information in very few numbers.

Allen Park22:38

Yeah.

Dan Biderman22:38

Very efficiently. Um.

Allen Park22:39

Mm-hmm.

Dan Biderman22:40

So the examples we like to give is that, uh, if you take a Llama 70B model and you load one article from Wikipedia, which is a few tens of kilobytes, and you have the model read this, uh, the brain state of the model when reading this few tens of kilobytes is like 80 gigabytes.

Allen Park22:57

Mm-hmm.

Dan Biderman22:57

80 gigabytes on, on the HBM of the GPU.

Allen Park23:00

Yeah.

Dan Biderman23:01

It's an insane amount. And the entire set of parameters of this model would be like 140 or, or, or so gigabytes.

Allen Park23:07

Mm-hmm.

Dan Biderman23:07

FP16. So those 100 or so gigabytes with some distortion represent the entire internet. And this one article about Taylor Swift is like same order of magnitude memory consumption on the GPU. So it's highly memory inefficient.

Allen Park23:24

Yeah.

Dan Biderman23:24

That's a systems problem. That's the, the KV cache monstrosity that the smartest people in the world are trying to solve from the chip side of things and from the, from the software and kernel side of things. Um, but yeah.

So I would say another way to look at what we're doing from that angle, from the systems angle, is basically, uh, if we, if we can get a bit more technical.

Allen Park23:46

Yeah.

Dan Biderman23:46

Instead of doing those, uh, prefills where the model's just reading and reading and reading a corpus, uh, we are kind of like the we're destroying prefill. We're scaling training compute in some other time so we can load the thing into the model and can just immediately start decoding.

Allen Park24:03

Okay.

Dan Biderman24:03

Or prefill a little bit. And this kind of goes hand in hand with trends in how data centers are built out, this aggregating prefill and decode and doing this on different specialized, uh, cards.

Allen Park24:15

Mm-hmm.

Dan Biderman24:15

And, uh, and this is part of our.

Allen Park24:17

Yeah.

Dan Biderman24:17

Part of our initial interest in the thing.

Allen Park24:20

Mm-hmm.

Dan Biderman24:20

Okay. So it seems like we're mostly ready in here. I think it's a good time for our.

Allen Park24:27

Have the wine.

Dan Biderman24:27

White wine.

Allen Park24:28

Okay.

Dan Biderman24:28

On both of them.

Allen Park24:29

Great.

Dan Biderman24:29

Uh, maybe.

Allen Park24:30

Do you wanna do that?

Dan Biderman24:30

Yeah.

Allen Park24:31

Yeah.

Dan Biderman24:31

Yeah.

Enterprise & Personal24:31

Allen Park24:32

And do you have any tangible examples on this when you're talking about.

Dan Biderman24:35

Yeah.

Allen Park24:36

Your approach.

Dan Biderman24:36

So if you think about, um,

let's think, for example, about like an enterprise, uh, a firm.

Allen Park24:44

Mm-hmm.

Dan Biderman24:45

Or, um, investment banking, a law firm or investment banking, they can have many client matters, uh, many.

Allen Park24:53

Like a lot of knowledge work.

Dan Biderman24:55

A lot. Yeah. So many client matters. They have various clients, clients do, uh, financing, mergers and acquisitions and things like this.

Allen Park25:02

Yeah.

Dan Biderman25:02

And take loans and do deals.

Allen Park25:03

You wanna put this on or not yet?

Dan Biderman25:05

Um, no, not yet.

Allen Park25:05

Okay.

Dan Biderman25:06

Will you wanna let it, uh, let it.

Allen Park25:09

Let it reduce a little? Okay.

Dan Biderman25:10

Yeah. And so.

Allen Park25:11

Yeah.

Dan Biderman25:12

For example, these are the kinds of things we, we work with Harvey.

Allen Park25:14

Mm-hmm.

Dan Biderman25:15

Very, very large file systems. Uh, and there's many queries that agents might run into. Either humans ask them or the agents have to solve them.

Allen Park25:23

Mm-hmm.

Dan Biderman25:23

Which are these kinds of like ambient hard questions that are not easily searchable with RAG. For example, if you wanna ask like, which M&A deals haven't we completed this year? So to actually solve this problem, you have to go client matter by client matter.

Allen Park25:40

Yeah.

Dan Biderman25:40

Read all the files. You can't read in any place that it was not completed. You have to take the gist.

Allen Park25:45

You have to understand.

Dan Biderman25:46

The, the thing hasn't the loop hasn't been closed.

Allen Park25:48

Yeah.

Dan Biderman25:48

Thing hasn't been completed. And now you can solve these tasks with, with frontier models and compaction.

Allen Park25:53

Mm-hmm.

Dan Biderman25:53

And when you ask them to do so, they will consume thousands of dollars for queries that we think are harmless.

Allen Park25:58

Mm-hmm.

Dan Biderman25:58

That every employee in the company would be able to answer.

Allen Park26:01

Mm-hmm.

Dan Biderman26:01

So this is just like one example, but these kinds of holistic things that are, you know, you can get the gist if you read everything, but you can't find a thing in one thing. The kind of queries where the whole is greater than the sum of its parts.

So that's where, uh, this kind of magic of training, uh, comes in. And that's the, the magic that, that, you know, Ilya and others.

Allen Park26:22

Yeah.

Dan Biderman26:23

Have shown us with pre-training,right?

Allen Park26:24

Yeah.

Dan Biderman26:24

You learn the entire web in these sets of weights, and suddenly the model can infer things, can generalize, can interpolate and extrapolate to new things. And that's the kind of knowledge we wanna do.

Allen Park26:35

Mm-hmm.

Dan Biderman26:35

And it, and it's not a, not a coincidence that we're training on the entire web and we're not just doing RAG over the entire web.

Allen Park26:43

Yeah.

Dan Biderman26:44

Or putting it in the catalog and renewing it in time. Because we do think that learning from a lot of knowledge somehow creates these associations in the model.

Allen Park26:51

Gotcha.

Dan Biderman26:51

So some of it is concrete problems of now. Other parts of it are bets that in 18 months from now, the scale of the data will require the methods that we know from pre-training work.

Allen Park27:03

Gotcha. And so are you doing parameter efficient fine-tuning for.

Dan Biderman27:07

Yeah.

Allen Park27:07

Specific companies like LoRA based off of the corpus?

Dan Biderman27:10

Yeah. Yeah. So our, our ambition in the long term.

Allen Park27:12

But also start to write?

Dan Biderman27:13

Yeah.

Allen Park27:14

Yeah.

Dan Biderman27:14

So our ambition in the long term, uh, is that every person has a, has a model or, or a part of the model or a set of weights.

Allen Park27:24

Mm-hmm.

Dan Biderman27:25

Uh, that.

Allen Park27:27

Yes, sir.

Dan Biderman27:28

That represents their knowledge, their expertise.

Allen Park27:31

Yeah.

Dan Biderman27:31

Learns from them. That the more time they spend with the model, the better it gets for them. The more data they give the model, the better the model is for them.

Allen Park27:39

Yeah.

Dan Biderman27:39

They control those sets of weights. It's theirs, um.

Allen Park27:43

And the rights?

Dan Biderman27:44

Uh, and they're incentivized to let, let, let's let it boil over, you know?

Allen Park27:48

Okay. Yeah.

Dan Biderman27:48

Um, so they're incentivized to we can, we can put the spices up.

Allen Park27:53

Mm-hmm.

Dan Biderman27:53

So they're incentivized to, to make it better in the same way that you would work with, with a Tamagotchi. The more you nurture it, the happier it is. And that's the kind of thing we would like to create with the model.

So that's the ultimate continual learning. And we think on the extreme scale, on the single human scale.

Allen Park28:08

Mm-hmm.

Dan Biderman28:08

Um, turns out this is not just a research problem. That's a major infra problem. And in the long, long term, I do think these things will actually run on people's devices.

Allen Park28:17

Yeah.

Dan Biderman28:17

And, uh, we're seeingright now the, the new hardware on personal computers is already, uh, you know, soon approaching the ability to run inference on close to trillion parameter models.

Allen Park28:29

Mm-hmm.

Dan Biderman28:30

Uh, which will be very interesting for that kind of personalization. But in the shorter term, we do think that, uh, the kind of large corpora of knowledge, very dense, with a lot of expertise, uh, can be found in the enterprise.

Allen Park28:43

Mm-hmm.

Dan Biderman28:43

Um, that's where people spend most of their time. That's where AI is, is actually being used the most.

Allen Park28:49

Yeah. And enterprises.

Dan Biderman28:50

Um, so we, we, we go there and that's our bet. And there we're, we're our bet is that parameter efficient fine-tuning methods like LoRA, like cartridges, like memory layers, different things we, we contributed to as, as a research team, um, can actually represent knowledge and can be combined with other methods that are like context management, traceable methods that, that people can, can use and understand and, and audit.

Allen Park29:15

Yeah.

Dan Biderman29:16

So it's a combination.

Allen Park29:17

Gotcha.

Dan Biderman29:18

So again, it's like, it's like the chef that has the cookbook and it has the recipes, has the diary, but also has the nervous system that learns from every session. Uh.

Allen Park29:28

Gotcha.

Dan Biderman29:28

So it's boiling here. It's all.

Allen Park29:30

Yeah.

Dan Biderman29:30

Pretty efficient here.

Allen Park29:32

Yeah. Very powerful stove.

Dan Biderman29:34

Yeah.

Allen Park29:36

Very powerful. And then medium, low, then you wanna put this on?

Dan Biderman29:40

Yeah.

Allen Park29:40

Salt?

Dan Biderman29:41

Let's mix it up. Yeah.

Allen Park29:42

Okay.

Dan Biderman29:42

Did you put some salt?

Allen Park29:43

Here, I'll put some salt in.

Dan Biderman29:44

Um, we're gonna have yellow rice courtesy of Allen.

Allen Park29:50

Yes.

Dan Biderman29:51

Um, great.

Allen Park29:52

Well, we're nearly there.

Dan Biderman29:54

Yeah.

We'll let the these guys cook for a little bit.

Memory Architecture30:00

Allen Park30:00

With what you guys are attacking, how do you determine what, you know, should live inside of the weights? What should still kind of be handled with RAG? Um, where things should be orchestrated?

Dan Biderman30:11

Yeah. So this is a major open question, I would say, for us and for everyone else. And it's not just an open question in, in, uh, AI and, and a startup. It's an open question in the study of human memory from its, its, uh, inception.

It's like what kind of knowledge should be internalized and what kind of knowledge should be externalized. Does it make sense for you to remember everything you've seen as a person? Uh, people who have that, uh, often are not enjoying that capability.

Allen Park30:42

Yeah.

Dan Biderman30:42

And it's, it can be very distracting. It can be, uh, at some times very scary. Um, so a certain amount of forgetting is healthy. Certain things you wanna remember in the weights and certain things you wanna put in, in text.

So I would say it's an open research problem. The way to work on it is to train models both to train models to manage it themselves. And that's an active area for us. Have the model know, like, without any explicit supervision signal to determine this kind of stuff I can pull from my brain.

Allen Park31:10

Mm-hmm.

Dan Biderman31:10

And that kind of stuff I rather keep in notes. And you can imagine these kinds of things, uh, relate to like the saliency of certain parts in the data, how, how often, the frequency at which they repeat, and the affordances of like what can you do when you know this fact from your brain.

Allen Park31:28

Yeah.

Dan Biderman31:28

Uh, you can imagine that, uh, remembering, I don't know, remembering, uh, the, the, your room number in a hotel for tonight is less important than remembering your partner's phone number or remembering your address,right?

Allen Park31:44

Yeah.

Dan Biderman31:44

So.

Allen Park31:45

So it's higher signal data to.

Dan Biderman31:46

Higher signal data. And now the thing is, if you start manually, heuristically saying this is in, this is out, then it becomes a whack-a-mole.

Allen Park31:54

Okay.

Dan Biderman31:54

Every, every person in every enterprise has different data, and you can't really very easily pick and choose what goes in, what goes out. So the holy grail is have the model learn for itself, have it operate with a notebook where it can take notes, have it operate with a, with a brain, associative parameter efficient thing that it can read from, and have it decide when to go to each and do this with training in an unconstrained way.

And we're working on it, and more breakthroughs are needed.

Allen Park32:20

Yeah.

Dan Biderman32:20

But I think that, that is the dream, that the model learns what, what comes in and what comes out, and we get out of the way.

Allen Park32:26

Gotcha. Okay. And so that's the angle you wanna achieve where it's completely autonomous, so there's no human in the loop to kind of tell it what to fetch. Um.

Dan Biderman32:35

The human in the loop can be the user.

Allen Park32:36

Mm-hmm.

Dan Biderman32:37

And if the user chooses to say keep this in or keep this out, we would like the model to listen to the user.

Allen Park32:42

Yeah.

Dan Biderman32:43

But we would lo not like to depend on the user,right?

Allen Park32:46

Okay.

Dan Biderman32:46

So user feedback is something we can learn from, and we can learn from implicitly, but we don't want a person supervising every step because that's not the way, uh, people enjoy using, uh, language models. But I would say the kind of models that we're building, uh, unlike other models where you, you do thumbs up and thumbs down, and you're basically like helping the provider maybe in the next version give you something that's more workable.

Allen Park33:10

Yeah.

Dan Biderman33:10

Here, if you give a thumbs up or thumbs down or you say something, you know that someone's gonna scale compute on what you said, and someone's gonna go and practice, uh, to get better at what you said. And this is kind of the thing we wanna get to, like building trust with the user that, uh, they're listened to, and they're making a model that's better.

And it's not generally better for everyone. It's better for them.

Allen Park33:31

Gotcha. So the benefit is it being a tighter loop. Um, compared to the gen, like you said, a general model provider, the UI having a thumbs up, thumbs down, you don't know if that's actually gonna contribute to the feedback.

Dan Biderman33:40

Yeah. There's tighter loop, and we use a different machinery.

Allen Park33:42

Okay.

Dan Biderman33:42

If, if we allow ourselves to use the machinery of training, we know we can, we can hammer that in. There's no uncertainty about it. And, and another important thing to say is not everything that a user tells you is ground truth,right?

Allen Park33:54

Mm-hmm.

Dan Biderman33:54

Not all of us, including myself, are Einsteins, and we can say things to the model where we think we'reright and the model is wrong. And increasingly, the models will get better, and increasingly, they'll know more things than we do.

So the model in some way has to learn and understand and kind of like discern what, which feedback is valuable and which feedback should be ignored.

Allen Park34:13

Mm-hmm.

Dan Biderman34:13

But there too, I think the holy grail is to get out of the way and have the model learn it if you define theright objectives for training.

Allen Park34:20

Gotcha. Okay. So get out of the model's way. Have there been any moments that have really kind of shown you what is still needed, like, or what's still possible? 'Cause it seems very ambitious to be in an end state.

Dan Biderman34:34

Yeah.

Allen Park34:34

Where a model autonomously can handle this all. Um, and I think like multi-agent setups and evenright now too, just constrained to coding still has problems.

Dan Biderman34:43

Yeah.

Allen Park34:44

Um, and so I guess when you get to knowledge work and to specific enterprises, it seems a lot higher stakes.

Dan Biderman34:48

Yeah.

Allen Park34:49

Um, and you can't make as many mistakes. And so have there been moments or even research topics that you're kind of seeing great progress inright now that can get you there?

Dan Biderman34:57

Yeah. Yeah. I would say like we're we, we will share more of our results.

Allen Park35:00

Mm-hmm.

Dan Biderman35:01

Like in the coming weeks and months.

Allen Park35:02

Okay.

Dan Biderman35:02

And we're kind of keeping, keeping the logs of it.

Allen Park35:05

Still confidential.

Dan Biderman35:05

Yeah.

Allen Park35:05

Yeah.

Dan Biderman35:05

But I would say the, the, the theme is and the kind of thing you're seeing, and we're not the only ones seeing it, but I would say like we're looking very closely into it is, is behaviors around token efficiency, the ability of the models to go where they need to go and solve things faster.

Um, and, and basically it's, it's this kind of view of intelligence where getting smarter means, um, exerting less energy, uh, to, to solve increasingly harder problems.

Allen Park35:35

Yeah.

Dan Biderman35:36

And these are the kind of things we're going for. Um, and, and a lot of it, it comes into, into life in the form of like, uh, efficiency and speed, but we will share more of that soon.

Allen Park35:48

Okay.

Dan Biderman35:48

Yeah.

Allen Park35:48

I guess on the token efficiency point, do you also heavily consider like model routing? 'Cause I assume there may be some cheaper models that may do the same job for much less, but there may be some tasks that it may take, uh, much cheaper open source model, like $100 worth of tokens, when a much smarter model may be able to do it in one shot, like very quickly for much cheaper.

Dan Biderman36:10

Yeah. I would say that, um, routing is an interesting thing.

Allen Park36:14

Mm-hmm.

Dan Biderman36:14

Uh, and, and it there's a reason so many, uh, enterprises and computer scientists are looking into it because, uh, we have models that are overkill for many things.

Allen Park36:26

Yeah.

Dan Biderman36:26

You, you don't need Fable to tell you how much like, um.

Allen Park36:30

Water to put to.

Dan Biderman36:31

Soft on to put in.

Allen Park36:32

Yeah.

Dan Biderman36:32

Rice and stuff. Um, at the same time, you do, you do need to know when to go to Fable when you are trying to crack something that's, that's, uh.

Allen Park36:39

Mm-hmm.

Dan Biderman36:39

That's above your pay grade. Um, so I think routing is a great direction. Um, routing too, it's not that easy as someone who's worked on it.

Allen Park36:49

Yeah.

Dan Biderman36:49

Um, you can train, uh.

Allen Park36:50

Or what are the main challenges having the experience of seeing the difficulties of routing?

Dan Biderman36:56

So I think routing will be part of the solution there for sure.

Allen Park36:58

Yeah.

Dan Biderman36:58

And I think many people, not just myself, say the solution is multimodal. It's not Engram taking over.

Allen Park37:05

Yeah.

Dan Biderman37:05

It is one model, and you teach it things.

Allen Park37:07

Yeah.

Dan Biderman37:07

And, and, uh, you can close Stargate. Um, that's not our approach here.

Allen Park37:11

Yeah.

Dan Biderman37:11

Um, the solution will involve some form of routing. Um, and in that part of routing, our, our thing would be like, you know, your best friend or your employee, you cannot fire someone who's been looking over your shoulder the whole time and can tell you where to look, where to focus, and can even go out and ask Fable things in a more targeted way with more context that Fable can then go and work on it for days and solve very high-stakes tasks for you.

Um, so I think routing is great. Uh, it's just, uh, it's just, uh, like any other research area, uh, it's an unsolved thing. There's been progress on it, but a lot of more work is required to actually route things to the, theright model, theright time, uh, theright cost.

Um, and models are continually updated. Uh, versions are coming out. So yeah.

Allen Park37:58

Yeah. Okay. That makes sense. It seems like there's still a lot of open questions and even some of the work you're doing is confidential, which makes sense.

Team & Hiring38:03

Dan Biderman38:04

Yeah.

Allen Park38:04

And has to release more. I guess now shifting towards a team.

Dan Biderman38:07

Yeah.

Allen Park38:08

Um, yeah. How is it like working with such a mixed group of folks, like some who had professors, um, while you were doing studying.

Dan Biderman38:15

Yeah.

Allen Park38:16

And some that you also met and have left their PhDs or finished their PhDs as well.

Dan Biderman38:20

Yeah.

Allen Park38:20

Like Jack.

Dan Biderman38:20

Yeah.

Allen Park38:21

Um, and yeah, having such like a diverse group.

Dan Biderman38:23

So I would say, uh, there's many strengths to our group. Um, I'm not sure diversity is one of them. So if you're watching this.

Allen Park38:30

Yeah.

Dan Biderman38:30

We have a lot of researchers.

Allen Park38:32

Yeah.

Dan Biderman38:32

Uh, we could, we could diversify.

Allen Park38:34

Research focused. Yeah.

Dan Biderman38:34

We could diversify a little bit. Um, I would say, um, for this niche that we're trying to solve, which is memory and continual learning, taking knowledge, shoving it into weights, I think our team is the most specialized team in that, in that kind of thing.

Uh, and our team, our team is fun. You know, Jack.

Allen Park38:54

Yeah.

Dan Biderman38:54

Jack is a, is a fun guy.

Allen Park38:56

Jack is a very fun guy.

Dan Biderman38:57

Uh, and, and Jack is teaching us a lot of things about how to, how to think clearly, communicate ourselves clearly internally and to the external world, just see in the same way. Very complementary kind of like approaches. Uh, she thought a lot about how humans and AI work together, uh, in, in this closed loop.

Allen Park39:18

Mm-hmm.

Dan Biderman39:18

Of a machine and, and human. Um.

Allen Park39:21

Yeah.

Dan Biderman39:22

Sabrina and I were tiny bit more on the, on the systems, uh, side. And I was, Scott and I worked on statistics.

Allen Park39:29

Yeah.

Dan Biderman39:29

So it's all researchers and not diverse in this kind of way, but it is, as you said, diverse in our inclinations.

Allen Park39:35

Yeah.

Dan Biderman39:35

Some of us are more mathy, some of us are more systemsy, some of us are more like AI leaders. And so it's been interesting. The way we try to do this is to take those PhDs with a lot of experience and pair them up with like those kind of, uh, you know, up-and-coming, rising, uh, cracked types that.

Allen Park39:55

Mm-hmm.

Dan Biderman39:55

That have joined our company. Uh, for example, Shizza or Dhruv from Stanford and Berkeley respectively.

Allen Park40:02

Yeah.

Dan Biderman40:02

They're, they have research backgrounds and have written papers, um, but they're kind of, um, you know, entering this field with a lot of momentum. And we like to pair them up with someone who's been, who's been around the field for a few years, has some intuitions, can warn them from the rabbit holes.

And this, I think, powerful combination of different levels of expertise, uh, different levels of freshness of thought, uh, is making our place an interesting one. And I would say.

Allen Park40:30

Yeah.

Dan Biderman40:30

Culturally, all of us from day one, uh, in the, in the fall, winter of '25, which was a crazy time in terms of frontier lab.

Allen Park40:40

Mm-hmm. Yeah.

Dan Biderman40:41

Uh, recruiting and, and, and AGI anxiety. I would say all of us entered into the startup world in a very sober way, knowing that it's not just a research club in here.

Allen Park40:53

Mm-hmm.

Dan Biderman40:53

It's not just a journal club in here. The thing that's missing is products.

Allen Park40:56

Yeah.

Dan Biderman40:56

And products need to be distributed.

Allen Park40:58

Gotcha.

Dan Biderman40:59

You have to earn theright to play, um, by, by selling things that people love. So we're working very hard on that as well.

Allen Park41:04

That makes sense. I guess more of a fun question.

Dan Biderman41:06

Yeah.

Allen Park41:07

Out of founders, who do you think has the best taste in food? I assume you guys sometimes like order food or even like go out. Like, are there some that stand out to you? 'Cause I'm the.

Dan Biderman41:15

Yeah. I would say, uh, my co-co-founder Sabrina. He's, he's consistently, uh, he loves life.

Allen Park41:24

Mm-hmm.

Dan Biderman41:24

He loves, uh, good food. He's very well known in the company as someone who's the surf and turf guy.

Allen Park41:30

Gotcha.

Dan Biderman41:31

Uh, he will have it for lunch and dinner. And, um, yeah, being on DoorDash with him has been an inspiration for me. Um, and I would say, I would say, uh, Jack and Jesse have, have good taste as well.

Allen Park41:43

Good taste as well.

Dan Biderman41:44

Uh, yeah.

Allen Park41:45

Great.

Dan Biderman41:45

It's, it's fun to be with them. And we're now in like a small office, like an apartment. So we have all the lunches and then often dinners together.

Allen Park41:52

Yeah.

Dan Biderman41:52

And it's, it's, uh, a bit like a family, you know? We at home, you see your parents and siblings. Sometimes it's too much, but it's like, you know, you're, you're not never gonna forget.

Allen Park42:02

Yeah.

Dan Biderman42:02

That period.

Allen Park42:03

No, it does sound very fun. Have you cooked for them yet?

Dan Biderman42:05

Um, have I? Um, maybe not enough.

Allen Park42:10

Okay.

Dan Biderman42:10

Maybe I should.

Allen Park42:11

Maybe this week.

Dan Biderman42:11

Maybe this weekend I'll cook some.

Allen Park42:13

Okay.

Dan Biderman42:13

Um, they're cooking other things. They're cooking stuff on the science, but yeah.

Allen Park42:17

That's true.

Dan Biderman42:17

Um.

Allen Park42:18

They're cooking on the research.

Dan Biderman42:19

Yeah. On the research and there.

Allen Park42:22

Yeah. That's great. Okay. The rice should be done.

Dan Biderman42:24

Should be done.

Allen Park42:25

Do we wanna give it a try? Let's see. But we'll see. First bite.

Dan Biderman42:31

It's good. It's nice.

Allen Park42:32

Oh, yeah. It is. Yeah. The bottom's definitely more cooked. Um, you can also probably let that sit.

Dan Biderman42:37

Yeah. Now, these guys.

Allen Park42:38

Yeah. Would you say this is ready?

Dan Biderman42:40

These guys are, ooh, probably ready. Um, we can perhaps open it up a little bit.

Allen Park42:47

Yeah.

Dan Biderman42:48

Let it, uh, concentrate a bit. We can let those guys concentrate for a couple minutes here. And then, um, and then.

Allen Park42:59

Yeah.

Dan Biderman42:59

And then maybe we can cut those things with our handsright now.

Allen Park43:02

Yeah. Do you have any call to actions? Are you guys looking for specific folks, hiring for types of researchers?

Dan Biderman43:08

Yeah. Uh, I would say like, um, what we're trying to build.

Allen Park43:13

Mm-hmm.

Dan Biderman43:13

Which is those systems that continually learn. Obviously, there are many open problems on the research side, how you do this without destroying the model, how you do it.

Allen Park43:22

Yeah.

Dan Biderman43:22

In a cost-efficient way. Um, what data do you learn from, um, and stuff like that. But it's also an extremely ambitious, um, infrastructure problem.

Allen Park43:33

Mm-hmm.

Dan Biderman43:33

If you truly believe in the possibility that there's gonna be trillions of tokens of, of, uh.

Allen Park43:39

Data within the company.

Dan Biderman43:40

Enterprise data.

Allen Park43:41

Yeah.

Dan Biderman43:41

Or even personal data.

Allen Park43:43

Yeah.

Dan Biderman43:43

And if you truly believe that.

Allen Park43:44

Yeah.

Dan Biderman43:44

We can get to the level where we have those kinds of parameter-efficient adapters for.

Allen Park43:49

Yeah.

Dan Biderman43:49

Every person and team, you suddenly think about deployments that involve millions of different endpoints stored in different places.

Allen Park43:57

Mm-hmm.

Dan Biderman43:57

That need to be efficiently read from disk to HBM.

Allen Park44:01

Yeah.

Dan Biderman44:01

Uh, and then use that inference time.

Allen Park44:04

Mm-hmm.

Dan Biderman44:04

Um, and swapped and updated. It's gonna be if things work out for us, this thing ha will have a massive compute footprint and many new questions on, on systems and, and balancing of AI workloads in new ways. So.

Allen Park44:18

Mm-hmm.

Dan Biderman44:19

The kind of people that I think, um, could enjoy them and help us a lot are those, uh, one, uh, LLM kind of, um, performance engineers, research engineers who know how to make things go burr.

Allen Park44:31

Yeah.

Dan Biderman44:31

We have some of them.

Allen Park44:32

Make things go burr.

Dan Biderman44:33

Um, and we have, uh, Cade Daniel, who, uh, was one of the inference, uh, leads at Databricks.

Allen Park44:41

Yeah.

Dan Biderman44:41

And one of the core contributors of the LLM.

Allen Park44:43

Mm-hmm.

Dan Biderman44:43

And we are all kind of like, uh, systems inclined, but we think infrastructures engineers, uh, people who know how to work, uh, and, um, and build those large APIs and databases, I think would have a very fun time working on questions that they can't find in other places.

Allen Park45:01

Mm-hmm.

Dan Biderman45:02

Um, yeah. And, and generally, like, we always are, are open to, to smart and creative people who think out of the box and, and are committed to, to, to working on interesting problems.

Allen Park45:14

Yeah.

Dan Biderman45:15

Um, yeah.

Allen Park45:18

No, that's, that's great. It seems like a very creative mix of researchers and people who also care about the infrastructure, as you said.

Dan Biderman45:25

Yeah. Yeah.

Efficiency45:25

Allen Park45:25

Um.

Dan Biderman45:25

And I also wanted to say something, I guess.

Allen Park45:28

Yeah.

Dan Biderman45:28

I was just thinking about it while I was talking before that a lot of the use cases for this kind of training and continual learning.

Allen Park45:34

Yeah.

Dan Biderman45:35

Uh, involve.

Allen Park45:36

Mm-hmm.

Dan Biderman45:36

Um, let's give it a teeny bit more.

Allen Park45:40

Yeah.

Dan Biderman45:41

A bit less fluids there.

Allen Park45:42

Okay.

Dan Biderman45:43

Uh, the quizzing.

Allen Park45:43

Yeah.

Dan Biderman45:43

And so.

Allen Park45:44

Mm-hmm.

Dan Biderman45:46

The, the point for me is the principle is any kind of like efficiency and intelligence, they cannot really be de-decoupled.

Allen Park45:53

Okay.

Dan Biderman45:54

Sometimes people think if you're building something that's more efficient that can save you dollars, therefore you're not in the premium category. You're in the, uh, you know, you're, you're making the, the cheaper product. And that and intelligence, this is just, um, you know, purely wrong,right?

Allen Park46:11

Mm-hmm.

Dan Biderman46:11

So the more you can do with less, the more ambitious tasks you can solve longer term. Um.

Allen Park46:18

Gotcha.

Dan Biderman46:18

So I think that, um, the current paradigm of scaling with AI has been doing more with more.

Allen Park46:26

Mm-hmm.

Dan Biderman46:27

Uh, and it took us extremely far, and it will keep being a valuable way to, to build intelligence.

Allen Park46:33

Yeah.

Dan Biderman46:33

But I think the next paradigm and many leaders of the labs are seeing it, uh, involves a certain element of doing more with less.

Allen Park46:39

More with less.

Dan Biderman46:40

Um, to take on longer horizon tasks and harder problems generally. So I think going beyond enterprises and going beyond efficiency, that's where I hope to go.

Allen Park46:51

Yeah.

Dan Biderman46:51

Uh, if we solve and, and, and the challenges that we're facingright now.

Allen Park46:55

Mm-hmm. Okay. That makes sense. I think thinking of both coupled does seem very important, especially as you want to, like you said, do more ambitious tasks.

Dan Biderman47:03

Yeah.

Allen Park47:04

Okay. Should we mix in the spinach or just let it steam on top?

Dan Biderman47:07

Yeah. You can try mix it up.

Allen Park47:08

Okay.

Dan Biderman47:08

And just try it. It will steam.

Allen Park47:11

Let's soften a little bit. And I could cut this lemon.

Dan Biderman47:21

Yeah. Not too big.

Allen Park47:29

And then just squeeze the lemon in.

Dan Biderman47:32

Yeah.

Allen Park47:32

Okay.

Dan Biderman47:32

The, the sauce,right?

Allen Park47:34

Yeah.

Dan Biderman47:36

Okay. Do you wanna squeeze this one? This one has a little less, but this one's up.

Allen Park47:42

Amazing. And where can people find you?

Dan Biderman47:45

Where can people find us?

Allen Park47:46

Yeah. Are you guys.

Dan Biderman47:47

They, they can find us at engram.com.

Allen Park47:49

Engram.com.

Dan Biderman47:50

Um, they can find us NSF. Um.

Allen Park47:53

You guys are pretty fun.

Dan Biderman47:53

They can write to me, uh, at dan@engram.com, um, to talk about things. Um, and yeah, I would say like as we're scaling up different parts of the company that involve engineering, that involve product and business, there's a lot of things to talk about beyond the, the frontier AI type research.

Uh, and there's a lot for us to learn from smart people.

Allen Park48:17

Great. Awesome. That's exciting. Should we try it?

Dan Biderman48:20

Yeah.

I, I challenge every, every leader of a NeoLab to come to my house in Noy Valley and cook things with me. And, uh, and I'm sure we can learn a lot from each other.

Allen Park48:39

Cheers.

Dan Biderman48:44

Good?

Allen Park48:45

Mm-hmm.

Dan Biderman48:46

I think it turned out well now.

Allen Park48:48

Yeah. Very good. Wow. That's very good.

Dan Biderman48:52

Uh, maybe a little bit Persian.

Allen Park48:54

Mm-hmm.

Dan Biderman48:54

I would say. Um, with the yellow rice. So great.

Allen Park48:58

I'm a big fan. Okay.

Sean?

Sean49:03

Rice is amazing. Rice and spinach is amazing on this one. Actually, I'm just gonna finish the whole thing after.

Allen Park49:08

Yeah.

Sean49:09

Mm-hmm. Or meatballs are a bit softer than that.

Allen Park49:11

Yeah.

Sean49:12

Mm-hmm.

Dan Biderman49:13

It's way too heavy.

Sean49:14

But it's so tasty.

Dan Biderman49:15

The entire show, I pretended to be the cooking expert here.

Allen Park49:17

Okay.

Dan Biderman49:18

It has to be pure, pure crazy. Just kidding.

Allen Park49:21

This is like a 8 out of 10, 9 out of 10. Mm-hmm. Great. But yeah, no, I think that's basically it. It turned out pretty well. I mean, how was it? Was it fun? Did you enjoy it?

Dan Biderman49:30

It was the funnest, uh, podcast.

Allen Park49:33

Okay.

Dan Biderman49:33

I ever had.

Allen Park49:34

Yeah.

Dan Biderman49:34

Makes you feel at home.

Allen Park49:35

That's true.

Dan Biderman49:36

Um, easier to talk about things.

Allen Park49:38

Thank you again for coming.

Dan Biderman49:39

Yeah.

Allen Park49:39

And hopefully it was a fun time.

Dan Biderman49:40

It was super fun.

Sean49:41

My weights have been updated.

Allen Park49:42

That's good.

Dan Biderman49:43

That's good.