LALatent SpaceApr 21, 2025· 34:02

Sleep-Time Compute — Letta AI (Charles Packer, Charlie Snell, Kevin Lin)

This episode covers Letta AI's new paper on Sleep-Time Compute, a scaling direction that applies compute during model idle periods (sleep time) rather than only at test time. Charles Packer, Kevin Lin, and Charlie Snell explain that sleep-time compute precomputes inferences from static context before queries arrive, yielding Pareto improvements on math benchmarks like GSM8K by shifting the accuracy-to-token curve leftwards. They distinguish it from test-time compute by emphasizing that sleep-time tokens incur no user-latency cost, and show that the benefit is largest when questions are predictable from context. The team also ties the concept to stateful agents and memory systems (building on MemGPT), releasing two implementations: one for low-latency chatbots and another for document-driven agents.

  1. 0:00Intro
  2. 1:13Sleep-Time Compute
  3. 3:02Stateful Agents
  4. 4:50Paper Motivation
  5. 6:54Sleep Analogy
  6. 11:00Toy Example
  7. 14:06How Literal?
  8. 17:30Evaluation
  9. 22:47Latency Matters
  10. 24:01Coding Results
  11. 27:26Memory Systems
  12. 30:54Closing

Powered by PodHood

Transcript

Intro0:00

Alessio0:00

Hey everyone, welcome back to another Latent Space Lightning Pod. This is Alessio, partner and CTO at Decibel, and I'm joined by my co-host Swyx, founder of Small AI.

Swyx0:09

Hello, and today we are in the remote studio. We're convening with the authors of Letta's new paper. I actually don't know the name of the, the new paper yet. But anyway, Charles, actually you, you are presented a workshop, but you're, you're most famously known as one of the co-authors of Mem-MemGPT, and two of your co-authors, Charlie and Kevin.

Welcome.

Charles Packer0:29

Yeah. Thank you.

Swyx0:31

Nice.

Charles Packer0:31

Yeah, so the new paper-- Yeah, I can just kick things off. Um, the new paper is called Sleep-Time Compute. I did see actually recently, you know, a clip of you on Twitter that was going around, you know, talking about sleep time-

Swyx0:40

I got so much hate for that. Oh my God.

Charles Packer0:45

Yeah, but I mean, I think that the name is, like, pretty self-explanatory. It's basically exploring this new scaling direction. Um, and it- there's been a ton of talk about, like, test-time compute as kind of the way to get over maybe some of the limitations and scaling of like the base model.

But the thing about test-time compute is it's only scaling compute at test time, right? And practically speaking, you know, machines, they're not like humans. They can be run all the time, and there's a ton of downtime, both in advance of, like, questions being asked, also, like, after questions have been asked too.

Sleep-Time Compute1:13

Charles Packer1:13

So I think beyond just scaling at test time, I think kind of the natural, like, very big missed opportunity is scaling at sleep time. So yeah, we're kind of dividing like inference time now or like post-training time into test time and sleep time.

I think that's like kind of natural breakdown. And the, the paper is all about, you know, what can you do when you do try to scale sleep time, kind of maybe what are the limitations and also, like, the advantages.

Swyx1:35

Yeah, sure. O- uh, one thing I also like, like to do is establish the different voices for people listening on audio. If, if each of you, Charlie and Kevin, you-- Charles, you wanna introduce a little bit in, in your relationship with the project?

Charles Packer1:47

Yeah. So yeah, I'm Charles. Uh, I guess, yeah, Sean gave a quick intro. So I was a PhD student back at Berkeley with Kevin and, like, at Berkeley, myself, Kevin and Sarah, um, who's also at Letta, we... The last year of our PhD, we worked on MemGPT together.

And yeah, now Kevin and I, um, we're at Letta and, you know, we're continuing a lot of both, like the research directions and like application directions to do with like LLMs, LMOs, building these compound systems around large language models.

Yeah, excited to kind of talk about our latest work more on like the research direction, but has, like, very big implications too for the systems we're building.

Kevin Lin2:18

Yeah. So hi, everyone. I'm Kevin. So I also graduated from the PhD last year. I shared one advisor with Charles, and towards the end of the PhD, we were working on MemGPT together, which I feel like was kind of, kind of what the initial version of what, like, a stateful agent looks like.

And I shared another advisor with Charlie, who has been working with us on this new Sleep-Time Compute work, kind of, uh, taking some of his ideas also from Test-Time Compute that he's worked on before.

Charlie Snell2:48

Yeah. So I'm a PhD student at Berkeley right now, and, um, I had done some work on test-time scaling previously and, um, so I was really excited to work on this, uh, direction of Sleep-Time Compute with the folks at Letta.

Swyx3:02

Charlie, your name looks familiar, but I think because one of your papers last year got a bit of attention on test-time compute as well.

Stateful Agents3:02

Charlie Snell3:08

Yeah, yeah. I did some work on test-time compute. I, I was a student researcher at Google for a while and, um, was working on test-time compute there and sort of did, um, some analysis of like what is the best way to scale test-time compute, and we looked at a few different ways of doing that.

It was, uh, that was one of my recent papers, yeah.

Swyx3:30

Okay. And now you're coining Sleep-Time Compute, so that's exciting. Um, yeah. Do-- any, any other context that's relevant? Uh, I noticed that you're no longer the LLMOS projects. You're just stateful agents. Is, is that a strategic shift?

Charles Packer3:45

Actually, I think with the-- we've been playing around with kind of more of like the, the copy, you know, like what kind of makes the most sense. Um, but I think, yeah, stateful agents, LLMOS, these are all like very similar ideas.

I think, you know, to, to have a stateful agent, you need like an LLMOS because you need something other than the LLM to kind of maintain state. So yeah, I think we've, we've been kind of like playing around more for like a general developer audience, like what really hammers home, like the ideas the hardest.

Is it like stateful agent? Is it LLMOS? But yeah, I would say like we're, we're definitely-- we use like both words to describe what we're building. I think LLMOS is kind of like the larger like... Talks about-- It's more about the infrastructure.

Like what do you-- what is the big system you're building? What, what is the code? What gets deployed? And then stateful agent is like the thing that you can build, the, the actual end product. You're building a stateful agent and you're like enabled by having this LLMOS style infrastructure.

Swyx4:32

Uh, cool. I, I guess we can get dive right into the paper then. Yeah. By the way, I also noticed that, and I'm rebranding your, um, AI engineer workshop to more stateful agents because I noticed that was your title.

Um, yeah. Uh, so, uh, what are the main findings or what, what was the, the background to starting work on Sleep-Time Compute?

Kevin Lin4:50

Yeah. So we've covered kind of like two, I think, really big ideas. One is test-time compute, so spending more time at test generally is going to get you better, uh, responses, be able to solve more difficult problems. And I think MGPT established kind of like what the future will look like in terms of agents will have some state, um, you know, your chatbot will have some memory, and we're seeing like that idea kind of being put, put in a lot of, uh, products and applications these days.

Paper Motivation4:50

Kevin Lin5:21

Um, so we're kind of thinking about test-time compute. Most of the evaluations and work in this area, they kind of assume you get all of your context at test time. You get a math problem, you get like the entire setup and the question at test time.

If you have a code base and you have an agent that re- tries to resolve like a PR in the code base, all of that is given to you at test time. Inherently, a lot of these applications are, are stateful.

Um, while you're not asking, uh, your code editor a, a question or like to do, uh, fill in the next line, it has all the context with, you know, your entire repository, the previous questions you've asked. What we're, uh, looking at in this paper is kind of like during the sleep time, so like when your agent, when the model has access to a lot of state, how to best represent that state.

For a code base, that could look like maybe like exploring the code base, mapping out what's important in the code base, looking for bugs and, and kind of the standard benchmarks that a lot of people are testing for test-time compute, kind of these mathematical reasoning tasks.

It could look something like you have some context that's like a setup of like I have, you know, like 100 tennis balls, half are blue, and then of those, uh, a quarter are marked with an X. Like what percentage are blue and marked with, with an X?

And while the agent has this context before it gets the final question, it, it can start making some of these precomputations and useful inferences before the user asks their question.

Sleep Analogy6:54

Alessio6:54

Can you guys maybe just explain the background behind the name? Uh, so I get like maybe the parallel to the brain of when we sleep, kind of like our memories settle into the brain and whatnot. Is that the initial inspiration?

Because you mentioned context as like a big part of what drives the sleep rather than like the runtime, so to speak.

Kevin Lin7:14

Yeah. I think there's a definitely like a cognitive analogy to, you know, when humans sleep, like you have memory that's being consolidated. I think also what's guiding our research a lot is kind of more the systems analogy of like when your computer, when you're not typing something, you know, like there, there's background like sleep time processes, right?

Like if you have a lot of data, there's, uh, you know, like indices, uh, that are built into databases, for example, that's like better repres- representing the data so that when it's queried, it's much more efficient. And similar in the test-time compute setting, you know, here the, the state is tokens and the kind of like sleep time, like indexing process is a re-representation of those tokens into something that is like more easily queryable and more flexible.

And this I think it like it very much depends on the application. So for like a chat application, you know, the re-representation will look something more like looking at all the memories, um, and the previous chats and kind of like building a hierarchy of things that are, uh, more important and like eliminating like inconsistencies in the past.

And for something like these test-time compute benchmarks, re-representation will look something more like computing subquantities or like rewriting and like doing some analysis ahead of time.

Alessio8:39

Yeah. And I, I wonder, I don't know if in the paper you have some sort of like chart or diagram. Uh, I, I think like I still... It's hard to understand when an LLM is sleeping. Like I understand that, but like if I were to use like, you know, look at the paper, it's like, is it...

Uh, uh, how do you define sleep time? Like when does sleep time start? When does it end? Is it on demand? Is it, uh, innate part of the LLM inference?

Kevin Lin9:03

Yeah. I think like we basically think that sleep time is like any time post-training that's not test time. I, I think that also kind of intuitively makes a lot of sense because like the LLM, if it's, you know, running in a data center somewhere, if you have your GPU live, like you could be using it, like inferencing on it at any given point in time, even when the user is not actually or like an event, an active event isn't happening.

So it's basically like, yeah, it's like every time that is not test time, you know, where you could presumably be doing some sort of compute, that's what we're calling sleep time.

I, I think, you know, to go back to the analogy thing, I think, you know, with like MemGPT, we also thought a lot about like, you know, is the systems analogy better or is like cognitive analogy better? Um, I think the cognitive analogy for MemGPT is like, oh, well, a human has like different levels of like representations of their brain and like, you know, maybe the core memories is going to be something that's like more high level and then a- archival is like something more abstract.

Um, I think similar to MemGPT, um, I think the cognitive analogies are almost like a subset, kind of like Kevin was saying, of the systems level analogies. And I think the system level analogies are just a lot sharper because, you know, at the end of the day, like with tokens, you-- it's like a, a memory hierarchy, and you have to have algorithms or systems that move tokens around.

But yeah, I think here, I think it's always fun to think about kind of the, the cognitive analogy. You know, like when humans-- when a human goes to bed, their brain does something, maybe defragments their memories. Who knows what's happening?

I guess, like we don't really know yet.

Alessio10:32

Uh-huh.

Kevin Lin10:33

And yeah, you could do that with, with like sleep time compute. Like if you activate sleep time compute on a chatbot like ChatGPT, it can like learn about you as you're not on chatgpt.com. I think that's, you know, kind of what they're probably gonna try to do.

That's the direction they're going in. But yeah, like Kevin said, I think like the systems analogy similar to MemGPT is just a lot-- It's, I think it's stronger. It's, uh, it's much more precise. It ties much more cleanly into like the real methods that are used, but I think both apply.

Yeah.

Toy Example11:00

Alessio11:00

Yeah. Now, now that you bring this up on screen, I think it's much clearer. I, I think for people that maybe are not watching the video, um, on the left side, you basically have the raw context. In this case is, um, a juggler can juggle eight hundred balls.

It's pretty impressive. A quarter of the balls are tennis balls, half, half of the tennis balls are indigo, and a tenth of the indigo ones are marked. So the idea here is that if you're then gonna ask how many marked indigo tennis balls are there, and the traditional kind of like test-time compute, you basically have chain of thought at test time and you try and come up with the answer.

And the sleep time is before you ask the question, you're kind of precomputing the answers in a way. I, I think one... O-obviously, memory is like one way to think about it, but to me, it's almost like a recommendation on questions.

You know, you're taking, you're taking the opposite approach. It's like instead of recommending the question, you're like recommending to yourself the answers and like precomputing them and like hoping that the user asks this to kind of save time as test time.

Is that a good way?

Swyx12:04

to think about it or is it kind of minimizing maybe the, um, the approach?

Charles Packer12:09

Yeah, I think for the purposes of like this paper, um, I think, you know, Kevin and Charlie can speak to this much more, but a lot of the experiments are driven by, you know, we want to be empirical and we want to be like scientific.

What are people studying in test-time compute? Like, what are the classical problems like, you know, the papers that Charlie has written, you know, about test-time compute, like what are the experimental domains? And it's obvious we can always like have, uh, fun discussions about, you know, imagine if your GPU was dreaming when it was like you're not around it.

But at the same time, I think if this is really kind of like a direction we can scale, I think we can kind of look at the problems in test-time compute and say things like pretty concretely about like how much we really can kind of take advantage of Sleep-Time Compute, which is effectively completely unmined today.

Like nobody's really scaling in the Sleep-Time Compute direction. And I think, you know, once test-time compute, you kind of reach the, the end of the line for how much we can really squeeze out of there. Sleep-Time Compute's like completely unmined, right?

So I think this, this paper is maybe it's thinking more about like the scaling principles. Like if you are interested in the principles behind test-time compute scaling, what can you do with Sleep-Time Compute? But yeah, maybe, uh, Charlie can rephrase that a little bit better or like closer to the, to the meat of the paper.

Charlie Snell13:25

Yeah, I think that's about right. Basically, we're interested in sort of the questions of like, you know, using sort of these math benchmarks. If we do Sleep-Time Compute, does that improve, you know, how we can scale compute at test time?

Is there like a Pareto improvement? Um, and then we also look at like if we apply additional Sleep-Time Compute in the background, can we shift that Pareto improvement even further? And so, yeah, we're, we're kind of in the paper trying to understand kind of these, uh, empirical questions about how, uh, Sleep-Time Compute can improve scaling in sort of these classical domains where test-time scaling has been studied.

Swyx14:06

Yeah. Got it. So quick question or pushback, I guess, on sleep. How literally do we take this, right? Humans sleep for one third of the day. Machines probably don't need that much time. Uh, I actually did even look back to the original MemGPT implementation, and I don't think there was...

How Literal?14:06

Swyx14:25

There was heartbeats, but there wasn't like a very heavy sleep cycle, right? So basically, like are we just rebranding sort of pre-computing some data or like data pre-processing as sleep? And is it, does it, does it get deeper than that?

Charles Packer14:41

Yeah. Well, I think, yeah, maybe you should talk about like the lower context component. Because I think with like it... So there's definitely the aspect of there's diminishing returns. So depending on what you're trying to do, you know, you will kind of re-reach a limit of how much you can re-represent the context.

I think in this case, you know, with these like GSM8K style questions or like AIME style, um, there's gonna be a limit. Like, you know, at c- a certain point, you will have kind of expanded the context to, to the point where you've covered like every single foreseeable question you could ask on, on that math problem.

But I think that's kind of what's cool about it, because I think that also means that it's inherently agentic. So like you kind of want to be intelligent about how much you actually apply Sleep-Time Compute. This isn't just something you're gonna like brute force and, you know, infinitely like get better results on like every single domain.

Kevin Lin15:32

Yeah. I think one thing to add, earlier you mentioned kind of like you kind of are hoping that you can guess some of the questions ahead of time, and I think that's like kind of one kind of precondition to a lot of this of like you have to know something about the query distribution.

And like similar in databases like kind of like with data cubes and stuff, a lot of it hinges on like you knowing like reasonably well, like what kind of query distribution you'll get. So we also did some analysis of like kind of trying to understand like where, like what kind of context it makes sense to like think more about in the background.

And I think that's like kind of the core part of what makes it agentic. So like one way to, uh, think about it is kind of like, oh, you just have like some background process that's like constantly running and churning the context and like creating more learned context.

And I think like the core of like what makes something agentic is like deciding like how much to do, how much to like think about it. And yeah, we have some analysis in the paper that's kind of... And Charlie can maybe jump in on this part, which is like kind of looking at like for questions which ones are more predictable from the context.

So the... We and we find that like when you have questions that are more predictable from the context, then it's like more helpful to think about them ahead of time. So I think what this will look like is more like when you have some state, then you can maybe have some very lightweight process that kind of like estimates like for this context like should I think about it a lot?

And if yes, then like, like really like push the LLMs, uh, to think about that. And if not, then like don't really do, do anything. And I think like, like kind of like the choosing to do something or not is like kind of like what is like the key of like what makes a system agentic or not.

Swyx17:30

Yeah. Um, so like a-actually let's, let's zoom out a little bit more in terms of like how you want to eval this, right? You're, you're testing on GSM here, but could, could you just talk a bit about like how you explore this as a paper, right?

Evaluation17:30

Swyx17:46

Because it's such a broad concept that with many implementations. So I just wanted to see like what empirical work did we do.

Kevin Lin17:56

Yeah, yeah, yeah. So I'll quickly talk about kind of the main core result here. So the main core result, so we take some of these standard benchmarks that people are interested in for test-time compute. One is, uh, GSM8K So typically you have just all this context which you see on the left here.

So like the query is just like a juggler and so on. So we-- to evaluate this, like Sleep-Time Compute, we separate it into kind of a state, so the context that's kind of setting up the problem, and then a query, which is the final part of the problem.

And then in the standard case, basically you just apply Test-Time Compute on the original query, where you have like all the information given at once. And in the Sleep-Time Compute setting, you're given the context first, and then the m- agent is prompted to like think about it, rewrite it as much as it wants, and then it receives the final query.

Swyx18:59

Yeah.

Kevin Lin18:59

Yeah. So this is kind of like, kind of the, the flavor of the main result here. So like with... So here like in gray, you kind of see the Test-Time Compute accuracy trade-off of the, of the standard test time setting.

So like you're given all the questions, and you see like with, you know, these like, you know, like jugglers with so many like, like clauses, like when you have on the left not much Test-Time Compute, then the accuracy is pretty bad, uh, you know, without doing like chain of thought on these questions.

But on the right side, you know, like the, the gray curve goes up. So you see like there's fundamentally kind of this trade-off between compute and accuracy. And then the blue is when the model is allowed to think about it ahead of time and think about it during sleep time.

And you see like especially when, you know, you're like being pushed to answer right away, you're telling the model like, "Go, go, go," like, "Tell me the answer," then there's like a huge gain from being able to think about these contexts ahead of time.

Swyx20:09

Yeah. Interesting. The right-hand side of the chart, you know, g- starts, it starts to converge a lot, which means that basically this is really about efficiency, right? If you're trying to get the most done with the least amount of tokens, then you want to do more sleep time.

Charles Packer20:28

I think there's also a trade-off where... Well, I think, yeah, I think it's important to clarify, like for the purposes of like the paper, um, the main way to really evaluate Sleep-Time Compute against Test-Time Compute is going to be like on this Pareto frontier trade-off.

Um, and I think the idea is that you want to kind of show, and this is what the like results are showing, that you kind of are like exclusively better. Like if you have access to Sleep-Time Compute, you are able to apply it, you should probably apply it.

Although if you also have access to kind of like log scale infinite Test-Time Compute, why would-- then probably you don't need to. You can just like always, you know, take as long as you want at test time. I think that the part about the right side of this plot is that it's very unfavorable.

I think, you know, it's a very... Just imagine even in like ChatGPT, like how many people start like a deep research and then leave, and like never return to that deep research? How many people, you know, turn off of like o1 and like go to like o3-mini or like, I guess like whatever the, you know, the past or low latency one is, um, just because they can't stand waiting that long for the a-answer to come back.

So there's a part in the paper where we kind of try to be a little bit, um, specific about this in terms of like, you know, assigning higher costs to tokens that come at test time. Because once you're at test time, it kind of implies that a user is waiting.

Something is waiting. It's either another process, a user, an event, and there is real cost to like every single token, um, or like every single second you wait. Whereas at sleep time, just by definition, there is no real cost.

The only cost is like maybe energy. So I think, I guess that's maybe like one axis that this plot isn't evaluating, but, but yeah, that's, you know... That at a high l- high level, I'm just trying to say like to evaluate cleanly against Test-Time Compute, I think you really need to talk about efficiency.

Um, but I think Sleep-Time Compute is a much broader concept beyond just these empirical results. So I think that's what you're seeing, that's why you're seeing so much interest in like the new ChatGPT memory features and like memory's getting shoved into everything, right?

Because there's a lot of interest in just doing more in the background of the user. Like not just having this completely like passive interaction where the user has to say something and the assistant answers back, but just having kind of the, the machines run all the time.

And I think that's another aspect of like what makes something agentic. Like not having to have a user send an event to trigger the machine to turn on, just allowing these machines to run all the time.

Alessio22:47

And do you see this being more useful when the initial response is actually much faster? Because if I think about the way I interact with LLMs, if it's streaming the response, I'm kind of reading as it comes out, and then I respond right away.

Latency Matters22:47

Alessio23:00

There's like not as much sleep time, so to speak. But when you have like very fast completion, then there's like reading time, and that's kind of like sleep time that you can, you can leverage. Yeah, I'm curious because you have 4o mini and you have 4o.

I'm curious like if you think that makes a big difference or not as much.

Kevin Lin23:17

Yeah, I mean, we also have some results on like the, uh, like o1, o1 mini, where those ones actually will be spending minutes. Uh, so we, we kind of have like a couple of different flavors of experiments, both with like the 4o style and the, the-

Alessio23:33

Nice

Kevin Lin23:33

... the, the like proper reasoning models. It, it probably, like the latency is probably more of a like a real problem with like some of these like o3 mini, where you're asking it like very hard math problems, and you have to wait, and they, they even like hide the reasoning chain so you can't read much while, while you're waiting.

Swyx23:53

I mean, you're just measuring tokens here, so that's fine, right?

Kevin Lin23:56

Yeah. So this is the output tokens. Yeah. At test time.

Swyx24:01

Okay. Yeah, like, uh, I think broadly cool. Any other results you care to highlight that people should know about ways in which you explored, uh, this?

Coding Results24:01

Charles Packer24:11

Yeah, I think the, the coding ones are pretty exampl- are, are a pretty good example too.

Swyx24:16

Okay.

Charles Packer24:17

There's some stuff on Sweet-agent, I think

Swyx24:21

Sweet Agent's always a fun one.

Charles Packer24:22

I guess not Sweet-- Yeah, like the Sweet Agent domain. Oh, yeah. I guess this is like an interesting figure to touch on. I mean, I think this is kind of what Kevin was talking about, where, you know, really trying to be scientific about like when do we think Sleep-Time Compute will actually help.

And I think it's, it's all about, you know, how much can you predict what's going to come next. And if you think you can predict what's gonna come next, then that kind of means that you probably have a good chance of like reorganizing your thoughts in advance to prepare for that.

But, a-and yeah, basically you, you can do this sort of evaluation, um, by kind of evaluating like log probabilities. But yeah, Charlie, I don't know if you wanna highlight anything on this graph in particular.

Charlie Snell25:05

Yeah. Basically, like we measured how predictable is the question from the sort of context, and we find that like as the questions become more predictable, the benefit you see from applying Sleep Compute kind of widens. So I think like this sort of indicates...

We, we don't do this in the paper, but this sort of indicates that like if you were to apply some kind of policy to prioritize when to apply Sleep Com-- more Sleep Compute and when not to, you might want to use some sort of metric like this to help you, uh, make those decisions.

Charles Packer25:37

Yeah. So it kinda indicates that you can make this agentic. Um, you can actually have like intelligent decision-making about when to stop, like when to finish Sleep-Time Compute and when to just kind of like yield your process.

Swyx25:48

Any, any observations between models?

Charles Packer25:51

Oh, yeah. I guess the, the o1/o3 stuff is pretty interesting, right? Because, yeah, I guess o1's kind of an outlier to a certain extent. Um, I guess we have like some hypotheses about why that might be the case.

Charlie Snell26:02

Yeah. I'm, I'm not entirely sure why o1 is like less of a Pareto shift than o3-mini. My guess would, like it could be some like detail in how they post-trained it or something where it's, you know, not as favorable on some of these Sleep-Time tasks.

Outside of o1, the other models, uh, like tend to show like a pretty significant Pareto shift on this task. Yeah.

Charles Packer26:27

Yeah. It's like pretty consistent across like both 3.7, DeepSeek, o3-mini, which all like... The way you actually scale the X-axis here is fundamentally quite different in each case. With 3.7 extended thinking mode, the parameter you provide to scale it is different from like o3 reasoning effort.

And then to, I guess, enable this sort of like... To send, to tell like R1 on the API whether or not to like go to 10K tokens versus like 2K tokens, um, also is like a, you know, different mechanism.

Kind of like irregardless of how you like push the agent to go further and also like kind of cross, across frontier mo-model categories, this sort of observation kind of holds that, um, yeah, it's kind of like a Pareto, Pareto benefit to apply Sleep-Time Compute.

Swyx27:14

Yep. Got it. Yeah. Perfect. Cool. I mean, I think everyone should go check out the paper. Uh, by the time when we release this, I think you'll have a full blog post and everything, and, uh, we'll be able to dive in deeper.

Memory Systems27:26

Swyx27:26

Uh, I guess as a parting thought, any thoughts on memory and sleep type implementations that you've seen out there? Uh, obviously, you mentioned the ChatGPT one. First of all, any, any thoughts on the, the improved ChatGPT memory? And then second of all, uh, any other memory and sleep implementations that you like out there?

Charles Packer27:44

Yeah. I mean, like K-Kevin, you've been like playing around with the ChatGPT memory stuff.

Swyx27:48

And by the way, we can, we can cut the, the screen as well to-

Charles Packer27:52

Oh, yeah.

Swyx27:52

Save the thing. Yeah. So just comments on ch- on memory and sleep im-implementations out there that you want people to, you want to point people to, including Letta, I guess. You know, obviously you guys have your own.

Charles Packer28:04

Yeah. Well, Letta definitely, I think, you know, part of what we're releasing with this paper is actually like a pretty robust implementation of these ideas. I think the average kind of developer, they're not necessarily interested in Amy or like GSM 8K, right?

What they're interested in building is like a really immersive, you know, AI experience. The median use case, you know, generally being chatbots, um, or like chat with your data or like, you know, personal assistants, that sort of thing.

Um, I think in those use cases, like this is a really intuitive idea. You know, unintuitive that you would not do this. I think if you brought like GPT-4 back in time 10 years and like you didn't really explain how it worked, and then you just kind of showed it to somebody, they'd be kind of surprised if about like all the inherent limitations, right?

About like context window management, about lack of memory, about, you know, why doesn't it... You know, this is clearly AGI, like why isn't it like remembering or like thinking when I'm not chatting with it, right? So yeah, I think it's, I think just generally this sort of design, like applying compute not, you know, at direct interaction time or test time is just gonna be kind of the norm, the norm.

Um, and I think that the algorithms for like what kind of compute you're going to do will definitely be varied per application. But also, I think there are just like general scaling laws that will hold. Like generally speaking, you know, we're gonna observe that like you, you kind of have, you know, limitations.

There's like Sleep-Time Compute across the board will improve, improve performance until you enter a regime where you just do not care about how much test time compute you burn, right? But yeah, I think the architecture like we're releasing, um, we have two different ones we're releasing.

One is like more chat focused, so it's basically like MemGPT V2. So chatting with an agent, but the main agent isn't actually managing the memory, and you have like a very aggressive like background process that can be scaled to like many agents.

It can be scaled. It can just be one, it can be like 10. Um, but those agents will aggressively like kind of place tokens in and out of the context window of the main agent. Um, and I think that idea is actually super cool for voice or like other domains where you have low latency because what that means is like this context optimization, this prompt compiler, this LMOS, whatever you want to call it, you can kind of add additional constraints on what you want it to do.

So you could say like, "Hey, this Sleep-Time agent should be assembling the memory of the main agent, but it should be doing it with a hard constraint of no more than 2K tokens because we want time to first token to be less than X."

So yeah, I think this, this general design has like a lot of really cool implications for chat and for voice. And we're also releasing another kind of like parallel implementation that's more for documents. And it's kind of the idea that you can click upload on a document, kind of like cloud project style, and immediately start chatting with an agent.

But you have a kind of an army of agents working in the background or just one, however many you want, that will parse the document kind of incrementally and update a memory block that the main agent can have access to at any time.

So it's also kind of like an anytime algorithms. We'll always have the latest refreshed copy even while like the subprocess, the Sleep-Time agent is running. I think those two-- Yeah, I think similar to MemGPT, I think they're, they're definitely like very good reference designs for just what's coming next.

I think this sort of thing is just gonna be like the norm in like a year or two years.

Alessio30:54

Cool. Yeah. That was great, guys. Thanks for the preview on, uh, apparently the paper is going live in three hours on arXiv, so, uh-

Closing30:54

Charles Packer31:01

Yeah, it'll be up soon.

Alessio31:03

Um, yeah. Any-

Swyx31:04

Yeah. Good luck with that

Alessio31:05

... yeah, any, any other call to action to wrap? I don't know if you're hiring at Letta or if you want people to contribute or help with anything.

Charles Packer31:14

Yeah. We're hiring both on engineering and research. So yeah, if this sort of stuff piques your interest, both from like an application perspective or like a research angle, um, definitely reach out. And yeah, I mean, I think, you know, MemGPT was like one cool idea we had.

I think the Sleep-Time stuff also we're like super excited about, similar to MemGPT, the actual practical implications. And yeah, it's just one of many to come. I think we have a lot more, a lot more ideas for like fundamentally the ways that we should be building kind of these compound systems, these agentic systems around LMs.

Um, so yeah, tons of more like really fun research to come.

Swyx31:46

Yeah. I, I think the, the-- in terms of messaging, the kind of message that Noam Brown did for Inference Time Compute was a very linear message. It was basically he was like, "One second of inference time scaling is worth 1,000 or 10,000 seconds of pre-train," right?

And so you might wanna get your-- distill your charts down, because the charts obviously have different things at different scales. But distill the charts down to like, okay, how-- what is the value of one more second of sleep versus, versus, you know, test time?

And, and that, that could be an interesting ratio for people, just for communication purposes. Obviously, it's more complicated than that.

Charles Packer32:24

Yeah, yeah. I guess that would be like slicing the charts on an X-axis point.

Swyx32:28

Exactly. Cool. Well, thank you so much. Looking forward to the post and also the rest of y-your work. I think there's just a lot of interest in stateful agents in general, and I think this is a core piece of it that I'm interested in exploring, and I'm sure there'll be more.

So we'll talk about that soon.

Charles Packer32:44

Yeah. I mean, I think, yeah, one last comment is, I think with the stateful agents thing in particular, I think one interesting thing about scaling Sleep-Time Compute is you need memory for it. Um, you just fundamentally cannot do this sort of scaling without memory.

Uh, memory is like, you know, kind of table stakes if you want to do this sort of thing.

Swyx32:58

Yeah. I think then the other thing was als- I also wanted to shout out like LangMem is, is relaunched from LangChain as well. That's another interesting implementation. I don't know if you have any differences or comparative with them.

Usually, people don't look at them, look at the competitors.

Charles Packer33:11

Yeah. Yeah. We-- I, I saw it in the context of your, the tweet, the one on, uh, Sleep-Time. Yeah, I think, you know, there's definitely a lot of implementations of like the background agent idea. I think, you know, similar to MemGPT, there's, at the time MemGPT came out, lots of like concurrent implementations of similar like function calling, memory editing ideas.

But yeah, I think that's part of like the reason with this paper, we want to be like pretty precise because we think it is like a scaling direction important enough to kind of var-- like be precise about similar to Test-Time Compute.

And I think, yeah, once you do that, then I think it becomes also more clear like how you would architect, like if you're LangChain, how are you going to architect the background agents to be very well engineered? Like, because now you kind of have a rough idea of like, you know, what the limits are, how far you can go, um, from more of like a scientific or empirical perspective, and that can like really inform the agent design pretty well.

Swyx33:57

Awesome. Cool. Thank you so much.

Charles Packer34:00

Yeah. Well, thanks.