Origins0:00
Okay. Here we're, we're here in a remote studio with Dhravya Shah of Supermemory. Welcome to Lane Space.
Thanks for inviting me.
Yeah. Uh, obviously you've been blowing up on, on the timeline for multiple years now. I just found out you launched Supermemory in 2023. It feels shorter-
Yeah
... than that, but also, uh, you've been doing this for a while.
Yeah. I've been in this space for way too long. Supermemory has been many different products, but now it's everything at once. I'll, I'll get back into that, but yeah. Lots of things.
Okay. So, so what, what is this lore? Can, can you... Uh, we looked back at your first product hunt launch. You can show us what was it originally?
Let's, let's walk back from there. So Supermemory launched as this consumer app that was supposed to be your kind of, uh, better bookmarking or note-taking tool with a few other features. So, like, for example, you can take your memories across LLMs, and this is from very long ago.
So we had a s- uh, an MCP launch, and then we had another launch for the app itself. So earlier, obviously, it was just the app, which was like a save... You can save things to the app, you can get things from the app, very simple as that.
And then we were like, "Okay, MCP is a thing now, so we should probably make an MCP out of it." And this was a time when I personally was doing internships. I was working at Cloudflare at the time, and this was a side project that I did as a part of this thing called Build Space.
And, uh, Build Space, uh, you might know about Build Space. It's f- it was like this place where a lot of builders used to build things nights and weekends, and they really pushed you to building something and, like, talking to users and stuff like that.
Yeah, it's like a Gen Z accelerator that, like, they... The founder just gave up and just, like, uh, moved on.
Exactly. So I was in both season three and season five of Build Space, and I open-sourced Supermemory during Build Space, and it, like, the project absolutely blew up. It got, like, you know, it was one of the biggest projects, like, fastest growing projects of 2024.
And I, I kept tweeting about how I'm running this on Cloudflare on $5 a month budget, and it was still a consumer app, but then I kept adding features to it. So earlier, the way we... the world used to think of memory and how we used to think of memory was just RAG.
So you just embed it, put it in a vector database, and that's all you need to do. And now, then I, like, dug down the rabbit hole and so f- so vector databases, like, at that time and even now, except a few, like, like TurboPuffer and Chroma, were not really scalable for the scale I was getting to, where they were either getting too slow or too expensive to run.
And so I had to, you know, find other creative ways of making it really cheap for me to run. And I kept tweeting about all of this. So my tweet was like, "I run Supermemory on $5 a month, and I have 100,000 users, and this is my architecture."
So it was, like, an entire diagram. And that blew up. That was the first kind of, you know, uh, famous Supermemory event, and that led to not people being interested in the app itself, people being more interested in the infrastructure that I was building.
Um-
Yeah, I remember you open-sourced it, and then I was like, "Oh yeah, I need to star and fork this-" "... 'cause I'm like, this is... Like, if I ever need this, I need, I need this, uh, setup."
Exactly. Exactly. So after starring it and, like, after that open source, you know, it, it blew up to, like, 10,000 stars in a few weeks. But then I realized, like, you know, a lot of companies were reaching out to me.
They were like, "Hey, can you help us set up our memory thing?" And I also went and worked at other memory companies, and it became, like, this entire thing that, like, I was doing pretty much a lot of consultation work around context engineering.
Um, and I realized, like, all of these have very similar use cases, but with slightly different flavors, where they, they not only need retrieval, they need, like, ac- like, true user understanding, and we'll get, get to that as well.
L- like, what is modern memory? You know, like, how, how should AI remember? And then I was like, okay, so to make an LLM truly good at understanding a user, you have to handle four things. One is knowledge updates.
So you have to invalidate stale knowledge and build on top of it, which is different from just storing vectors. You have to have some sort of temporal reasoning, so you have to have, like, give the agent an understanding of time and how things have passed for personalization, and you need some sort of forgetfulness to forget things that are not relevant anymore at all.
But you also need this concept we call user profiles to always, like... Memory is not just a retrieval call. The LLM has this very small profile of the user that it will utilize on every single turn. So I, I was building all of these components into Supermemory and, and all of these apps, and I realized that, okay, there are some really difficult infrastructure challenges to be solved here, which is both on the extraction layer, which is how do you truly, like, you know, how do you train a small model that can extract things about the user and maintain the second layer, which is the knowledge graph?
And the way we, we built our knowledge graph is different from how people usually build knowledge graphs. So people usually have this concept of triplets-
Triplets
... where they have-
Object, predicate, object, yeah.
Exactly. Object, predicate, object or entity, relation, entity. And we think that triplets actually lead to worse performance because you have to traverse them a lot- ... to get to any information. And, like, you can just think of this-
This is the, this is the learning of every graph guy ever.
Exactly. Exactly. So you have to, like, to, to, to find out, like, okay, what do, what food does Dhravya like? You have to first get to the entity Dhravya, and you have to traverse one or two layers deep to find out, like, okay, Dhravya likes food, Dhravya enjoys this thing, et cetera, et cetera.
And that leads to extremely slow and, like, this is just unnecessary. So we built our own knowledge graph. We built our own extraction pipeline, and I was like, "This consumer app is probably not useful anymore. This core infrastructure is more useful now," and that is our core business now, and that's what we offer.
Amazing update. Um, I think- One of those things where, like, I think we can get more into it. Uh, what, what is Supermemory today? Like, is it a cloud business? Is it a VC-backed company? You know, you're, you're a young founder.
Like, tell people more just about the company itself and, and what it is today, because obviously we spent a bunch on the-
Yeah
... the background. Yeah.
Yeah. Uh, well, I, I should have done this before, but, uh, all right. I'll give a little bit of an intro about myself first. So I'm Dhravya. I'm currently 20 years old. Before this, I was working on agents and databases and a bunch of other things at a few companies, including Cloudflare, High Fury and others.
I had two companies acquired and I came to the US. So I built Supermemory, again, as this thing, raised $3 million, um, last year, and since then we have been a VC-backed cloud business that's mostly open source, actually, and we are trying to build a context infrastructure for the AI era.
And so what do I mean by context infrastructure? Whatever external context or, like, that the agent would need, including just retrieval to, like, user understanding, memory, et cetera, like self-improvement, self-learning, personalization, we provide all of these in this one single tool, which is Supermemory.
Because the way we structure the data, we can automatically provide all of these facets because, again, every single memory use case would require a different way of implementing it. Um, and so we do that for developers. We provide an API for people to use, but we also still have a consumer-facing app, which the, the idea is that you have your own memory that you can utilize across everything you use, or you can add Supermemory to your agent to make it much better at remembering things.
So, for example, the OpenClawing was blowing up or Cloud Code blew up as well. It got like 2 million impressions. Um, both of them were kind of this consumer-facing, like, plug-ins that we offer.
Yeah. It's, uh, it's really impressive to watch. I think a lot of people... I'm gonna put OpenClaw in the, in the title of this thing. Let's-
Yeah
... let's go right into the OpenClaw thing.
Let's do it.
You know, there's, there's, there's, uh, there's more to go into Supermemory. I think more people can check it out. You've published a lot of stuff, but the topic of the moment is, I think, was it, uh, Peter Levels or something?
Somebody was having problems-
Yep
... with OpenClaw memory.
Everyone. Everyone has been. Yep.
And then they're like, "Well, okay, how do you do it?" And then, and then there's, like, all these different answers and I'm like, "Wow, this is really not solved."
Yeah, exactly.
Okay. So, so tell everybody what, what, what's going on.
Memory Sucks8:26
All right. So let's first talk about how OpenClaw currently handles memory. And I have, like, this little claw artifact that I prepared for my article. But so the way OpenClaw has with QMD or with whatever memory plug-in you use, it inherently relies on tools to search through these memory.md files that it prepares.
So, you know, like, what did I decide about the API? Then agent will decide to search, and sometimes it won't decide to search, which is probably the biggest problem here. And then it will get some results, it will do some scoring, and then it will return it back.
So, and the search right now is through the memory.md files or the memory folder. So if you-- if the agent decides to not remember things explicitly, it will never show up in the results because you don't know what to remember at ingestion time.
So there's a bunch of issues here, and because these files are static, they don't handle updates, they don't handle, uh, like the forgetfulness I was talking about. There's no temporal reasoning. You cannot look at particular parts of the file because you don't know what is where and other problems like that.
This, like, you know, this is the reason why everyone is complaining about OpenClaw's memory.
This to me is surprising because when people complain about OpenClaw's memory or when OpenClaw even did the memory, people didn't, like, look into how exactly and why it's bad. So we did all of this, like, digging in and we were like, "Okay, what is the right way to do memory?"
And we already had these plug-ins for Cloud Code, for Open Code, so we had a lot of learnings about how agents should be stateful. So it... We basically turned the tools-based approach to hooks-based approach that actually is reading from the Supermemory graph, which keeps the content fresh.
So it handles updates, uh, et cetera, et cetera. So now, you know, for every tool call or for every user message, a hook automatically will put less than 2,000 tokens worth of information dynamically. Um, so the entire context will not have more than 2,000 Supermemory tokens, and optionally, it could also do a tool call if it needs more information.
And we have, like, contradictions resolution, we have temporal reasoning and all of that built in. So the information is always truly fresh no matter what. That's the gist of it, um, but we can dive more into it, uh, because, like, there's some things that are not in this diagram that we also do.
So yeah, uh, that's kind of the simple why Supermemory is just better.
QMD is, is from Toby Luca, which is, like, really super impressive. I, I think it's a very interesting... So basically, Toby wants to do everything locally. You-
Yeah. Makes sense
... you obviously are Cloudflare maxi, so you're gonna do on Cloudflare.
Yeah.
So obviously that's, like, a fundamentally different thing. I think hook-based is probably right because obviously models make mistakes, and I get very, very frustrated when, like, I use Groq in xAI and it doesn't search tweets. Like, I'm... Dude, there's only one reason I use Groq, is to search tweets.
Yeah. You have to, like, ask it, like, "Please use these tools." And most of the OpenClaw users don't even know that tools are a thing, so they don't want to think about the tools. So yeah.
Do you think... Do you perceive that they're non-technical? I mean, is-
So-
You know
... especially, you know, a lot of people are, have been seeing they are technical in the engineering sense, but they're not technical in the agent's world. I think we are in a huge bubble right now.
Yeah.
And, like, they won't know that, okay, this is exactly what's going on behind the scenes. Like, they will not even look into it Um, so yeah.
It- it's just curious because as an engineer, you should... This is, obviously you and I think that this is-
Yeah
... the only thing int- that's interesting, but-
Yeah
... uh, you know, there's many other engineers who have realized. Yeah, so, so, okay, I, I think other than this, there's another memory discussion that I think Cursor and Claude, those guys have, have been saying, which is like sort of file system-based memory, right?
Like where-
Yes
... where you just, you put, you dump a bunch of whatever .mp into your file system and then you grep over it. And OpenClaw has some compaction, right? Like some, some data-
Yep
... journal, some heartbeat stuff. What are your takes on all those things?
So I think file system memory is actually the right way to do things in a lot of cases. Like, for most Claude Code users, that is probably the right thing to do. But you still have the same problem of you have to know what to remember in order for it to remember things.
So you have to explicitly mention that, "Okay, please remember this," or, "Please bring this up." And it's limited by the files themselves. So the files can get longer and longer, so that becomes a problem, so you have to split between many files, and the files still don't have any update logic, et cetera, et cetera, unless you explicitly have the agent do that.
They're also really slow to traverse through, so you have to do agentic discovery to find out any context to answer any question. Like for example, if I have something like, "Implement this workflow," it will have to, you know, first look through everything it knows in the memory, all the workflows, et cetera, et cetera, and then it comes with the answer.
Now, I think this is a very hot take, but I think that they could, like, one is personalization, which I'm talking about, but then there's also code indexing, and they could index the code and Cursor does, but Claude Code does not.
And I believe this is also because they are not incentivized to utilize less tokens by the end user because they're just offsetting the cost of indexing, et cetera, et cetera, to tokens and time that the user takes to run the agent, which is, which is what they want.
So theoretically, you could add like code indexing, you could add personalization. It would be much better and faster. And I'm not just saying it. We actually ran a benchmark just yesterday. I published this blog. I'll share my screen again.
Um, I published this like extremely technical report, which also kind of blew up, but we ran a benchmark on LongMemEval against three things. Oh, the screenshot is in one of the code tweets, I think. Uh, how do I view codes?
Twitter UI keeps changing. Here. Um, so the Claude Code one performed the worst, and OpenClaw slightly more than that, and Supermemory is the highest. And you can see that, you know, it's like a pretty significant difference, like almost 50% in most cases.
Yeah. Let me, let me clarify. This is your reproduction of the Claude Code method and your reproduction of OpenClaw, or you actually use Claude Code?
Yes. No, it is our reproduction, so it's like pretty much exactly-
Yeah. I think it's important to say. Yeah.
Yeah. It is pretty important. And, uh, you know, we also think that benchmarks should not be trusted and cannot be trusted. So, like, you know, you should not listen to me, actually. So we built this open source kind of benchmarking evaluation platform, like Groq's OpenBench, it's called MemoryBench, which I'll also share my screen.
So we... Like, this is our implementation of exactly how it's doing. In fact, we also literally, like, went into it. We, uh, implemented OpenClaw's hybrid method with the same formula, with the same everything, and you can just view the PR to find out exactly what's being done.
But in short, this MemoryBench is a way to benchmark any providers against any benchmarks against any judge on the same kind of the base rules. And this lets you quantify all the things. So not only the quality of how the memory system performs, but also the latency, cost, and other, like, you know, heuristics like the top K recall, NDCG, et cetera.
So this is like a really good way of doing, you know, benchmarks for memory, and it makes it really simple as well. I actually ask every single competitor to use this because it actually makes it easier for our competitors to compare against Supermemory as well.
Yeah, I think that's, uh, very good infrastructure. Okay, yeah, people should try that out. Other than, uh, just, just while we're on the topic of benchmarks and all that, there's LongMemEval. Um, what's not so good? Whatever you want to say.
Yeah. So I think we are still not there in the memory benchmarks space, and there's a lot of reasons behind this. I might also write a detailed post about this, but I think that LongMemEval is a very good benchmark because it calculates, it tests for the right things.
Benchmarks16:30
It tests for the right things. The data set is great. Everything is really good in LongMemEval, except for the fact that you can win LongMemEval by remembering as many things as possible. Like, there's a direct correlation to the number of memories extracted and the score you get on LongMemEval.
Um-
Okay. That doesn't seem bad
... because, yeah, it's not too bad, but like, in real world, you would spend more cost. Like, in real world, you don't want to extract everything because extraction is not free to win against the benchmarks. So it's good for, like, some use cases, but in the real world, that's not how people expect memory to work.
So we think that that's fine, but LongMemEval also does not test for forgetfulness, because how would you test if something should be forgotten or not? It tests for updates, but not forgot-forgetfulness, and a few other things. So we still use LongMemEval because that's the best we have.
Then there's LoCoMo, and LoCoMo is pr- like really bad because like... And, and people actually, it's not just me, like pretty much everyone knows that LoCoMo is not the right thing to, to judge on memories.
No, no, I don't know why. Why? Tell me why, tell me why.
So the, the reason is, like LoCoMo is essentially testing for retrieval capacity, and it's not testing for the real memory things. Like, you can essentially just context dump everything in LoCoMo and get like 100% in the latest models.
So one is the, there's only 30 questions, like the, the data set is really small. Second is it does not test for like knowledge update, temporal reasoning, multi-session, all of the, the other things, like the important things, like preferences.
It doesn't test for anything but retrieval. And, and like in Gemini Flash or whatever, like you can just literally... The, the questions, the sessions themselves are short enough that you can put the entire session in these context lens and it will work right away.
So I think, and it's also very old now, like in that age it, it probably did make sense, but now it's been three, four years since that, and the models have gotten much better. There's also ConvoMem, which I actually like.
It's a really good one. It's by Salesforce, and it tests for more assistant-y use cases like, like assistants should do both RAG and memory, et cetera, et cetera. So that's pretty good. And sometimes you need both, which I also want to come to actually, but that's what I think about the, the, the benchmark space.
There's one important thing, which is there's no concept of benchmarking user profiles, so all of these benchmarks are very retrieval heavy.
Like factual information.
Yeah. Like you, you essentially want to, you know, fetch something, get an answer, answer the question. But that's not all use cases. I- I'll show you one, a slide real quick, which, which will make this very clear. I have it ready right here, luckily.
So traditionally, you have this retrieval thing that happens, so you get a question, you retrieve something, and an answer is generated based on that. But we think that it will not work for most non-retrieval questions. Like, find the best monitor for me.
If I've never spoken about a monitor, which I probably have not, it will not return anything. It'll be like, oh, it'll give some generic answer. But agent probably needs to know a lot of default things about me at all times, like user is a CEO, but it also needs to know some episodic information like user recently raised fund- fundraising, moved into a new office last week, et cetera, et cetera.
And in which case it will be like, "Oh, you code all day and you just moved into an office, so this is probably the one that you should buy. It's a little expensive, but this is the one for you."
And this can only be done if you really have like a very extremely small, like 500 token information about the user that captures both the static things and the episodic things, like some recent episodes has, that's been going on.
And we also have this built into Supermemory.
Yeah.
But there's no way to benchmark it.
Yeah. It's basically personalization, right? So it's a bit hard. Okay. Uh, you know, I, I just want to close the loop on anything OpenClaw related. I think people are using a lot of things. What other issues crop up?
Like things like secrets, things like context confusion if they've done like similar things, you know, like there's a, there's a pr- trade-off between retrieval and accuracy or like recall-
Yeah
... and accuracy, right?
Yeah.
Do you, do you, do you find these kinds of things and, like what are people doing?
Yeah. So I think like, so the way OpenClaw does it is it essentially sends back the last 15 messages in the conversation, and it essentially uses that back and forth. And I mean, the approach itself is not ideal because you will, like you are not doing any, like you're not utilizing any cache tokens.
And because you're not utilizing the cache, you are essentially paying like insanely, like 10X more than if you would do it the right way. So that is always a problem. So like we think so that... But that's like a OpenClaw issue.
You, you know, we cannot really do anything about that. The other thing is people are giving it access to their secrets and stuff like that, and we have to, like, as a memory system, we have to like, make sure that no secrets end up to us, like in any case.
Um, a lot of users in OpenClaw, Claw also have like really, like browsers attached and stuff like that, and heartbeats and heartbeats are all going on, you know, every hour, and the heartbeats are also ending up in memory causing some pollution.
And some people want to use their OpenClaw memories along with their Claw code memories, and we actually want-- don't want to do that. Like, it should be siloed, like it's their own things. But people want that, so we are offering configuration.
Recall-wise, like you know, Supermemory sometimes does not end up giving the right results, and that's why we recently shipped a new thing for OpenClaw, so it was a lo- a huge learning experience for us. It's called hybrid mode.
Hybrid Memory22:26
There's always this trade-off between what you want to remember in a memory system. You always have to choose exactly what you want to remember. Which is important because you want to keep them updated, you want to keep them fresh, you want to forget things.
But sometimes the user, you don't know what the user will ask about. Sometimes the user will say something that was not remembered or was not very significant. So how do you handle those cases? So the way we do it is we have, because we do both memory and like managed RAG, we can show up the memories if they exist, but we have a fallback to RAG answers.
Like it will just return the raw chunk if Supermemory did not remember it. And then in the background, it will add that to the memory as well. So this, this essentially gives you like, you know like, we say that Supermemory should almost never forget anything because of the fact that we are always giving the agent the right context as long as the user is asking for it.
Um, so yeah.
Amazing. Are we going to the hybrid stuff?
Yeah. So this was the hybrid stuff. So essentially you, we'll return the memories first, and then if there's any raw chunks that match up, we also return those to make sure the agent knows just enough information to answer the question.
I, I-
And no other provider does this right now. So yeah.
How s- how important is cost? Like, my perception is people are very cost insensitive, but you obviously like... Anyone who is like a real engineer um, does care about cache hits and all that. Uh, I, I don't know.
I don't... Do people care?
Yeah. I, I think people care a lot about cost, actually.
Okay.
Like I think we, again, we are in a huge bubble and- ... a lot of our users, especially like, you know, they do care about the cost because, like, because people would compare a memory system against the LLM they use.
So if the memory system is, you know, $10 for a million tokens and the LLM they use is $2 for a million tokens, it will be like, oh, like this does not make sense to use. It's too expensive.
I should not use this. And this happens to us, and we want to make sure that, like we are giving them a lot of, you know- ... lot of satisfaction on, like, they should be happy about using Supermemory and it should not be something that, you know, they have to be concerned about.
So for us, we have optimized our infrastructure down to it costs us two cents for a million tokens because we have our own model, our own database, our own everything. So we have to price our users also really low, which we do.
So yeah.
Okay, cool. That's a really good sort of masterclass in, uh, just everything sort of memory, uh, state of memory, in- including a side tangent on, on evals and all that. What's next for Supermemory? Like, what is sort of the open directions, uh, for 2026 that people should think about?
Outlook25:05
Yeah. So 2026 was the year we realized that the prosumer aspect of Supermemory is something that people are really excited about. Like, this idea of having this one memory that you can connect everywhere, you can interop between providers and stuff like that.
And we are continuing to do things like that, like we're going to launch the plugin with Cursor today, and we already have the Claw, OpenClaw, OpenCode, all of those. That's good. Next up, we are going to do a launch on voice agents, partnering with pretty much every voice agent company because of the fact that Supermemory is so fast at, at the work it does, so it's really good for voice.
Um, that would be amazing. We are doing more on evals. We're building our own eval. We are going, going deeper into personalization because we think that no one has done a lot of work in personalization yet and a lot of other ex- Like, we are building a new version of our, of our database.
So yeah, it's, it's going to be really exciting.
Cool. Well, thanks for jumping on. I, I think this is, uh, a great introduction, at least for Latent Space listeners, to Supermemory. Obviously, like, you've been plenty famous on your own with your, with your projects, but I, I think I already invited you to Worlds Fair.
We'll see you, like, do more talks-
Yeah
... for, for that kind of stuff. And I think just generally I'm starting to see real startups and people taking memory seriously compared to two, three years ago when it was all, like, vector database.
It was a joke.
Uh, I think it's, uh, it's really good. Uh, there's, there's more coming, you know? Like, I think LangChain has their own take.
Yeah.
You know, I think all, all the others, uh-
There's Letta, LangChain, Memzero, Zepp. Yep.
Both Letta and Zepp have done workshops, so I think it's your turn. Uh-
Yep. Amazing.
So it's, uh, looking forward to that. Okay, cool. Um, good to catch up and, uh, I'll see you online.
Great. Thank you so much.





