LALatent SpaceApr 18, 2026· 48:53

⚡️ How to turn Documents into Knowledge: Graphs in Modern AI — Emil Eifrem, CEO Neo4J

Emil Eifrem, CEO of Neo4j, joins host Shawn Wang to argue that AI systems need more than top-K chunk retrieval—they require graph-shaped context for accuracy, explainability, and developer productivity. GraphRAG combines vector search with graph traversal, starting semantically then expanding through relationships, which he says yields higher accuracy and auditability than opaque vector spaces. Neo4j now powers production AI at companies like Pfizer (over 60 million documents), Novo Nordisk, and 20 of North America’s 20 largest banks, with a mortgage lender seeing a 20% conversion lift. Eifrem describes four data sources for agentic systems: operational databases, cloud warehouses, agentic memory, and context graphs—the latter encoding institutional decision traces. He notes a recent shift where enterprises lead with generic text-to-Cypher instead of specialized functions, and highlights Neo4j's new `create-context-graph` starter kit for 22 industries, built to bootstrap context graphs and agent memory.

  1. 0:00Intro
  2. 1:55Why GraphRAG
  3. 11:06Use Cases
  4. 18:42Cypher & LLMs
  5. 23:18Context Graphs
  6. 32:58Startups vs Enterprise
  7. 35:01Getting Started
  8. 45:04Closing Thoughts

Powered by PodHood

Transcript

Intro0:00

Swyx0:04

Okay, we're in the remote studio with Emil Eifrem, uh, CEO of Neo4j. Welcome.

Emil Eifrem0:09

Great to be here.

Swyx0:11

Uh, Emil, you-- I've been waiting for this for a while. Uh, you were one of our top speakers at the first Worlds Fair, and then, then we did the full GraphRAG track, uh, with the rest of your team last year, and we're, we're back this year, uh, focusing on GraphRAG.

How do you introduce Neo4j today?

Emil Eifrem0:24

It depends on the audience, right? Like, I think we're most well known as a, as a database, right? As a graph database specifically, right? But it's a broader platform than, uh, just a database these days. And I tend to start with the, you know, we transform data into knowledge, right?

So it's a platform to do that, and then of course leads down the path of what do you mean with knowledge? 'Cause that's kind of an ill-defined term, and I talk about how actually it's ex- about extracting the signal out of the noise and express that in a way that is a knowledge-dense way, and the graph is, is, is one of them.

And it... At the core of it, of course, is the, is the database, but then, you know, we- e- there's many other things in part of the platform as well now these days.

Swyx1:06

You f- you guys have been around for a bit. You're used by basically everyone, including Transport for London. Uh, which will be fun.

Emil Eifrem1:12

It's a very graph thing. Like, you'll, you'll see when you get there, the tube network-

Swyx1:15

Yeah

Emil Eifrem1:15

... is a graph.

Swyx1:17

Oh. Oh, yeah, of course. Yeah. Well, you know, depending how graph-tilled you are, everything's a graph, right?

Emil Eifrem1:22

This is true.

Swyx1:23

And I've been down that, that rabbit hole as, as well myself. Yeah, when I was... So when I learned to code, uh, 10 years ago, I did a, did a coding boot camp and there was a workshop at some conference I attended that introduced me to Cipher and Neo4j, and I think that-that's how a lot of people maybe first come across, uh, Neo4j.

We wanted to talk about, you know, what, what's happened since and what, what is graph intelligence just generally, right? Like, what, what is the, um, I guess the, the, the rest of the system that you've been building?

Why GraphRAG1:55

Emil Eifrem1:55

Yeah. So let's first talk about kind of the why behind it, right? And so, you know, if you come in at it from like an AI engineer's perspective, right? You mentioned the word GraphRAG before, and that tends to be a popular way of describing retrieval that the R path you include- ...

the knowledge graph, right? You know? And there's plenty of reasons to do that, right? The ones that I hea-hear most loudly from, from users is higher accuracy, because you have a very rich representation of your data. And then actually, surprisingly to a lot of people, but improved developer productivity.

And there's a built-in assumption there, which is you have the graph, and we can talk about that. But when you have the graph, and if you compare that to vector space, it's very opaque vector space. If you search and find the top K documents, you don't know why.

It's like .7 in some cosine Euclidean space kind of thing, right? Compared to in the graph, it's explicit, and you can even visually inspect it and look at it, and I have an apple and an orange, and they're related because of their fruitness, right?

Whereas like an apple and, I don't know, a tennis ball might be .7 in Euclidean space, right? Because they're both round or is it because they're green? You don't know, right? So that's kind of the second thing. And then the third reason that we hear very loudly is around explainability.

People love the fact that they can actually audit why I chose these top K documents compared to just, again, opaque vector space.

Swyx3:29

Yeah. Um, I don't hear anything about query speed, which I guess it's, it's, it's part of it, but always think about graphs in terms of like, well, okay, the reason you wanna do this is because you can, you can walk a graph, you can traverse a graph, and that's probably better than doing a bunch of queries and joins.

Is that an old school way of thinking about it, or is, is that actually just another way of saying the same thing that you just said?

Emil Eifrem3:51

Yeah. I think it's probably built into the accuracy thing, right? Because frequently you can kind of cover a lot of ground, like you can, you can traverse and touch a lot of documents or nodes by virtue of doing it really fast, right?

I guess it's an interesting observation. I guess we hear less of it, and even I hear less of it.

Swyx4:08

Yeah, you never said speed. Yeah.

Emil Eifrem4:09

Yeah, exactly. Right? Like in, in kind of thinking about it from an AI engineer perspective, right, AI-ish type use cases, I guess we're also used to kind of the LLM eating up so much latency anyway that it's maybe f- kind of rarely pops up for that reason.

But it's like a fundamental reason why we can get the higher accuracy, I think. So I think that's probably why it relates.

Swyx4:31

Yeah. Yeah. I mean, if you, if you have speed, you can just throw more time at it to get more accuracy.

Emil Eifrem4:35

Exactly.

Swyx4:36

Yeah. Okay, cool. And then I think the, the other question that comes up a lot is what happened to vector databases in your mind? You know, as a, as a neutral sort of, uh, database CEO, how come that has been a less, I guess, how should I call it, persistent or independent category?

Emil Eifrem4:52

Yeah, durable maybe.

Swyx4:53

Everyone has vector indexes now, but, like I would... I think it's fair to say vector databases are, as a standalone category, are over.

Emil Eifrem4:59

Yeah. I don't know. Maybe that's overstating it. I've been... So several years ago, I was very kind of on the record saying that I don't believe this is a durable database category, right? At least not as a database category.

Maybe it feels much more like search, right? That is kind of my, my statement a couple of, uh, of years ago. I guess I've been surprised that there's still, I think, some kind of long tail thing going on where we still see a lot of experimentation where people try with vector databases early on, right?

And then at the, uh, like a super high end, then like some of the vector databases are still better than the vector search features of other databases. But like you call me a neutral party, I'm not neutral here, like, 'cause we also have vector search as part of, of, of Neo4j, and it's not as good-

Swyx5:45

Yeah, but-

Emil Eifrem5:45

... as the dedicated vector databases

Swyx5:46

... so does everyone.

Emil Eifrem5:46

Exactly, right? But like every quarter, every year, kind of the line moves up, right? And so there's less and less oxygen for them. You know, I obviously- Watched, uh, with, with pleasure, you know, your, your conversation with Simon a couple of weeks ago as we're recording this, uh, from, from TurboPuffer.

Swyx6:05

Yes.

Emil Eifrem6:05

And he described it as a search platform, I think, or search-

Swyx6:08

Yeah

Emil Eifrem6:08

... search tool or something like this, right? Um, I think that's generally where people are going right now who used to be a vector, vector databases. Um, and I think just the oxygen, you know, between everyone else adding it as a feature, they're...

Like, the good enough ends up being good enough for most situations.

Swyx6:23

Yeah. Got it. Um, I, I think, I think that's true too. And, uh, I think something that everyone should dis- disclose or disclaim is under what size and what complexity of data or cardinality or whatever, right, are you, like, truly excelling, right?

Where, where your solution is leaps and bounds above e- everyone else's because you are architected for, like, in, in a, in a specific way. At small size, who cares? Just throw it in whatever. But, like, at scale, it s- it really starts to matter.

Emil Eifrem6:51

Yeah, I agree. And, like, the way that it fits in for us is... So you have, call it your RAG corpus, right? And you have some kind of pipeline to get that in a state where you can query it and get it kind of out and ultimately to the LLM, right?

And then we populate the graph based on that ingestion pipeline and embeddings in our, you know, vector search index, right? But, and then at query time, we typically tend to use it to find the starting points in, in the graph, right?

And then we traverse f- from there. So in the classic kind of customer support type use case, Swix, you know, just got a new laptop and the permissions don't work, and so you go to, onto the Apple, you know, support site and you type in something.

Then, you know, from that natural language, right, there's typically a vector search, typically combined with some kind of BM25 type search as well, right? And that ends up finding some documents, right? Call it 100 documents or something like that.

And then you traverse from there to get the full context. And then you end up saying, "Well, okay, it's not just whether these documents were found by the vector search, but also it turns out that they were written by this author who's highly ranked," maybe through a PageRank or maybe just very simple, like stars or, or some kind of signal like that.

And, and you get to the top K, you know, something like this, right? And so it's not like graph or vector search, it's like vector search in combination with traversing the graph. That's the typical-

Swyx8:23

Yeah

Emil Eifrem8:24

... kind of pattern that we see.

Swyx8:26

Yeah. And, and, like, the engineering's actually, like, really hard here because there's so... It's just trade-off, trade-off, trade-off. Like, you can do anything if you have unlimited budget, but you don't. And so

Emil Eifrem8:35

I agree.

Swyx8:35

It's-

Emil Eifrem8:36

I d- I do think there's a general trend, though, like broadly Neo- setting aside Neo4j, like trying to extract more signal out of the ingestion pipeline to make the kind of down-

Swyx8:47

Like pre-process?

Emil Eifrem8:48

... downstream querying. Yeah. Well, you could call it-

Swyx8:51

TL?

Emil Eifrem8:52

... pre-process, but just extract more signal, like in the... When you add things into your, uh, like vector search or whatever it might be, right? In the vector databases, you tend to call that kind of metadata, right? But it really is structured data, right?

And we're part of that broader trend, and you could say that graph is, like, a very rich type of structured data. And I think that generally tends to be the right way to think about it. Like, the more work you can do kind of upstream, the easier it becomes at, at runtime or at query time.

Swyx9:21

Since you mentioned the Simon episode, uh, that was a really fun one. He's a very charismatic guy. I'm curious if you have any, like, technic- like, you know, any pushbacks, anything sounded weird to you. I can ask you about like the sort of blob S3-centric v- view of the world that he has, which a lot of database CEOs, including Neon, uh, you know, betting heavily on.

Presumably, you don't have a strong opinion there, but just... And I just want to let you respond, whatever you want.

Emil Eifrem9:46

No, I think broadly it's, it was this really strong articulation. Now, this was like years ago that I listened to it, because it was like three weeks ago or something like this.

Swyx9:54

Yeah, yeah, yeah.

Emil Eifrem9:55

Right? So it feels like-

Swyx9:56

Yeah

Emil Eifrem9:56

... forever since to kind of-

Swyx9:57

Well, yeah, yeah

Emil Eifrem9:58

... boot, boot it up.

Swyx9:58

Maybe you can't perfectly remember, yeah.

Emil Eifrem10:00

No, I can't perfectly remember it, but generally it was like a very strong articulation, what sounded like very just pragmatic and savvy trade-off and, and choices, right? And like, of course, as a database geek, I enjoyed the conversation around compare and swap, right?

And even as you and I were DM'ing around, it was like Raft versus two-phase commits and stuff like that.

Swyx10:22

Yeah.

Emil Eifrem10:22

I just kind of loved all, all that because I just remember f- for, for us, right, I used to do compare and swap like on a CPU level way back in the days, right? And then it was a big unlock for us when it...

Because we're on, on the JVM, when it got exposed through the JVM, it allowed us to do kind of, you know, you know, basically lock-free concurrency, so optimistic concurrency or optimistic locking in the da- in database, right? And it's interesting to see like how he, 15 years later, can do the same thing, but just, just based on, based on S3.

It seems like a really good set of trade-offs for, like, the use cases that, that he's targeting, which is of course very s- like very different from what, what we're targeting.

Use Cases11:06

Swyx11:06

Yeah. Cool. Uh, let's talk about some of the use cases, right? Like I, I think, uh, one of the surprising things last year was I think the Pfizer talk, which is, uh, kind of cool and, and, uh, just, uh, you know, you have lots...

You have so many customers here. Just who are the sort of AI forward customers that you, that people might be surprised are using graph databases or using Neo4j for something?

Emil Eifrem11:30

Yeah. So lots of things to talk about there. Like, that's one of the biggest things that have changed certainly since the first talk that I gave, which is now two years ago. It was, uh, June of 2024, I think, or May-

Swyx11:43

I put-

Emil Eifrem11:43

... maybe May or-

Swyx11:43

I put you on the landing, yeah, the landing page, uh, yeah, down here.

Emil Eifrem11:47

Yeah, yeah, yeah, yeah. So and a huge one of this, of course, is that people now have put it into production. It was such early days back, back, back in tho- those, those times. You men- you mentioned Pfizer.

We see a lot of adoption inside of life sciences. Right? And so this is broadly kind of scientific intelligence that researchers at these big life science companies use on a daily basis. This is access to not just internal peer research, but patents, like externals published, you know, academic papers, you know, tho-those kind of things, right?

Novo Nordisk is one of the, like, public case studies. We have here over 60 million documents, you know, billions of nodes and, and relationships. Use lots of kind of savvy NER and ER, so that's named entity recognition and entity resolution, which is such an under-discussed area right now in, in AI engineering, by the way, which I don't understand wh-why it is.

But, but that's, uh, like, key for them to make sense of all, all that data, right? And if you think about, like, a life science company, right, like extremely PhD heavy, right? And this is, like, 100% on the critical path for them in order to, like, improve their pro-productivity and, and output.

So that, that's an example. We have lots and lots of recent uptake just in 2026 in banking. An example here, um, is like, so 30% of our AI conversations these days, like in 2026 so far, has been with like global banks-

Swyx13:23

Mm

Emil Eifrem13:23

... which is pretty, pretty amazing.

Swyx13:25

Yeah.

Emil Eifrem13:25

One example, which I don't know if we have on our website, is like a massive mortgage lender company. Um, and, um, what they're doing is they have a ton of, let's call it bankers. They call them agents, which makes it very confusing.

But human beings that they hire-

Swyx13:43

Human agents.

Emil Eifrem13:44

Yeah, yeah, exactly. Human agents, right? And, and they are kind of mid-20s, low-20s kind of demographic a-age-wise, right? And like their normal tenure is super low. It's like less than a year. And so a huge part of that game is to ramp them quickly, and as they do outreach to their customers, how can you get kind of the bottom quartile and move that up, up, up the stack, right?

And so they built this huge system that looked at all prior, like kind of the, the best case path and what actually converted in the past and helped them pull that all together. And they actually talked about it publicly and didn't mention Neo4j, but that it increased conversion rates by, by 20%.

And so that all happened, uh, you know, last year.

Swyx14:30

How much money is that if you convert it to money?

Emil Eifrem14:32

Yeah, I don't, I don't know. Way more than they pay us. Um, and but the really cool thing is, and now what they're starting to do this year is they're starting to now kind of automate that, right? So prior to that, it was just kind of serving up the draft to the banker so that they can send the t- send the text message or the email or something like this.

And now of course, they're kind of removing the human in the loop, and they just send it out in an automated fashion, right? And that's very cool to see how quickly people are starting to put these things in production in customer-facing ways.

When we reviewed our kind of customer p-portfolio last summer, almost no one had customer-facing things in production, but that has shifted very, very dramatically in the last, call it three months.

Swyx15:17

When you say three months, is it... Uh, so there, there's this, uh, thesis that I've been investigating, right? Something happened in December 2025, where, uh, a lot of curves infected and-

Emil Eifrem15:28

Opus 4.5 and all that

Swyx15:29

... all that. Are you seeing that too?

Emil Eifrem15:31

We are seeing that. I cannot believe that it's related to Opus 4.5. That, that seems like a, like a separate, separate thing, right?

Swyx15:41

Uh, GPT as well. But, uh, no, no-

Emil Eifrem15:43

Yeah

Swyx15:43

... like every chart in both databases. So obviously I see a lot of these sort of industry-wide stats and numbers. I'm an angel investor in like 30 companies, um, and I, you know, obviously through, through cognition and, and AI, I, I see a bunch of stuff.

Every compute chart is like this, every database chart is like this, and obviously every coding agent chart is like this. Only like, yeah, something's going on.

Emil Eifrem16:04

Something is going on, right? And to us, like I'm not sure if it's like purely the model quality, but yeah, we're seeing it too. And, and a big shift is what I just mentioned, right? Which is, you know, it used to be just draft me the messages.

Now it is kind of send me the messages, right? And so that complete automation in a customer-facing way, it's now gotten to some kind of tipping point around trust, right? Where they, where they're doing that. Another example of this is if-- So that's kind of high level kind of what the companies are doing.

But if you go down to the micro level and like an single developer, a single application, which was kind of the scope of my GraphRAG talk two years ago, right? Like that was kind of the, um, the, the boundary.

And if you think about that, right, like how do people actually write agentic kind of graph applications r-right, right now? Okay, so, so the graph is a tool, kind of vector database might be a tool, you know, that kind of thing, right?

And then you have some kind of English coming in, some natural language coming in maybe, right? And then it used to be that what the recommendation was, take the most, the kind of the hotspot of the type of questions you might get.

Again, if it's a, like the customer support portal is just this most obvious example, right? But we can choose plenty others, right? So if it's Swix, then okay, Swix goes to Apple and he asks some questions, take kind of the hotspot of those questions, express that as a function or a tool, right, in Cypher, and then use text, generic text to Cyph-Cypher as the default kind of backstop if that doesn't work.

And then what customers ended up doing was like, "Okay, I'm gonna do this." And then they sat down and they looked at the log of everything that kind of had a fallback to, to text to Cypher. The stuff that didn't work, they would extract into a separate function or a tool call.

So that was kind of the, the general way that people were doing things, right? And you saw similar with other databases and, and agentic applications about a year ago. And then in the last three to six months, what has changed is that it used to be lead with the specialized functions, fall back to the generic, and now that has flipped So now it's like, okay, just start with generic text to Cypher, right?

And then when that fails, you extract the edge cases into the specialized functions, right? I think it's just across the stack, there's a number of things that have happened like that, that I think caused it to, to flip, call it, you know, three to six months ago.

Swyx18:39

Because it, like, can single shot most things now as well.

Emil Eifrem18:42

Exactly.

Cypher & LLMs18:42

Swyx18:42

Yeah. Uh, y- you know, this, this triggers a, a, a common talk track that I have as well around LLM coding. You know how, um, one of the reasons... I did the workshop with Cypher, and then I didn't, I didn't really use it other than that, b- mostly because it's a DSL, and I'm like, "Uh, do I wanna learn a DSL?"

But now DSLs, you know, a good reason is they are optimized for their use case, and they are so concise, they can handle so many things correctly. But then the, the downside is you have to learn them. Now you don't have to learn them.

You can just...

Emil Eifrem19:13

Yes.

Swyx19:13

And, and yours happens to be a DSL that has a ton of training data. So it's like you, you, like, made the cutoff where you've been around long enough, you survived through ups and downs, and now, you know, people actually can freely use you, um, and obviously optimize if they, if they need to.

But other than, other than that, you, you know, you can mostly one-shot these things, which is super nice.

Emil Eifrem19:33

I, I agree. A couple of more thoughts. One is we had the benefit of becoming a complete actual ISO blessed standard, right? So it started out-

Swyx19:42

Yeah. Aura

Emil Eifrem19:42

... as openCypher actually back in 2015, which after so many back and forths and all kinds of, you know, standards wars and whatever, right? Or not even wars, but standards bureaucracy ended up becoming GQL, right? Which is the first sibling language to, to SQL.

And I actually think that helps as a signal for the kind of LLM training data, right? So that's one thing. Now, having said that, internally... So we have a bunch of in- kind of internal text to Cypher things that we use across the product portfolio, like a old school kind of copilot-type thing.

If you go into Neo4j, into our browser where you can type your Cypher queries, right? Then of course you can use a copilot to translate that, and you can run agents on top of us, and then... So we, we need that as a primitive in our, in our platform.

We actually still fine-tune then that one, and so whenever we, we default to Gemini, and we still do, like, some fine-tuning and even some post-processing.

Swyx20:39

Uh, when you say fine-tune, are you s- are you saying fine-tune a custom model f- to produce, you know, your, your text to, to Cypher, or you're fine-tuning the playground and the output?

Emil Eifrem20:48

No, no, no.

Swyx20:48

Like the prompts.

Emil Eifrem20:49

We're, we're fine-tuning an actual model.

Swyx20:51

Okay. It wasn't clear.

Emil Eifrem20:52

Yeah.

Swyx20:52

Because you can-

Emil Eifrem20:53

Yeah, yeah

Swyx20:53

... I don't think you can fine-tune a Gemini model, right? You can't fine-tune an, an Anthropic model.

Emil Eifrem20:57

Yeah. So this is where, like, we probably use some of the, one of the derived ones internally. I actually don't know, kn- kn- know which one. But we then even have a post-processing step where we do some, like, real just imperative code, like regex-type stuff sometimes to, like, switch some arrows around.

You know how, like, in Cypher you can describe it in relationships in different directions and, and stuff like that, right? And so it actually is a little bit messier than, you know, it's not that the models out of the box are good enough for all situations right now.

I certainly hope, like, what are some bitter pill kind of thing, right? That over time it's gonna be good for even 99% of all situations, but it's not quite there yet.

Swyx21:43

I mean, I, I, I think, you know, that's where, that's where, like, experts will come in, having a good ecosystem will come in and, uh, you know, that, that will still be the domain of human experts, at least for now.

Um, and I wanna come back to, like, um, use cases and cool things that people are doing, right? I do think that there's a, just a ton of work obviously in, like, traditional recommendation systems and fraud detection and all that.

I did start a LLM Rex systems track last year where, um, it looks like LLMs are eating Rex Systems, and I th- I think there's, there's interesting usage of, of graph databases there as well. Uh, apparently all of YouTube, um, you know, the, the YouTube Rex Systems is LLM based, and they-

Emil Eifrem22:20

Is it really?

Swyx22:20

... they, they, they obviously... Yeah.

Emil Eifrem22:21

That's cool.

Swyx22:22

Um, it-- They, they re- they tokenize every video, every video, and put it in a code book, and then they train a LLM on it to r- and then feed in your, your context just like a regular LLM and ask it to predict the tokens of the next videos you should watch.

It's crazy.

Emil Eifrem22:38

That, that is pretty cool actually. I had no idea.

Swyx22:40

The Rex Systems people I know, uh, Eugene, who's, uh, Eugene Yan, who's a, you know, big figure in the Rex Systems world and, and all that, is very bullish about this. Apparently it's, it, it'll-- that's also powering the new X algorithm, that's also powering Pinterest.

It's, like, all the rage in, in their world, and it's kinda cool. You know, it definitely feels uncomfortable. I definitely have some weird recommendations in my algorithm that, like, normal systems would never recommend, but whatever, right? Like, it, it's a new world and everyone's trying to get signal.

And I'm sure if, if anyone A/B tests your stuff, it's YouTube. So, so yeah, ba- basically, like, uh, I guess what are the new workloads or use cases that you're seeing that you want to encourage people to check out?

Emil Eifrem23:18

Yeah. So there's plenty of experimentation going on, which I love, right? Like, again, you know, for the last 10 years we've been very focused on the Global 2000, right? I mentioned several examples there, life sciences, financial services, and I can keep going on and on and on about that.

Context Graphs23:18

Emil Eifrem23:33

But we've also a little bit gone back to our roots in the last kind of, call it year and two-- or two, with, uh, like we have a startup program right now and our, and our cloud service, which is called Aura, you know, has a form factor and a price point that works for startups, right?

And so that's, like, really, really phenomenal. And of course, a very popular one here is agentic memory, and so there, there's a lot of people who wanna do kind of memory on graphs. Um, people don't talk about this, but, like, actually the initial MCP release includes like a tiny little in-memory-

Swyx24:06

Graph database

Emil Eifrem24:07

... this is a hard thing to say

Swyx24:07

200 lines of code

Emil Eifrem24:09

Uh, yeah, I thought it was 300 lines of Python. Yeah, yeah, s-some- s-- yeah, yes, something like that, right? And so it's a toy, it's tiny, right? And it's an in-memory, memory implementation, right? But it's graph-shaped, right? Which is, which is cool, right?

Of course, I love that. And so there's, there's just a lot of people who, who kind of naturally do that. And then, of course, there was all this discussion around context graphs, uh, that happened over the last, you know, kind of three, three months.

And so we see plenty of people doing things like that, which is also, uh, I think a great use case for graphs.

Swyx24:43

Yeah. Okay. So, uh, memory is a whole topic. I don't know if we'll have the, the time to cover that. My, my quick cent- two cents on it is, like, in theory, graph databases are fantastic for memory. In practice, it might be overkill.

Most people's memory doesn't go that far, and I don't think, like, we have figured out the struct- the true structure of memory that, that, like, really, really works over a super long period of time. And probably it's like, yeah, it probably fits in a, a single file.

Like, like m- I don't put, I don't put out that many tokens, you know. But context graphs, yes, right? Like, so, you know, this is where, like, what I always say is, like, I don't care how long your LLM context is.

We, we took, you know, three years to go from 100K context to one billion context in every frontier model, but we're not going to a billion. You're not going to a trillion with context graphs, uh, context lengths. What is your take on the context graphs discussion?

I mean, it's really blown up. We had a, we had a short pod with the authors about it. I, I think they're kind of leaving it open to interpretation. I think a lot of people will build context graph systems, and we'll figure out what it is, but what are your early takes?

Emil Eifrem25:45

Yeah. So my view on it is that in my head, it completed the quadrant of the types of data sources that are required to reach what I've been talking about as kind of escape velocity for agents in production.

And so what I mean by that is I think it's exactly four types of data sources, and I would love to hear your take on it, right? If you can identify a fifth or a sixth one. But I think there's exactly four data sources that are required, and I don't think it's as easy as you have to have all four of them, but, like, the more the merrier on some level.

And the first data source for agents, I think, is operational data stores, right? And so this is the, like, the system of records, and I think of them as systems of record for the present, right? So this is like, okay, how many customers do I have right now?

Or what's the value of this particular customer right now? You know, that kind of thing, right? And, and I think in the conversation, we flip between calling the application on top of the r- like, operational database. We call that the system of record, so Salesforce is the system of record.

That's one way we talk about it. The other way is, like, the actual database is the system of record, right? But no matter what, I think that's kind of the operational database is in the first-- is, is the kind of the first quadrant.

The second quadrant to me is the cloud data warehouses, right? So this to me is people say that they're not systems of record. I actually think they are. They're a system of record of the past, right? So if operational is, is system of record of the present, then this is the system of record of the past, right?

And so how much revenue did we get from the LATAM region in Q3? You know, that kind of thing, right? And then, like, the, the third quadrant to me is agentic memory.

Swyx27:21

So that's OLAP. Yeah.

Emil Eifrem27:22

That's OLAP, right? So OLTP and OLAP for like... Yeah, yeah, yeah. You, you said DSL before, which is probably a data term to your domain-specific language, right? Like, but OLTP and OLAP, right?

Swyx27:31

No, I think it's like a programming, yeah, programming term, yeah.

Emil Eifrem27:34

Yeah, yeah, yeah. Um, and then the third one is agentic memory, right? And so, or maybe, like, agentic state, like system of record for agentic state, maybe short-term, long-term state, something like that. And then the fourth one is the context graph, right?

And so what is it then? So this is the why behind we made, like, uh, the particular values in one, quadrant one and two, right? And okay, so, you know, the classic example is the, you know, I, I sold into this customer at this price point.

That's a 20% discount off of list price, and our policies say that we're only allowed to discount by 10%, right? Okay, so the why is, well, I wanted to break into this vertical or into this geo, and I got an approval from my sales VP that was on the phone, over Slack, by email.

It's not recorded anywhere. And those decision traces is what they call it, right? End up constituting a graph that they call the context graph. And if generally we're in the kind of space right now of trying to shift decision-making from, you know, kind of human brains into agentic brains, right?

It used to be from wetware to software. I don't even know what to call kind of the LLMs now, right? Into latent space maybe, right? Um, then, you know, then having access to that institutional knowledge of how things actually happen, the actual decisions in big organizations feels really valuable.

And so those are the four quadrants, and then I can talk about kind of more context graph-specific stuff. But first of all, do you agree with the four? Do you see a fifth or a sixth?

Swyx29:11

I feel like the, the, the three out of the four is obvious. Sorry, I strongly agree with. And then the, the fourth one is agentic memory, actually, where I'm like, that's a little weaker. That's a little less established.

That's a little smaller. It's like the, the one, two, and four are very strong, very large categories where I know exactly how to architect it and everything. The third one is like the, uh, I don't know, something, something memory.

And I feel like there could be more... A- when I would think about two by two, right? OLAP, OLTP is great. That's like a dimension of how wide you're querying, how, what volume of transactions you're doing. The other dimension should be ideally-- uh, the other axis should be orth- orthogonal, and I don't necessarily know what that axis is.

Uh, and so it doesn't exactly fit a normal two by two, which usually means that maybe there's a third dimension that's kind of being sort of mixed into the, into the, into the equation here. Because agentic memory, let's call it, is probably mostly personal, maybe some organizational Whereas context graphs is fully organizational.

Um, uh, and yeah, so, so that's my, that's my initial reactions, but I'm happy to just only talk about context graphs unless, unless you have more sort of-

Emil Eifrem30:20

No, but I, I, I... It was really interesting to hear that. I think, like, agentic memory is some kind of, "I wanna understand what happened in the past for that individual," I think is gonna be an important source for retrieval for many use cases, but maybe not all.

And-

Swyx30:35

Yeah. Yeah, so, so, so maybe, okay, maybe one, one sort of... I'll put this in my own words, is agentic memory is, is sort of the stuff that you've done with the agent, and then the context graph is the stuff that you've done with everyone else.

Because I think actually there's a lot of work by all the agent companies on querying their own past trajectories, and memorizing knowledge, and self-improvement from, from there that really has nothing to do with the context graph. It is just the agent should get better as I use it, right?

Uh, it's, uh, yeah, of course.

Emil Eifrem31:01

I think that's right.

Swyx31:02

Yeah, yeah.

Emil Eifrem31:03

Now, on context graphs specifically, right? So I think it makes a lot of sense to somehow encode that institutional knowledge, right, in some digital form, right? And it's back to the PhD-level intern that you get that wake up every day, and they don't know the past, you know, that kind of thing, right?

And so, so that, that makes sense. I guess the trick is, like, how do you actually instrument your organization to get there? I think we can all see this kind of future state where, okay, I have my agents in production, like my example from before.

Like, I have some kind of an agentic process that reaches out to my customer and offer them mortgage discount or, or something like that, right? And if they do that, they should be instrumented and record all of their decision-making throughout that.

Use that as a source for improving themselves going forward, and you can kind of sense the compounding effect or the flywheel, right? Okay, that's great once we're there, but we're not there now. We don't have that in any digital form.

Most typically, it's that sales rep calling someone up from a car, asking for a discount, getting a yes, then selling the deal, and at best, it's recorded in Salesforce, right? And that's kind of how it happens today, right?

Like, so how then do we bootstrap the context graph? So that has been almost all the conversations that I've had with big enterprise customers. With startups, the bootstrapping of the context graph is by virtue of getting adoption of their product, right?

And so that's then not the problem then. The pr- the problem there is, like, how do I even get adoption of my product, right? But inside the kind of the enterprise context, no pun intended, that ends up being like, how do I even start this instrumentation so I get enough decision trails that I can bootstrap the process?

That ends up being the, the interesting discussion.

Swyx32:58

So how valuable are startups to you? Like, you, you deal with, like, 80 of the Fortune 100. You, you, you target to the 2000s. Startups, you know, free tier, whatever. Like, they're not that worth they're not worth that much.

Startups vs Enterprise32:58

Emil Eifrem33:09

Yeah. I mean, maybe I'll take the, the kind of founder, CEO kind of-

Swyx33:13

Yeah

Emil Eifrem33:13

... point of view on this, right? Like, like a decade ago, right, we looked at our revenue, and it was like a third, a third, a third, and, like, a billion and above was a third, and then there's like some kind of a mid-market kind of thing, right?

And then a third was startups, right? And we just saw all the underlying metrics of the, of the billion and above segment was just so much better. And then we looked at all database companies at scale, also known as Oracle, but even DB2 as well and stuff like that, and 80-plus percent of their revenue was from the Global 2000.

So clearly that is where database companies monetize, right? And so we kind of added all these things together and we said, "Okay, 100% what we need to do is lean in on the enterprise." And that was the right business decision.

My CEO hat, love it. It was super successful. That's why I can sit here today, all that good stuff, right? My founder hat was a little bit sad, right, because startup people are my people, right? And, and they represent the future and all of that kind of stuff, right?

And so now that we have this new generation of AI-native startups happening, right, we're like, "Oh, now this is really important for us to be built into those architectures," right? And not necessarily maybe from generating tons and tons of ARR.

I will always make more money from Bank of America as a kind of generic example, not saying anything specific of Bank of America. But except we do have 20 of the 20 biggest banks in North America are customers, so you can derive from that what you want, right?

So we have all of them, right? But I'm always gonna make more money from the enterprise segment. But I think it's really important to just be part of the zeitgeist. And from that perspective, like, getting the next generation of startups built on Neo4j is, is really crucial and important.

Swyx35:01

I-- presumably startups are easier to onboard because they have less context than the, the, the larger companies. So I guess, what have you prescribed so far in terms of like, you know, whichever segment you want. Uh, honestly, both are interesting to me.

Getting Started35:01

Swyx35:15

What are the best practices so far? Like, you know, how, how do they even get started?

Emil Eifrem35:19

It depends on which altitude you operate at, right? Like, if it's a, if it's a single startup, it's not just there's one product, but the product is the company. So that's then everything, right? And that's very different if you look in- inside of the enterprise context, right?

Because then we engage in, in two levels, right? One is the scope of a single application, a single project, right? And the other one is more enterprise-wide. This-- Let me talk about the, the, the last one for a moment, 'cause this is a big change that has happened since, certainly since the two years ago, the, the, the GraphRAG talk.

A clear pattern that we're seeing happening is we, we tend to call it a knowledge layer, but we bump into a lot of enterprises. And the problem is, okay, every single data source is gonna have some kind of an MCP endpoint here, right?

Okay, I wanna give my agents access to that, so, like, what can I do? Either I give my agents access to all of the MCP endpoints and let it figure it out, right? The problem that they run into is that, yes, that works, but, like, some of the data's gonna be conflicting, right?

And so we just talked about we have a cloud service, right? Even for me, like, when I try to figure out, like, how many customers do I have on my cloud service, like, I go to the actual cloud platform and I ask, and we have, let's call it 3,000 customers, right?

That means how many are active, have an account with a running database. And then I go to, like, the finance system and I have 2,800 customers because those are the ones that were, like, with a credit card and the...

You know. Okay, so, like, so how do you make sense of that, right? And I think we're learning that an LLM is gonna give you an answer wi-without doubt, right? You just have no idea if it's right or wrong, right?

And it's very hard to find that out, right? And so what we're talking to a lot of big enterprises about now, right, is you wanna own that layer or they wanna own that layer, right, where they consolidate the metadata around where data sits, right, inside of their enterprise, and that gives you the consistency and the trust and the explainability, right?

And so what is that metadata that they wanna own then? It's basically the data asset landscape, right? So what are the schemas that sit in your relational database, right? What are the buckets in your S3 and so on and so forth, right?

Express that in a graph form, marry it up with a business-facing ontology. What is a customer, right? Like, how does a customer relate to a supplier? And what, what are the concepts that I have?

Swyx37:50

That's so hard. You talk to five different people, you get six answers.

Emil Eifrem37:54

That's exactly right, and this has been the key problem that has, has been holding this back in the past. What is happening now, of course, is I need to solve this to make my agent successful for this particular use case.

So that forces them to kind of converge on something that is good enough to work for those agents, right? And then the third piece is kind of the mapping between, between... People talk about semantic layers. People talk about context layers.

I like the term knowledge layer. They all mean kind of the similar but distinct things. But this is a very, very popular use case that we have right now. If you, you were on our website, if you Google kind of leading media company or something like this, um, and Neo4j, you're, you're gonna see this, and it has, like, a nice little architecture in green that, that you can scroll down, and if you can screen share that, you'll, you'll see it.

And so you have, like, the disparate data silos at the, at the bottom, and they, they call it the semantic layer, right? And then they have their agents running on top of that. There's a couple of flavors where you do kind of zero copy.

So what the, what the knowledge layer will give you is the map of where I should actually go and query, or sometimes we inline the data, like partial data, into the knowledge layer so you can just do the queries right in there.

Swyx39:18

Yeah. What, what, what is zero copy again?

Emil Eifrem39:20

I wanna say Salesforce has coined or at least popularized this term. This is where it's like, okay, instead of replicating your data in every single, like, between Snowflake and this and that, right? Instead, you have, like, a virtualization-type layer where it points to the original data source.

So in-

Swyx39:42

Okay, foreign data wrapper.

Emil Eifrem39:43

Yeah, exactly. Exactly. Right.

Swyx39:47

You can tell I'm, like, very Postgres-centric.

Emil Eifrem39:49

Yeah, yeah, yeah. That's right. And so the query path, you end up going to this layer, and then you figure out where the data sits, and sometimes it's materialized into that layer, and so you can just query that directly.

And so this is, this is a very popular use case kind of enterprise-wide. But I think your original question was like, okay, what are the best practices for, for getting started? We talked about context graphs. Uh, there's a, a, a really neat Python wrapper called UX, like, whatever, Create Context Graph, right, that gives you a context graph out of the box for, I think, 22 different industries.

Uh, it was just published, like, a few, a few days ago, and it stands up a full, full Neo4j with a, with a front end, and you can-

Swyx40:36

This is cool.

Emil Eifrem40:38

It is modeled-

Swyx40:39

We haven't did it

Emil Eifrem40:39

... and you're gonna love this, Wicks. It's modeled after, what was it called? Create React App.

Swyx40:44

Of course.

Emil Eifrem40:44

Right?

Swyx40:45

Of course.

Emil Eifrem40:45

And so you can have, like, this interactive, it'll help you with your own data. But you can also choose, you know, out of, um, existing domains. And then, um-

Swyx40:55

22 domains. Wow.

Emil Eifrem40:57

Yeah, exactly.

Swyx40:57

Okay.

Emil Eifrem40:58

Integrates with eight or nine different agent platforms.

Swyx41:02

Only wanted one, and it's the one that's not here.

Emil Eifrem41:05

Oh, no. Which one?

Swyx41:07

Social media.

Emil Eifrem41:08

Interesting. Yeah. Yeah, we can add that.

Swyx41:11

Yeah, yeah, yeah. Social, social graph and social media. Basically, I have a view that every SaaS basically eventually becomes, has a, has a social graph inside of the SaaS with the auth system. You know what I mean? You want people with teams.

You want people who can message each other. You want to send notifications. You want a feed. Elements of social media find their way into every single one of these domains. It's, like, not even a, a, a vertical domain.

It's just a, a, a feature set that I always c- you know. Um, I actually looked into doing this as a, as a business before, and I, I found that other people have done it and it, and it didn't work as a business.

But basically, you know, drop in social media, uh, social, I guess, network into my user base as a feature.

Emil Eifrem42:00

Into your end user base or into your employee base?

Swyx42:04

No, uh, end user.

Emil Eifrem42:05

Yeah, yeah, yeah. Yep.

Swyx42:06

So you know what I mean? Like, like, because, because every one of my users, they work in a team, they follow, they have, they want a feed, they want notifications, they want DMs, they want even blocking, uh, whatever, right?

They want a homepage. It's all social features that none of these people really would, would build themselves, but Probably they want something. Maybe, maybe here, I don't know. Um, but yeah, anyway, this is really cool. Uh, co-congrats to William.

Uh, I, I know he's, uh, he's done a workshop at one of the AIEs before.

Emil Eifrem42:35

Yeah, he's, he's phenomenal. And, um, so what, what it does is it just helps you get started, right? And it creates some s- either synthetic data. It also integrates with, like, a bunch of SaaS tools like, uh... I saw you were excited about the, what is it called?

The Google Workspace CLI or something like this, right? So it can get data-

Swyx42:55

Everyone is excited about that. I mean-

Emil Eifrem42:57

Yeah, I know

Swyx42:57

... you see the number of likes on that one.

Emil Eifrem42:59

I know. It's shocking, right? And-

Swyx43:01

Do you know how... Okay, so I got so excited about Cloud Cowork because I could use it to navigate the GCP dashboard. I was like-

Emil Eifrem43:09

Yeah.

Swyx43:09

... "I don't wanna do this. This is so bad."

Emil Eifrem43:14

But so, so it can, can get you real data, or it can generate synthetic data for you.

Swyx43:20

No, it's fantastic.

Emil Eifrem43:20

And-

Swyx43:21

Yeah

Emil Eifrem43:21

... it, it, it uses... We have an, like, an agent memory kind of toolkit, which has the short-term memory, so that's kind of the conversational state type stuff. It has long-term memory, so that's, those are all the entities and all the concepts in the domain, and it also has the decision traces, so the context graph.

So it kind of puts it all together with a little graph visualization, kind of UL- UI, and all that kind of stuff, right? So that's a great way to get started.

Swyx43:48

It's, it's a great idea. I can't believe no one else has done it. I can't believe this is only six days old. Yeah, you should, you should, uh, promote this a lot.

Emil Eifrem43:55

And, and, and you can imagine, it's the same story. Like, he, he hacked it up one Sunday afternoon, like this, so I guess five days ago-

Swyx44:01

Yeah, yeah

Emil Eifrem44:01

... six days ago-

Swyx44:02

Yeah, yeah

Emil Eifrem44:02

... Sunday afternoon ahead of a context graph meetup that we, that we did, and people love it.

Swyx44:07

Yeah. Yeah, yeah. Okay. O- I- if, if I have one feedback looking at this, it, it, uh, almost does too much, almost. It's like-

Emil Eifrem44:15

Yeah, yeah

Swyx44:16

... well, um, w- where does it end, and where does my app begin, right? There's, there's a lot here, and so I-

Emil Eifrem44:23

Yes

Swyx44:23

... I'm already feeling a little bit overwhelmed with this README. So some focus is, is, is valuable. You know, it's, it's one of those things where, like, open code, you know, the OpenCode CEO, Dax, was like, "Look, in a age where any feature you can just vibe code it within days, actually restraint and focus and, like, you know, fit for purpose is, is, is a thing."

Because yeah, you can add it, but, like, does it actually add anything, or is it just more stuff that just fills up my context window? And I'm just like, "Okay, what, what is the, the job that you do for me?

And just, just do it really well," you know? So I, I do think, like, this is one of those things-

Emil Eifrem44:54

It's back to the kind of Steve Jobs, "I'm the editor of the product," kind of thing. And l- of course, the most important part of the editor is saying no, like removing things.

Closing Thoughts45:04

Swyx45:04

You know, that, that's a, that's a quick tour. Uh, obviously there's a lot more going on in your world, uh, that we'll, we'll learn about. Yeah, any, any parting thoughts? Outside of your, your official hat, your, your, your hacker at heart, what else are you excited by?

What are you gets you super excited?

Emil Eifrem45:18

Well, I mean, we have... It's like this era that we live in right now is simultaneously the most exciting time and the most shocking time. We haven't talked at all about kind of SaaSpocalypse and all of that kind of stuff.

But, you know, with the other hat as a CEO, I have to pay attention to, to, to things like that. And of course, we were kind of front and center in, like, starting to see the shift of the buy versus build, right?

Because, you know, Klarna was one of the first examples of, you know, shutting down transform-

Swyx45:51

They both tried it and then also rolled it back.

Emil Eifrem45:53

That, that is the public view, right? They certainly, I'm not gonna speak for, for Klarna, but I, I think there is overstatement on both sides of that, of that probably.

Swyx46:03

I agree. Yeah.

Emil Eifrem46:03

But, but clearly there's a shift in that buy versus build, right?

Swyx46:07

Yeah.

Emil Eifrem46:07

But then for me personally-

Swyx46:09

For, for me, I pay 200,000 a year for conference management software that I hate, and like, you know, maybe two, pay, I pay 2,000 to, to make one that I like.

Emil Eifrem46:19

Th- that's shocking that you-

Swyx46:20

But we-

Emil Eifrem46:20

... of all people haven't fixed that.

Swyx46:22

I'm busy running, you know-

Emil Eifrem46:23

Yeah

Swyx46:23

... like printing badges and, like, freaking-

Emil Eifrem46:25

Yeah

Swyx46:25

... figuring out my Magician costs and everything, you know? Like I, this, this is not, like, top of my list. And I have a team that isn't as agent native as me. So I cannot just, like, unilaterally as, like, as, like, a CEO, like, just decide, "Oh," like, "my whole company's gonna do this now," and expect that every single one of my employees will know what to do.

Because, no, they need to be trained, and actually sometimes the, the familiar, buggy, slow, whatever system, with all its flaws, actually still is better.

Emil Eifrem46:52

And, and it's, it's back to, like, what, what is the value of these things? And a big part of it is understanding processes and being prescriptive about those processes, right? And this takes us some while to iterate there, and you could do it if you're focused on it.

But in the meantime, you're gonna pay for the 200, right? So, so, so I get it. But, but so that's kind of on the terrifying side or, like, the shocking side or something like this, right? And then on the other side is just, man, the ability to build.

And remember, like, I've, I'm kind of post technical, right? Like, so I couldn't write the, the compare and swap stuff-

Swyx47:25

Post technical is a good way

Emil Eifrem47:25

... is like, that's the kind of stuff that I know, right? But I haven't been a modern, real pro- programmer for 10 years, right? And now, again, I can, like, software is malleable again, which is just so phenomenally exciting.

Swyx47:38

Yeah, but I think one, one failure mode of the post technical, you know, like, manager, which you are, which I am, is that you think you, you, "Oh, AI can do it." And then you, like, you know, you vibe code something, you throw it to your, your employees, and then you expect them to just pick it up.

No. So, like, I, you know, I don't, I don't wanna encourage the leaders among us to also just be respectful of, uh, the fact that not everything is vibe coded, and your employees still have to clean up your mess sometimes if you do this.

Emil Eifrem48:04

Of the real craft, I completely agree. Especially, like, as a database company, like, man, that's- ... there's some, some real craft in that.

Swyx48:12

Yeah, but you guys test very well, you know what I mean? Like, the database testing is, is-

Emil Eifrem48:16

Yep

Swyx48:16

... top-notch. Most software is not testing this.

Emil Eifrem48:20

Absolutely.

Swyx48:20

You know? So yeah, actually, I... No, be- because you can test so much, then you can vibe code more-

Emil Eifrem48:24

Right

Swyx48:24

... because it's so, so testable. You know what I mean?

Emil Eifrem48:27

Yes.

Swyx48:27

So anyway. Anyway, thank you, Emil. Uh, this was a really fun chat. Um, and, uh, I'm excited to catch up in person. Uh-

Emil Eifrem48:34

Always a pleasure

Swyx48:34

... I think you're coming to, to San Francisco. Um, so we'll see you there.

Emil Eifrem48:38

All right, my friend. Talk soon.

Swyx48:40

All right. Bye.