LALatent SpaceJun 24, 2026· 1:10:06

The Agent Cloud: Databricks’ Bet on the Future of AI — Matei Zaharia and Reynold Xin

Databricks cofounders Matei Zaharia and Reynold Xin argue the company is moving beyond the lakehouse into a full data-and-AI operating system, anchored by two new initiatives: Omnigent, an open-source meta-harness for combining coding and enterprise agents, and LTAP, a unified storage layer that gets most HTAP benefits without collapsing query engines. Omnigent provides a common API for agent sessions, files, streams, tool calls, and cancellation, solving portability, collaboration, and security issues across Claude Code, Codex, Cursor, and custom agents. LTAP writes transactional data directly in columnar Parquet format, eliminating brittle CDC pipelines—Reynold jokes CDC means 'continuous data corruption'—and enables instant analytics without overloading the source database. The episode details Databricks’ culture of rapid prototyping, where an engineer built the LTAP prototype without a formal design doc, and the thesis that traditional software will be rewritten once data is in the right place with agents on top. They also cover Mosaic’s shift from general frontier models to specialized fine-tuned models like document parsing, internal agent usage, and security features such…

  1. 0:00Intro
  2. 3:27Omnigent
  3. 19:17Agent Security
  4. 28:56LTAP
  5. 53:58Snowflake
  6. 59:03Mosaic Models
  7. 1:04:25Data Thesis
  8. 1:08:23Closing

Powered by PodHood

Transcript

Intro0:00

Reynold Xin0:00

One of the thesis we have is actually once you can get the data in the right place, the AI models are becoming pretty good. The generic agents are fairly... I mean, Ali talked-

Swyx0:10

Yeah

Reynold Xin0:10

... about AGI is already here. They have pretty good reasoning capabilities. Actually, I think many of the traditional software will be sort of, uh, rewritten, uh, with this new paradigm, which is just get the data to be there, and then just slap some agent on top.

Magic will come out.

Swyx0:24

Yeah.

Reynold Xin0:25

Um, but without the right data, you can't really do that.

Swyx0:29

Before we get into today's episode, I just have a small message for listeners. Thank you. We would not be able to bring you the AI engineering, science, and entertainment content that you so clearly want if you didn't choose to also click in and tune into our content.

We've been approached by sponsors on an almost daily basis, but fortunately, enough of you actually subscribe to us to keep all this sustainable without ads, and we wanna keep it that way. But I just have one favor to ask all of you.

The single most powerful, completely free thing you can do is to click that subscribe button. It's the only thing I'll ever ask of you, and it means absolutely everything to me and my team that works so hard to bring The In-Space to you each and every week.

If you do it, I promise you, we'll never stop working to make the show even better. Now let's get into it.

Matei and Reynold from Databricks, uh, welcome to The In-Space.

Reynold Xin1:20

Hey, thanks for having us.

Swyx1:21

Yeah.

Matei Zaharia1:22

Yeah, thanks so much.

Swyx1:23

Uh, thanks for taking time out. You- you have your Databricks, uh, Data + AI Summit going on. You were just telling me how the first summit that you guys ran was just 50 people.

Matei Zaharia1:30

Mm-hmm.

Reynold Xin1:31

Yeah, it was a-

Swyx1:31

In Berkeley

Reynold Xin1:32

... little meetup at Berkeley, I think-

Matei Zaharia1:33

Yeah

Reynold Xin1:33

... put together and-

Matei Zaharia1:34

We were doing these tutorials and, yeah, just teach people Spark.

Swyx1:37

Yeah. You know, obviously now it's like, I think like, uh, the, the headline number's like 100,000 people around the world, 30,000 in person.

Matei Zaharia1:44

Mm-hmm.

Swyx1:44

Uh, it's a crazy-

Matei Zaharia1:45

Amazing

Swyx1:45

... community. Well, I mean, I just saw the keynote. Ali's just... Did you know that... Was it obvious when the, uh, back when that Ali would be, like, such a great, like, CEO? Like-

Reynold Xin1:56

Oh

Swyx1:56

... such a great presenter?

Matei Zaharia1:57

What do you think? Uh, I mean, I think among our group of founders, it was clear that, uh, I think he'd be the best at this.

Swyx2:04

Yeah.

Matei Zaharia2:04

And, uh, and yeah, it turned out great. And he's, uh, I mean, he's ramped up on so many topics growing our company. He would just go in and, like, study it and, you know, be g- talk to all the experts, like, even if he can't hire the person, you know, learn enough about, like, finance and sales and whatever it was, um, and, uh, you know, and go from there.

Yeah.

Swyx2:23

Yeah.

Reynold Xin2:24

I mean, he's obviously very high IQ and a very high EQ, but it wasn't... Like Ali today is quite different from Ali from, like- ... 10 years ago. I think he- there's a lot of work that he put in to, uh, get to this point.

Swyx2:35

Yeah. I mean, no, I mean, to me the, the most appealing thing about him is that he's funny. And like, it, you know, it's, it's-

Matei Zaharia2:40

It's true, yeah

Swyx2:41

... uh, it's hard to make jokes about, you know, data warehouses-

Reynold Xin2:44

About serious topics

Swyx2:45

... and security and-

Matei Zaharia2:46

Mm-hmm. Yeah

Swyx2:47

... what have you.

Matei Zaharia2:47

Oh, yeah. That's for sure.

Swyx2:48

Yeah. So you, uh, you guys launched a whole bunch of things. Uh, I- I'll just do a name check briefly, uh, the stuff because we- we're not gonna cover everything. Omnigent, uh, your, your baby. LTAP, your baby, your dream engine.

Uh, we're also gonna cover Genie, cover Customer Lake, uh, you acquired Panther-

Matei Zaharia3:06

Yeah

Swyx3:06

... Open Sharing, and there's Unity AI Gateway. A lot of these, I think, like, are things that you would expect a Databricks to do. It's, it's like part of the-

Matei Zaharia3:14

Mm-hmm

Swyx3:14

... the roadmap. Everyone in your category has, has similar things. But I think, uh, probably the two of you are leading the two most unique and differentiated initiatives-

Matei Zaharia3:23

Mm-hmm

Swyx3:23

... uh, on, uh, in the landscape. Maybe we'll start with, uh, with, uh, Omnigent, and-

Matei Zaharia3:27

Mm-hmm

Swyx3:27

... we'll, we'll, we'll, we'll, we'll go into it. I do think that a lot of people are exploring this sort of meta harness concept.

Omnigent3:27

Matei Zaharia3:35

Yeah, totally.

Swyx3:35

What led you to it?

Matei Zaharia3:36

Yeah. There were actually a couple of, like, converging lines, which is, I think is a good sign that you need something new. So on the one hand, there's all the coding agent info internally. We have really great, uh, dev info team.

Uh, they built something called Isaac, that's basically like a wrapper on Claude Code and, and Codex, and, uh, lets you use them either on the web in, like, sandboxes or, uh, just on your dev machine or, or on your laptop or whatever.

And then, you know, they were adding all kinds of stuff there. And, and we saw all the, the m- sort of more advanced engineers like, uh, were building their own workflows with tons of agents, and they were building their own UIs and stuff on top of, even on top of that.

And then the other one was, like, us building agents. We ship this, like, data science agent called Genie on the research team, which I, I, I co-lead basically. We also build a lot of internal ones for various things, and then we have all the customer ones.

And all of them were running into this thing of like, "Oh, I need to switch model and, and harness and so on, uh, you know, every few months." Plus, the agent is, like, completely useless if you can't share sessions with someone and have history and have search, and all this, like, layer on top of it for collaboration.

I thought a bit about it from both contexts and, uh, at first people thought it was weird. They're like, "Why are you doing coding agents and custom agents in the same thing?" But I said it's, it's, it's basically the same problems and, um, you, uh, you, you just wanna build the stuff that lets you deliver the agent, maybe control it if you care about security and, um, make it portable across things.

And then we prototyped some things as experiments. We saw, yeah, actually, we can make it work, and then we, you know, we sort of built that for real.

Swyx5:21

I'm wondering if this kind of, uh, m- let's call it architecture-

Matei Zaharia5:25

Yeah

Swyx5:25

... maps to anything in your careers in the past. You know, like I always-

Matei Zaharia5:28

Mm

Swyx5:28

... think about how a lot of things actually just tie back to operating systems.

Matei Zaharia5:31

Mm-hmm.

Swyx5:33

A lo- a lot of operating-

Matei Zaharia5:34

Yeah

Swyx5:34

... systems tie back to databases, so-

Matei Zaharia5:35

So I-

Swyx5:35

... or the other way around

Matei Zaharia5:36

... so, so the thing I do think it ties a lot to, like, network protocols, you know, internet protocol. Uh, we, we also did-

Swyx5:43

Communication between entities.

Matei Zaharia5:44

Yeah. We did stuff with, like, data sharing also, which is probably, you know, most viewers probably won't know unless they're-

Swyx5:50

Yeah, open protocol is the term.

Matei Zaharia5:52

Yeah.

Swyx5:52

Open sharing. Open sharing.

Matei Zaharia5:53

Open sharing.

Swyx5:53

Yes.

Matei Zaharia5:54

Yeah. So it's like you have a company, you maintain some kind of table, like, like let's say like a Walmart or something. They have like the, you know, inventory and what's been sold in each store. And then you also have suppliers, and they would love to produce more things and ship them, like, exactly the moment you need them.

So they would love, like, real-time access to your table. So instead of like sending emails around or Excel sheets or phone calls, why can't you share like a view of that table in real time with them, then they query- They, you know, join it with their data, and they decide what to send.

So i- it's, it's one of these things where you, you, like... Like you, you, you might ask, like today, since we can vibe code anything so fast, why do we even need to design like protocols or APIs or software, right?

Can't you just vibe code things on demand? But actually, for this type of interoperability where multiple parties that are moving at different speeds are building stuff and you still want some layer on top to coordinate, um, you do wanna design it and build it.

So it, it reminds me of that, like agents talking to each other and, uh, users talking to agents and uh, and tools.

Swyx6:56

Reynold, any other comments or-

Reynold Xin6:58

Um-

Swyx6:59

... alternative viewpoints?

Reynold Xin7:00

I think, by the way, we had a debate on exactly which set of benefits would, uh, matter a lot, and I, I think around the time we decided to do this thing-

Matei Zaharia7:08

Mm-hmm

Reynold Xin7:08

... um, I was telling Matei, hey, it just happened to be there's a particular week that I was coding nonstop, uh, from the moment I woke up to like the moment I went to bed. I was like looking at my Claude sessions, my Codex sessions, and uh, and one of the things particularly annoying was having to keep my laptop open.

Matei Zaharia7:25

Mm-hmm. Yes.

Reynold Xin7:27

Um, I was actually driving to a doctor's appointment, and I remember because I wanted to make sure the whole thing continues working-

Swyx7:32

But by the way, it's so comforting to hear you say that because I'm like, I don't know if I'm a clown and I'm doing this, or like

Reynold Xin7:39

Yeah, yeah. Like, honestly, I, I was driving, and I was tethering my laptop to my phone.

Matei Zaharia7:43

Uh-huh.

Reynold Xin7:43

Keeping it on the side. Whenever I hit a red light, I started looking at what's going on on my laptop.

Matei Zaharia7:49

Yeah.

Reynold Xin7:49

And I just felt that was ridiculous.

Swyx7:51

Yeah.

Reynold Xin7:51

It felt like we went back to the dark ages of-

Swyx7:54

Yeah, yeah

Reynold Xin7:54

... programming. I mean, the productivity you gain from all this coding age is amazing, but, um, yeah.

Matei Zaharia7:59

Have you heard of Claude?

Reynold Xin8:01

Yeah. It, it was crazy to me.

Matei Zaharia8:03

Oh, the, the thing we were working on was the sandboxes, or was this before that?

Reynold Xin8:06

It was a sandbox.

Matei Zaharia8:07

Okay.

Reynold Xin8:08

I was work-

Matei Zaharia8:08

So you, you were in

Reynold Xin8:09

So I was approaching from a very different angle. I wanted to, "Hey, we're gonna have cloud sandboxes that actually doesn't shut down. You can get one very quickly," but not just for running agentic sessions.

Matei Zaharia8:20

Yeah. Right.

Reynold Xin8:20

It's actually also for running development. So I was actually personally building that that week, and through building that, I ran into all these issues, and then I wrote-

Matei Zaharia8:29

Yeah

Reynold Xin8:29

... actually a document for Matei. It's like, "Here's my wish list of what the actual environment should do." And I think he actually ended up almost implementing-

Matei Zaharia8:36

Yeah

Reynold Xin8:37

... every single one of them.

Matei Zaharia8:37

Yeah, I remember Reynold saying, 'cause my first prototype of this had just chats with your agent, and he said, "I have to be able to open a t- a, a shell, like my own shell, and like list files and like tail them and stuff."

So actually-

Reynold Xin8:50

Basically SSH into a mainframe.

Matei Zaharia8:52

Yeah, actually it has that now.

Reynold Xin8:53

Tailing my log.

Matei Zaharia8:54

Yeah.

Reynold Xin8:54

Um-

Matei Zaharia8:55

Yeah

Reynold Xin8:55

... and also another thing I think I asked was, uh, I, I, I had... I still use Cursor for the sole purpose of rendering markdown files.

Matei Zaharia9:02

Uh-huh. Yes, yes.

Reynold Xin9:03

So I said, "Just give me a way to see my markdown files and render them-

Matei Zaharia9:07

Yeah

Reynold Xin9:07

... properly. I don't need a separate tool anymore."

Swyx9:10

Yeah.

Reynold Xin9:10

And I think you also built that in.

Matei Zaharia9:11

Yeah, we, yeah, we did that, yeah. Yeah, we had a lot of engineers building, you know, their own vibe coding setup. But then the other thing they all said is like, "Hey, I built something that's amazing for me, but like no one else on the team can use it, 'cause I, I don't have a server to collaborate."

And, and this is, this is why we tried to set up O- Omnigent, so you can have a server and have the security, uh, set up in there. So, you know, like log in with Google or whatever and like actually securely share stuff.

Um, would... And that's where we've seen a lot of other agents like hit things. Like people think they prototyped an awesome agent, but you know, it's not allowed to connect to like some really important data or whatever because of the security team.

Swyx9:52

Yeah.

Matei Zaharia9:52

So, yeah.

Swyx9:53

Yeah, at this point, uh, so for those watching along on YouTube, we're gonna put- putting up a image of the structure here, uh, and we can talk through a little bit of the architecture. I think I just, I just wanna have p- people understand, 'cause like when we're talking about software, it can be very abstract.

Matei Zaharia10:06

Mm-hmm.

Swyx10:06

And like, uh, like here is actually what we're talking about. You've worked out in open source this entire platform, basically, and there's a runner component, a server component-

Matei Zaharia10:15

Mm-hmm

Swyx10:15

... uh, with a sort of uniform API that you've, you figured out. Um, any other sort of element, and obviously you can plug in all this, uh, persistence layers-

Matei Zaharia10:23

Mm-hmm

Swyx10:23

... uh, and, and compute layers. This is a whole cloud. It's an Agent Cloud.

Matei Zaharia10:26

Yeah, it's, I mean, it's got these components to work with it. The, you know, a lot of the action happens like on the machine where you deploy your agent to. So whatever you've got on there, you can run.

But yeah, it's, I think it's sort of the minimal thing you, you want to have hosted, like collaborative agents and to have that server. And one of the reasons we open source it is, uh, anyone building agents, this gives them an app they can start with and customize, uh, which w- we were seeing in Databricks too.

Like someone would make a nice, you know, agent app, and then other teams would ask, "Oh, can I just use yours for my agent?"

Reynold Xin10:59

Yeah, I think we had like five or six different agentic frameworks-

Matei Zaharia11:02

Yeah

Reynold Xin11:02

... built by every different team. They do all do more or less the same thing.

Swyx11:05

Yeah, you need to... Basically, people wanna take something that works and fork it-

Matei Zaharia11:08

Yeah

Swyx11:08

... and you might as well have something open source. Yeah, which, which also was another question, uh, which is-

Matei Zaharia11:13

Mm-hmm

Swyx11:13

... interesting for Databricks, like what do you choose to open source? What do you choose-

Matei Zaharia11:15

Mm-hmm

Swyx11:15

... to make it proprietary? It's in... I mean, this goes back to Spark, right?

Matei Zaharia11:19

Yeah. One, so I mean, one of the reasons to open source something is if you think it's a layer that will actually... There'll be some network effect, it'll benefit from many, uh, people collaborating, um, on it. So, uh, for example, with Spark, I don't know if you, if you know when, when, when Spark came out, we, we also focused a lot on letting you have libraries on top.

So like there used to be different-

Swyx11:42

Ecosystem

Matei Zaharia11:42

... distributed computing engines for like machine learning and graph computation. We said they should all be libraries that you can compose, and we made it super easy to add connectors to data sources too. And then we benefit because, you know, we, we don't have the time to write like connectors to like, you know, a thousand like different databases and, and file formats.

But we can just use the ones people make, and of course, they benefit from joining, uh, you know, kinda this, uh, this thing. So that's like one of the reasons. Another way to think about it is like imagine, you know, uh, we d- our thing wasn't open.

We had some kind of agent hosting thing, but it's not open, and then there is an open one. I- i- if you're... Which one's gonna win in the long run? So like here, because there is this benefit from like people writing integrations, it'll be, it'll be that.

And then there are other things that like you just can't, uh, even deliver as open source that are things the company does. Like, for example, how do you make sure you're like streaming, you know, jobs or your, your Lakebase database doesn't like, you know, lose all your data at night?

Well, that requires an, an operational team that's gonna sit there. There's no way it has to be a service. So like we wanna make sure as a company, we're really good at those infra services, and then we're as open as, as we can in terms of like what you build on top.

Reynold Xin12:56

I mean, speaking from a benefits, I think we are already seeing pull requests-

Matei Zaharia13:00

Yeah

Reynold Xin13:00

... of all kinds of ecosystem integration, even though it was only released on Saturday.

Matei Zaharia13:04

Yeah, Saturday. Yeah. So someone's-

Swyx13:06

Let's see, let's see what's going on

Matei Zaharia13:06

... March.

Swyx13:06

Yeah.

Matei Zaharia13:07

Yeah, you can look at the merge ones. Um, I actually asked Sam Najm this, this morning about the- 400 merge already? Yeah. Uh, ma- I, I think- Recent ... quite, I would guess around half are not from our team.

Uh, but for example, someone added support for running it on Kubernetes. Uh, people added, uh, many cloud sandboxes, so this can launch a cloud sandbox and run your agent in there, which is great for sharing too, 'cause it's not, like, on your laptop and someone's, like, running scary code on there.

Um, so yeah, many startups have put those in, and, uh, we expect to see more of them. We also have more agent harnesses already. Cursor, CLI, and Antigravity also. Yeah. That's all, uh, beautiful. I, you know, I, I feel like the last time this happens, there was the rise of the modern data stack.

Mm-hmm. I don't know if it's that useful. I'm, I'm actually kind of curious- Mm ... and your, your postmortem. Mm-hmm. I, I think most people- Agreed ... will agree that it is finally dead. Uh, but maybe this arises to a new modern AI stack that, like, does the same thing.

Mm-hmm. I don't know. I mean, I think the modern data stack was a pretty useful thing, probably even up until this day. I, I think what, uh, maybe for the audience who don't actually understand the history, I think the modern data stack is effectively decomposed into you need a layer to ingest the data in, you need a layer to transform your data, and then all of this are run...

And then you need a layer to maybe visualize your data, and all of this runs on some sort of data warehouse, or later on, uh, as we're doing data warehouse, also lakehouse. Mm-hmm. I think that concepts are all very powerful and very useful.

They sort of enable a lot of workloads. What people eventually run into is kind of a question of unification and consolidation is, hey, do you really need to chop all this into different pieces and work with so many different vendors and platforms in order to get, like, a very simple visualization done, right?

So I think, like, over time, everybody started realizing that customers are pushing us. We started, we can realize that, so we started building more and more capabilities and trying to consolidate. And at the end of the day now, customers don't have to worry about having me hook up five different systems in order to- Yeah ...

produce a chart. But the... I, I think honestly, something like this is probably happening, um, in how many different frameworks do you want to hook up together in order to produ- like, do a very simple agent. Just to be clear, I would say the core of this is this common API on top of all the harnesses.

So the API is basically, like, you've got an agent session, and you can send in a message or, like, a file, basically. That's what you can send in, and then you get out, you know, these streams as it's streaming text or as it's doing tool calls and...

Or the other thing you can send in is you can, like, tell it to cancel a turn. So that's the API. Now, the thing we did is we, we could get you that on top of, like, Claude Code running in a terminal, Codex, you know, Pi, OpenAI SDK, all that stuff.

We map them all to that same interface. So that is something that you'd have to maintain yourself if you built your own, like, agent orchestrator, and then whenever Claude changes its API, you gotta, you know, tweak your thing or it's gonna lose some messages.

So that's the thing that's valuable to maintain. Then on top of that, like, we built a few apps. I think we, we built a pretty cool UI and stuff, but that's, um... A- and, and we built a security and control piece, which, which I'm excited about.

But it's that common interface. So we don't... We... That doesn't try to be a stack, and in fact, you could plug in your, your own UI on top of this, uh, server. That... Uh, that's one of the use cases we care a lot about, 'cause we want to use this in, in our own products.

Yeah. It should be everywhere. Yeah. I think one of those things that is, is really interesting to me is, like, well, first of all, I'll, I'll endeavor to do everything and not call it the modern AI stack because like I think we need- Yeah ...

a different name. But like, yes, like, you know, um, so one of the first people that told me about compute, uh, sandboxing was Nikita from Neon. Mm-hmm. Because a lot of people think about Ni- uh, Neon as like, well, it's serverless PostgreSQL with, like, the separation of compute and storage and, uh, you know, instant branching and all those things.

But actually, every database company is also a compute company. Yeah. Yeah. And so he was actually showing to me his whole, his sandboxing solution. I don't think he have ever launched it. So our sandbox solution, the reason we could build it so quickly was because we realized if you just take the actual Lakebase architecture- Yeah, yeah ...

and remove the database from it, by the way- ... coming from Neon- Exactly, right ... you have this- Every database has it already, yeah. Now, there are some differences. For example, in the one to support this particular workload, it's important to have local persistence, um, because you want- Yeah ...

your state to persist. Your libraries, you don't have to install your library every time, right? Mm-hmm. Um, whereas the Neon architecture, because of the separation of storage from compute, you don't need persistent local disk. Yeah. So there's some differences.

Yeah. But the, uh... At the end of the day, yeah, it's, it's, uh- Yeah, so this is when you run, like, a, a coding sandbox, like if I use it, uh, uh, we have the dev environment internally at Databricks.

There's, like, many, many, like, tens of gigabytes of data just for, like, all the source code and, like, artifacts and stuff that I built, and I want that to come back next time, so. Yeah. But yeah. Before the show, we was talking about some sta-statistics that might be surprising at the adoption.

Mm-hmm. It could be internal, it could be external, whatever- Mm-hmm ... comes to mind, just to impress people the scale this is happening. So we, on the analytics side, I think we launched

Reynold Xin18:20

Maybe 50 or 60 million virtual machines a day across all three clouds, so we're sort of one of the biggest compute orchestrators out there. Stuff for sure for CPU compute.

Swyx18:28

Yeah.

Matei Zaharia18:29

Yeah.

Reynold Xin18:29

Um, the... And all of this process, I think exabytes of data, I, I joked about depending on which time zone you are, typically before you have breakfast, Databricks would have processed, processed exabytes of data already on that day.

Um, and on Neon, it's, it's actually pretty interesting, too. It's launching, I think, 13 million databases-

Swyx18:48

Yeah

Reynold Xin18:48

... a day now. Um-

Swyx18:49

Yeah, to me that was like a big-

Matei Zaharia18:50

Mm-hmm.

Reynold Xin18:50

And that's just like a-

Swyx18:51

Like, what do, what do you mean?

Matei Zaharia18:53

Yeah. At that point.

Reynold Xin18:54

And a lot of those were thanks to agent- agents and branching experimentation-

Swyx18:58

Yeah

Reynold Xin18:59

... because we made it so easy and so quickly, and thanks a lot to Nikita's team, to launch databases. It's, uh, the- so it's changing the way people use databases.

Swyx19:08

Yeah. Okay, we're gonna go into more database talk in a bit, but I wanna make sure we close up anything on Omnigents. Uh, you mentioned, uh, you were excited about the security and-

Matei Zaharia19:17

Mm-hmm

Swyx19:17

... control side.

Agent Security19:17

Matei Zaharia19:18

Yeah.

Swyx19:18

Uh, a lot of companies are figuring that out right now, uh, as well as the spend side.

Matei Zaharia19:23

Yep, yep.

Swyx19:23

Um, what have you found there?

Matei Zaharia19:25

Yeah, so I spent quite a bit of time talking to internal users, uh, developers, security team, you know, uh, uh, managers, a- and also lots of customers, and there's a few things. Like, first of all, one thing, you know, that immediately was...

became obvious is for security, uh, you know, there's this tension between, like, usability and security. And, um, the way people do, like, a lot of coding agents today have very basic things like you can tell me which tool patterns I'll allow or disallow or whatever.

It's like yes or no. But that puts you in a very tough spot. So just as an example, like, should my agent be able to read, you know, some confidential documents or what? Or, or let's say should it be able to install new packages from npm, which, you know, maybe it's, uh, it's compromised.

Yes or no? Like, maybe, maybe I wanna allow it. Should my agent be able to publish stuff to the company website? Well, if I'm using it to code on the website, yes. But should it be able to do both so it can, like-

Swyx20:24

Hmm

Matei Zaharia20:25

... grab a confidential document and be prompt injected and leak it? Probably not. So the thing we decided we need is stateful or what we call contextual policies, where you keep track of the state of that session. It's not like is it allowed to push to the marketing site or not, but like, hey, if it did a risky thing, like it installed, you know, a one-day-old package from npm, or it, it read, like, 1,000 confidential docs, then no.

Then don't, don't do it. Otherwise, maybe it's okay. That's one example of, like, moving that trade-off so it's both more secure and more useful by having a more powerful engine, essentially. This requires tracking sessions. The other piece that was interesting there is, like, there are these very low-level events it's doing, and you want some libraries on top that parse them.

Like, for example, we have a, uh, MCP server on Google Drive internally. It's got 60 API calls. Um, like, how do I know which of those, like, will share a document with stuff on the internet and which ones won't?

It, it's, it's annoying. So we designed i- in Omnigent the policy layer so that it's, it's functions and you can have libraries. Like, someone can make something that maps the low-level events to high-level ones, and then you write a policy about the high-level things that came out.

Um, so and that was-

Swyx21:40

This is related to the Panther, uh-

Matei Zaharia21:41

Yeah, P- Panther is... will help with that. Panther is-

Swyx21:44

Yeah

Matei Zaharia21:44

... kind of a similar idea on the event processing side, and it's Python-based versus a weird custom language. Uh, this is sort of more, as in real time.

Swyx21:54

I didn't know that.

Matei Zaharia21:55

Yeah.

Swyx21:55

Those things are happening, yeah.

Matei Zaharia21:56

Yeah. So yeah, but, but these are the cool things. I think the contextual or stateful part, and then the, the way it can be libraries, and that was another reason to make it open source, because others will write libraries and, like, we and our customers can use them.

And the final thing, because it's stateful, one of the states we track is how much you spent in that session. So I can... I've had, like, I, I ask an agent to debug something, and it spent $500 because it decided to read a lot of log files and-

Swyx22:24

Ooh

Matei Zaharia22:24

... burn a lot of tokens. Uh, but I can literally say, "Okay, launch a sub-agent to do this and cap it to spending $5." "Like, ask me for permission if it needs more." And because we're counting that within that session, it'll pop up and tell me, "Okay, you spent five, $5.

Do you wanna go on?"

Reynold Xin22:41

So important context here. Matei spent the last five years, a lot of his time was architecting Unity Catalog at Databricks-

Matei Zaharia22:48

Yeah

Reynold Xin22:48

... which is the governance layer for data.

Matei Zaharia22:49

That's right, yeah. Yeah.

Reynold Xin22:50

And he's sort of combining expertise at that layer together with all the AI governance he uses.

Matei Zaharia22:55

Yeah.

Swyx22:55

Do you-

Matei Zaharia22:55

But I also spent a lot of time being annoyed by coding agents and, and getting prompts. And also as the CTO-

Swyx23:02

Welcome

Matei Zaharia23:02

... I don't wanna end up on the front page as, like, I installed some weird npm package and leaked-

Swyx23:07

Yeah

Matei Zaharia23:07

... all the codes. So I'm especially paranoid, but also I have very little time, so I don't wanna sit there approving, like, do you wanna run a 20-line, you know, uh, uh, bash script, uh, yes or no? Um, so that's why I spend a lot of time figuring out, like, how can I make it as safe as possible and not annoying?

Swyx23:24

Yeah. Is safety and sec- uh, m- let's call it security, a bigger concern than token maxing or token budgets? You know, which one is, like-

Matei Zaharia23:33

Oh, yeah. They're both there. I mean, I don't know. I, I, I guess it depends on the type of company you are. So I think, uh, some companies, like, the budget is, is, uh, limited and, uh, you know, they, they really care a- about that, or-

Swyx23:48

I mean, you can be Uber and still be concerned, you know?

Matei Zaharia23:50

Yeah. Oh, yeah, totally. Yeah, yeah, yeah. If, if you have-

Reynold Xin23:52

I mean, for us, security is-

Matei Zaharia23:53

Yeah. Yeah

Reynold Xin23:54

... super paramount.

Matei Zaharia23:55

For, for us, security is, is absolutely critical as a, as a, you know, cloud provider. It's, it's the most important thing, and, uh, token maxing, you know, we, we're not so worried about it yet, but, but I've seen the op...

Like, for example, I talked to some consulting companies. They have, like, 100,000 employees who are all coding for customers. If those each spend, like, an extra $1,000 a month, that's, that's not fun.

Swyx24:17

Mm-hmm.

Matei Zaharia24:17

Um, you know-

Swyx24:18

Yeah

Matei Zaharia24:18

... we have, like, only a few thousand engineers.

Swyx24:20

What's the policy in Databricks? Is it, is it just unlimited or what's-

Matei Zaharia24:22

It's, it's unlimited, but we do... Um, you know, we use our own product to, like, analyze the traces and stuff, and we have a team that's, you know- You know, looking to optimize and, and to see if anyone's doing something weird.

And, uh, we actually had some really cool insights just from analyzing current tracers.

Swyx24:38

Yeah.

Matei Zaharia24:38

Like which models are better at, say, Rust versus like s- you know, TypeScript or whatever. So yeah, at least in our code base.

Swyx24:45

Yeah.

Matei Zaharia24:45

Yeah.

Swyx24:45

Amazing. Obviously, I have to ask the token question, obviously.

Matei Zaharia24:48

Yeah.

Swyx24:48

I think it's a-

Reynold Xin24:48

Yeah

Swyx24:48

... it's a key thing. But, uh, but yes, se- uh, security and control above that, and figuring out a sane layer that you can have some autonomy, but, uh, not too much.

Matei Zaharia24:57

Yeah. Yeah. And we wanna make it super easy. As an engineer, you should set a thing. So in Omnigent, you can ask your agent, "Set a policy on yourself to do this." So it can actually-

Swyx25:06

So if, if there's something I should be showing-

Matei Zaharia25:07

Yeah

Swyx25:07

... I, I don't, I don't see it on the GitHub, but, uh-

Matei Zaharia25:10

Oh, yeah

Swyx25:10

... you know, there's just-

Matei Zaharia25:10

Well, in the docs there somewhere.

Swyx25:11

Yeah, this is it.

Matei Zaharia25:12

You, you can look at it later.

Swyx25:13

Yeah, yeah, yeah.

Matei Zaharia25:13

Just look in the docs on-

Swyx25:14

Yeah, yeah

Matei Zaharia25:14

... contextual-

Swyx25:15

Yeah, yeah

Matei Zaharia25:15

... policies if you wanna see. Um-

Swyx25:18

I just like to sh- point people-

Matei Zaharia25:19

Uh, look at the built-in policies.

Swyx25:20

Yeah

Matei Zaharia25:20

Yeah.

Swyx25:20

If you want to, you know, follow up on this, this is exactly where to look, right? Yeah.

Matei Zaharia25:25

Yeah, yeah. Uh, yeah, and the story of these is like, I just wrote, uh, you know, like I wrote a doc with like 10 ideas for things before- ... as you were working on them. Well, th- that was like my wish list of things people asked, and I told the team like, "Hey, can you do like at least five of these for the launch?"

And then they just got back with all of them, so.

Swyx25:43

Oh, wow.

Matei Zaharia25:44

Um, so you can come up with, with more. But them- m- some of them are just meant to be examples. C- really you can intercept like any event the agent is making, and you can then either block or force it to ask the user or like allow, and you can update state to keep-

Swyx26:00

Yeah

Matei Zaharia26:00

... track stuff.

Swyx26:01

Yeah, 'cause y- you know, ultimately you're... I think, I think of you as like a systems designer.

Matei Zaharia26:04

Mm-hmm.

Swyx26:04

You let people plug in, right?

Matei Zaharia26:05

Yeah.

Swyx26:05

That's the whole modus operandi of what you do.

Matei Zaharia26:08

Yeah, yeah.

Swyx26:08

It's like-

Matei Zaharia26:08

And we, and we care a lot about also compose a... B- like, can someone else write a library that others use? Which-

Swyx26:13

Yeah

Matei Zaharia26:13

... this is meant to.

Reynold Xin26:15

There's also a batteries included philosophy here.

Matei Zaharia26:17

Yes.

Reynold Xin26:17

Probably very similar to how you did Spark, which is you could just start using.

Matei Zaharia26:20

Yeah, that's right. It has to be good out of the box at certain things, and then you can build your own things on top that like we, you know, we don't wanna do. But f- you know, in Spark, if you just wanna like, I don't know, like read a, a table or do like aggregation, it, it should be awesome at that s- out of the box.

Swyx26:37

Yeah. People want to catch up on Omnigent, they should watch your keynote.

Matei Zaharia26:40

Mm-hmm.

Swyx26:40

Uh, they should go through the GitHub and the docs. If they wanted to contribute or they want to build on this ecosystem-

Matei Zaharia26:46

Mm-hmm

Swyx26:46

... where would you call out as the most high leverage places to-

Matei Zaharia26:49

Mm-hmm

Swyx26:49

... get involved?

Matei Zaharia26:50

Yeah, do get involved in the Discord and in GitHub. Our team is, is there, is monitoring. And, uh, some of the things people ask for, we just built ourselves. Some of them, you know, we're, we're collaborating with, with them to build it.

Uh, and also tell us like-

Swyx27:03

Yeah, this gonna be very help-

Matei Zaharia27:03

... how you would like to use it. 'Cause I think especially for developers, like everyone wants it to work their own way, and a really good developer tool, uh, like you have to hear the feedback on all the ways and figure out the abstractions and how to let people customize.

So we love to hear, like if you think, "Hey, I, I, you know, I don't want it to work this way," uh, tell us. We really just wanna get that compatibility layer across agents and then let you do stuff on top.

Swyx27:28

Yeah. Um, is there any, uh, you know, in, in terms of like the startup side, I'm, I'm a founder.

Matei Zaharia27:32

Mm-hmm.

Swyx27:32

I want to-

Matei Zaharia27:33

Yeah

Swyx27:33

... I see an opportunity, I wanna get in front of you. What's your request for like a startup that like, you know, I wish some-

Matei Zaharia27:37

Oh, that you want to integrate with this?

Swyx27:38

... someone was working on this.

Matei Zaharia27:40

Oh, for a startup?

Swyx27:41

Yeah.

Matei Zaharia27:41

Mm-hmm.

Swyx27:42

Like, you know, you're, you got your own startup. It's doing well.

Matei Zaharia27:44

Yeah.

Swyx27:44

But like, you know, if you weren't working on your own startup, what, what is like obvious that you should...

Matei Zaharia27:49

Mm-hmm.

Swyx27:49

You advise many startups too, obviously.

Matei Zaharia27:51

I mean, I do think, uh, uh, just as a company with a lot of engineers, like anything that helps me make sense of how people are using-

Swyx28:00

Spend

Matei Zaharia28:00

... coding agents and, and-

Swyx28:02

Yeah. Analytics

Matei Zaharia28:02

... spend, but also quality or like you should write, you know, you should add this skill, or you should write this thing, or your agents are really horrible at tasks involving this survey, so like go spend time. That would be nice.

Um, yeah.

Swyx28:14

Yeah, the closest I found is, uh, this team, GitAI.

Matei Zaharia28:17

Mm-hmm. Oh, cool. Yeah.

Swyx28:19

They, uh, they started with like we will just do, uh, code and human attribution, but they're basically-

Matei Zaharia28:24

Mm-hmm

Swyx28:24

... building the analytics layer on top of them.

Matei Zaharia28:26

Yeah.

Swyx28:26

Uh, I do think like there are a bunch of like artificial analysis-

Matei Zaharia28:30

Mm-hmm

Swyx28:30

... is obviously, um-

Matei Zaharia28:32

Yeah, they have their ventures

Swyx28:33

... doing super well, uh-

Matei Zaharia28:34

Yeah

Swyx28:34

... with, with their stuff. Um, so there's, there will be people. I think, I think-

Matei Zaharia28:37

Mm-hmm

Swyx28:37

... this is like s- the domain of consultants first, but then people-

Matei Zaharia28:40

Yeah

Swyx28:40

... will actually build software that let's say have like the management plane-

Matei Zaharia28:44

Yeah

Swyx28:44

... for coding agents.

Matei Zaharia28:45

Yeah, I think there'll be a lot of insights there. You have it in other areas.

Swyx28:48

Okay. Well, uh, and then the other, uh, big thing is your dream engine. Uh, if you want to tell the, the story of, uh, of, um, you know, LTAP.

LTAP28:56

Reynold Xin28:59

So, and background with... Uh, I'm, I'm gonna make people listen to our Ankur Goyal ep- episode-

Matei Zaharia29:04

Mm

Reynold Xin29:04

... where we talk about single store-

Matei Zaharia29:05

Yeah

Reynold Xin29:05

... HTAP, and now all that history.

Matei Zaharia29:07

Yeah, yeah. The LTAP idea is actually pretty simple. Um, so if, if people have heard of the Ankur's, uh, talk about HTAP, it's effectively the world of databases... Sorry, there's like maybe a lot of context needs to be injected here.

The world of databases-

Swyx29:20

I am happy to be the database podcast that I'm- ... forcing people to like learn your databases, guys.

Matei Zaharia29:25

Uh.

Swyx29:25

You cannot vibe code with just markdown files.

Reynold Xin29:27

Yeah.

Swyx29:28

Like, okay.

Reynold Xin29:28

It's one of the most important fundamental- ... systems technologies out there. But the world of database effectively split into sort of roughly two halves. There's what we call OLTP databases, which are transactional, and think of your Postgres, your MySQL, your Oracle databases, and the other side is what we call analytics, and sometimes might have heard the term OLAP.

And the difference is, is on OLTP, you typically have maybe run some transaction on some event that looks up at one specific row, you update that row, right? It's a very row-oriented sort of, uh, data structure. And on analytics, you're trying to reason on the data.

You're trying to compute, "Hey, what's my revenue per store? What's my... How's my website doing every day?" And then you, uh, eventually want to probably end up running anal- uh, machine learning on it to predict, "Hey, how will my maybe sales be going in the future?"

Um, they are so very different architecture and everybody start with OLTP databases. Every app, when you become serious enough, that needs more than markdown files, you need to have a database, you want to lose your data, you want to have some transactional consistency.

But once you want to reason on the data, if you only have like- A hundred rows, it's probably okay to run it on your Postgres or your, on your sort of MySQL database. But once you have more data and want to run more complicated analysis, the very analysis might crush your sort of, uh, Postgres database.

So you start doing, getting data out of the OLTP database-

Swyx30:50

Replication.

Matei Zaharia30:50

Mm-hmm.

Reynold Xin30:51

Replicate them into the analytic systems.

Swyx30:54

Yeah.

Reynold Xin30:54

And just start-

Swyx30:54

Which for people, Elasticsearch is, like, a big-

Reynold Xin30:56

Yeah. So some of them actually get into Elasticsearch for, like, blocked analysis. A lot of our customers obviously get into Databricks to run more sophisticated things.

Swyx31:05

Yeah.

Reynold Xin31:05

And there's this term called CDC, uh, which is-

Swyx31:08

Change data capture

Reynold Xin31:09

... change data capture. Um, and what it does, it reads the binlog of the database, and if you don't understand what binlog is, it's fine. The, uh, but it's the little delta of the data, and it reconstructs based on the delta, the state of the database, um, on the analytics side.

But CDC is, like, a very painful thing. It, it's how basically standard in the industry, everybody uses it, but, um, it ends up being sort of... I think many data engineers ends up being waken up at, like, 3:00 a.m., um, because there's some pipeline thing.

Swyx31:36

Mm-hmm. You know, my, my explanation is, like, Airbyte is like a, you know, may- became a $5 billion company just doing CDC.

Reynold Xin31:41

Yeah, exactly. Yeah, CDC is, like-

Matei Zaharia31:43

It's hard

Reynold Xin31:43

... a very... It's one of the most boring- ... but one of the most fundamental operations, like, powering modern society.

Matei Zaharia31:51

Uh-huh.

Reynold Xin31:51

But it's so brittle that, uh, we joke that it's should be called continuous data corruption, because you might change your schema on your OLTP database, and then the CDC pipeline fails to handle-

Swyx32:02

Yeah

Reynold Xin32:02

... the schema change.

Swyx32:03

Yeah.

Reynold Xin32:03

And then everything goes out.

Swyx32:05

And I mean, there's all sorts of tricks that you can do, like, uh, you add in, like, some versioning or whatever, but yeah.

Reynold Xin32:09

Yeah. But it's a very, in general, very complicated... Like, I think at my keynote, I asked the audience put up their hand if they love their CDC pipeline. Only, like, maybe two people put it up. So if single store, like, about maybe a decade ago, I think the industry had this idea, hey, what if I built a single database that can handle both workloads?

Now I-

Swyx32:26

Which, like, by the way, every database person ever has ever-

Matei Zaharia32:29

Mm-hmm

Swyx32:29

... always dreamed about this.

Reynold Xin32:30

Yes. Yes.

Matei Zaharia32:30

Mm-hmm.

Reynold Xin32:30

This is the s- the holy grail of database engineering-

Matei Zaharia32:32

Mm-hmm

Reynold Xin32:32

... is why not build a single system that can do both of this? But it ends up just being a lot of compromises. Um, one, I think one of the first issue is that, hey, each... they say Postgres has a massive ecosystem, right?

You want to be using the tools that's built for Postgres. And Spark, for example, had a massive ecosystem. There's a lot of libraries you want to use. If you were to create now a new thing, you don't have a ecosystem.

You tend to create a new, smaller proprietary API, and you're lacking both, and it's also very difficult to make it performance-wise to be, uh, sort of comparable on either side. So it ends up being actually sucking on both.

And our whole idea of LTAP, it's kind of obviously a wordplay on the term HTAP, is that we think this is HTAP done right. HTAP wants to build a single engine for both. We think you can get 99% of what you need by unifying the storage, and just have a single storage layer.

And once you have the single storage layer, if your Postgres databases are writing data in a column-oriented format, everything analytics can just go read that data directly without any delay, right? There's no pipeline in between, so all the data will immediately be available for reasoning analytics.

I think I was telling some customers earlier, hey, when we talked about this is gonna be super useful a, for agents, I actually at first didn't really believe in it myself, even though we wrote that positioning.

Swyx33:53

Yeah.

Reynold Xin33:54

But then last night I was having dinner with a Australian customer, and they actually told me, "Oh, hey, one of the big issue we have is we have all these logs from our services, and we see SLA dips and want to investigate.

But then there's no way for those agents to even understand what's going on in the actual databases themselves. All we see is just, like, product telemetry of the database and the services." It would actually make those agents 10 times more powerful if understand, for example, who's actually placing those orders.

Um, what is happening? What exactly are they doing? So now I'm actually sold on our own message.

Swyx34:28

Yeah.

Reynold Xin34:28

Um, I think it's really kind of... It, it gets you basically the almost all of the benefits of the HTAP holy grail, which is, hey, make the data available immediate for reasoning analytics-

Swyx34:40

Yeah, I think, you know-

Reynold Xin34:41

... without any compromise

Swyx34:42

... in, in the way that humans are generally intelligent and want to have the ability and access to, to query anything-

Reynold Xin34:49

Yeah

Swyx34:49

... uh, e- while they do the work, they also need history and need context.

Reynold Xin34:51

Mm.

Swyx34:52

And, like, where else does they get context? That's ana- is the analytical workload.

Reynold Xin34:55

Exactly.

Swyx34:56

Yeah.

Matei Zaharia34:56

Yeah. And I, I remember when we had incidents with, with our databases, and engineers said, "Well, I can't just run a giant query on it to see what's going on because that's gonna bring down the database and hurt it even more."

Like, that's the kind of stuff that this gets rid of, because you spin up a whole separate fleet of machines that's doing the analytics. You're not overloading, like, the main database-

Reynold Xin35:16

Right

Matei Zaharia35:16

... that's still trying to serve stuff.

Reynold Xin35:18

Yeah.

Matei Zaharia35:18

Yeah.

Reynold Xin35:19

So this has been a dream for a while. Uh, what had to get done in order to get to today? Like, you know, um-

Matei Zaharia35:25

Yeah.

Reynold Xin35:26

I, I feel like, uh, you have announced variants of this several times, but it wasn't as clear as LTAP.

Matei Zaharia35:32

Mm.

Swyx35:32

Yeah.

Reynold Xin35:32

I think LTAP is like a-

Swyx35:34

So-

Reynold Xin35:34

Like, okay-

Matei Zaharia35:34

Yeah

Reynold Xin35:34

... we've got it, guys.

Matei Zaharia35:35

Testing, yeah.

Reynold Xin35:35

I was talking to somebody at Meta, and he was asking me, "Hey, what's the catch? Why is it possible now?" And I think the reality is we took a lot of time to actually work on the Lakebase architecture.

I mean, obviously a lot of it came from the Neon team, which is a separation of storage from compute. And it turned out it was just a tiny little step away, going from that to this LTAP idea, which is, hey, we just...

I- in the Neon architecture and in Lakebase architecture, we're writing data in row-oriented format to the open data lake. But in there, we're writing in Postgres pages. Actually, Ali and I were spending a lot of time debating, hey, can we actually just-

Swyx36:12

Mm-hmm

Reynold Xin36:12

... change that to write in column-oriented format? And we're just debating, and one day, one of our engineers who's, like, super smart came in, he's like, "Hey, I actually prototyped it. It works."

Swyx36:21

Wait, so prototype what?

Reynold Xin36:23

Prototype instead of storing the data in the data lake in the row-oriented format-

Swyx36:29

Column

Reynold Xin36:29

... like Postgres pages-

Swyx36:29

Yeah

Reynold Xin36:30

... write them in Parquet.

Swyx36:31

Yeah.

Reynold Xin36:32

Um, and he just make the observation that, hey, our storage fleet is, has a lot of extra idle CPUs. And we could use those CPUs to do the transcoding from row to col, where row is good for OLTP, but column is good for analytics.

Um, so let's do that transcoding at that time. And as a matter of fact, once you transcode the data, the data compresses better. So from those services writing to, for example, S3 or other data lake, like object stores, you can actually write them faster 'cause now they are now smaller.

Matei Zaharia37:03

Yeah.

Reynold Xin37:03

So there's no overhead, it's no compromise in performance-

Matei Zaharia37:06

Some CPU overhead.

Reynold Xin37:08

Yeah, b- because, but-

Matei Zaharia37:09

Yeah

Reynold Xin37:09

... we had extra CPUs anyway.

Matei Zaharia37:10

We had that fleet anyway, yeah.

Reynold Xin37:11

Um, so the, the debate ended. I mean, it's, it's one of the classics-

Matei Zaharia37:15

Mm-hmm

Reynold Xin37:15

... of a tech, uh, issue, of a lot of debate, but then somebody actually went ahead-

Matei Zaharia37:18

Yeah

Reynold Xin37:18

... and just tried to prototype and it worked.

Matei Zaharia37:20

But, like, something this strategic and important to the company, I expect there to be, like, a kickoff thing, like a design doc. Nothing like that.

Reynold Xin37:27

Nothing like that, Pi. He just... We, we were debating-

Matei Zaharia37:29

Yeah. Right

Reynold Xin37:30

... in many, many meetings-

Matei Zaharia37:31

Yeah

Reynold Xin37:31

... and then we're just debating whether it's possible or not from first principle.

Matei Zaharia37:34

Yeah.

Reynold Xin37:35

And then, uh, somebody just did it.

Matei Zaharia37:37

Yeah. I mean, if you set yourself up so people will do that, that'll be great. And that happened a bit with Omnigent too. I think if, if I just had a doc on, like, we can make these together, everyone would, uh-

Reynold Xin37:46

Yeah

Matei Zaharia37:46

... you know, would think, "Oh, what about this? What about this?" But then you- if you try it out, it, it helps. And then if you have real users and they bash it and, like, it's still working, or in, in this case, if you have the workload, you, you know what the workload looks like, you can just test the same pattern then.

Reynold Xin38:01

Yeah.

Matei Zaharia38:01

Yeah.

Reynold Xin38:01

Tech aside, which is very cool, this is, like, the most important thing, the culture of innovation, and you don't have to ask my permission. You don't have to-

Matei Zaharia38:10

Mm

Reynold Xin38:10

... like, do a whole form- formal process. Just do it, you know?

Matei Zaharia38:13

Well, especially these days, I think with-

Reynold Xin38:15

Yeah

Matei Zaharia38:15

... uh, AI, it's actually easier to build-

Reynold Xin38:16

But so, like-

Matei Zaharia38:17

... a prototype

Reynold Xin38:17

... I think you are very ra- I mean, I made a lot of C-suite of, like, large companies and, like, I think that at scale, things slow down, and I, I'm sure you felt it already. But somehow you have this core of people that, like, are exempt.

How? I think we hire and we work with really, really good people, and that's a very important part of it, and empowering them, but also spending a lot of time, maybe us in the trenches-

Matei Zaharia38:41

Mm-hmm

Reynold Xin38:41

... matter a lot also.

Matei Zaharia38:42

Yeah, I think, uh, I mean, I think first people c- can adapt to being in the larger company, um, so that helps. And we, we wanna make sure they know that they can try stuff, and, and settle debates and, and have a lot of examples of how it was done before, or launch a thing in beta or whatever.

Um, and then the other thing, I, I do think as a company, like, despite the size, we don't launch that many, like, products. We try to keep it pretty coherent. That's, that was actually the whole, like, sort of theory of the company, was like instead of having, like, 20 Amazon services, you need to set up like a analytics and machine learning stack, you just have one, and it's, like, the same API, the same semantics across all of them, the same copy of the data.

So that requires, like, unification. And then we basically added one more thing at a time. Like, we added storage with Delta Lake. We didn't use to do any storage. Then we added SQL, you know, we added, uh, machine learning platform stuff.

So but yeah, don't, don't do too many, but do those things well and, and, uh, that also helps, you know, helps keep it manageable.

Reynold Xin39:47

Yeah. The other thing we kind of encourage a lot is instead of building, sort of boil the ocean for everything, let's figure out how do we do it incrementally, how, how do we do it very quickly. Like many of our products-

Matei Zaharia39:57

Yeah

Reynold Xin39:57

... they're built in the span of weeks, and then we go to, hey... Like, usually my first question to whoever team is building is, "Who's the target customer? Who are you working with? Are you on a first name basis with them?

Are you texting with them?" Um, I think having that very tight loop.

Matei Zaharia40:14

Can you bring up another launch that comes to mind when, in this kind of thing? I just want to s-

Reynold Xin40:18

Omnigent itself-

Matei Zaharia40:18

Give, give examples

Reynold Xin40:18

... happened that way.

Matei Zaharia40:19

Mm-hmm.

Reynold Xin40:20

Yeah.

Matei Zaharia40:20

Who's the customer? Yeah. Omnigent was, uh, was, uh, more of an internal thing ac- actually, because we would use that for our, our developer and like, yeah, basically the whole, like, AI team got access to it and was using it, and we made sure it works from the beginning with our internal code base, which is a monorepo that's, like, enormous.

We gave them, like, some infrastructure. We gave them lots of, like, token, uh, capacity. Um, so it's all the developers. Yeah.

Reynold Xin40:45

Yeah.

Matei Zaharia40:45

We had others, I don't... This is, I think, a, a public story about-

Reynold Xin40:49

I was gonna ask, Marketplace, Open Sharing, all of them had... I, I just don't-

Matei Zaharia40:52

Yeah

Reynold Xin40:52

... remember exactly which ones publicly referenced.

Matei Zaharia40:54

Yeah.

Reynold Xin40:55

Yeah.

Matei Zaharia40:55

They had o- Well, very, very early in the company there was, like, a Delta Lake, which is the transactional-

Reynold Xin41:01

That's a good one

Matei Zaharia41:01

... storage layer we did. Um, we had, um, our largest customer at the time said like, "Okay, I need some... I want something in the cloud 'cause, you know, I, I... if the rest of our network is compromised, like this thing needs to be separate to store and, and query the events."

And then, uh, talked to us, he said, "Okay, this is the rate of events per second. This is, like, the freshness I want. Can you do it?" So that was, like, way larger than any workload we had, and we had our, um, engineer, uh, working on that, Mi- Michael Armbrust, and he worked just to make this work.

And once it worked for them, you know, it worked for everyone else. Yeah. This was early in the company, probably like four years in or something.

Reynold Xin41:38

20, 2018.

Matei Zaharia41:40

Yeah, '17, '18.

Reynold Xin41:42

Food, food company did.

Matei Zaharia41:42

Do you have other examples?

Reynold Xin41:44

I mean, there's-

Matei Zaharia41:45

Maybe you have others

Reynold Xin41:45

... Yeah, Clean Room, which is basically how you share data in a way without sharing-

Matei Zaharia41:49

Yeah

Reynold Xin41:49

... underlying data, but you allow specific operations. Those were done effectively initially just for two customers. I, I think the industry has a sense of, hey, maybe if you overfit to like one or two customers, it's gonna be really bad for you.

But I think the, uh, downside of overfitting is much smaller than the, the upside itself. And if you sort of try to be too ambitious and boil the ocean, it's a much bigger problem.

Matei Zaharia42:12

Yeah.

Reynold Xin42:12

'Cause you might end up actually having no customer.

Matei Zaharia42:14

Yeah, that's more, that's the more likely outcome.

Reynold Xin42:16

Yeah.

Matei Zaharia42:17

Uh, than, than you can sort of pivot from there. I, I do think there is such a thing as a bad customer that sometimes you should fire.

Reynold Xin42:22

They could exist sometimes if you drive... Uh, well, one of the challenge I think we, we probably see, and maybe many AI, so newer generation companies are seeing is, so tech companies are very, very different from non-tech companies or traditional enterprises.

Matei Zaharia42:36

Yeah.

Reynold Xin42:36

And, um, if you optimize everything just for tech companies, you might have various challenges-

Matei Zaharia42:41

Oh

Reynold Xin42:41

... scaling them outside of tech companies.

Matei Zaharia42:42

Mm-hmm. Okay, what like-

Reynold Xin42:44

Yeah

Matei Zaharia42:44

... what like top three differences that you always think about?

Reynold Xin42:47

Governance is a big one

Matei Zaharia42:48

I think, yeah, a, a big one is like, yeah, security, um, you know, data privacy, governance, all that stuff. So usually if you're building some kinda like B2B or developer tool, like your biggest market is gonna be enterprises, but it's just very different.

A company that's existed for like, you know, it's had some form of IT for like 30 years, they have so many legacy systems or they operate in a regulated space. Uh, whereas a startup or, uh, even like a, you know, like sorta m- m- more recent tech company, all the...

everything is new and sort of pristine. So yeah, it's just different, and if you've never worked with enterprises or been in one, you, you just won't know about it.

Reynold Xin43:27

Yeah.

Matei Zaharia43:28

Yeah.

Reynold Xin43:28

And the procurement process is probably quite different. There's actually far more stakeholders.

Matei Zaharia43:31

Yeah, that is one. Yeah.

Reynold Xin43:32

Uh-

Matei Zaharia43:32

Another piece that's interesting is I think some tech companies, uh, you know, people, uh, will say, "Oh, I can build that myself," right? I, I, I'll just build that myself. So, so then you go, but-

Reynold Xin43:42

I don't think people say that about Databricks, but, uh-

Matei Zaharia43:45

Um, yeah, but it depends on-

Reynold Xin43:45

They do, they do.

Matei Zaharia43:46

They do? I mean, the, the... Yeah, and it depends on the, the teams and things. So, but, uh, on the other hand, like many of the enterprises say, "Actually, I don't, I never wanna be in the business of building that.

Like, I don't want my, you know, whatever, I'm a retailer or something, I never wanna-

Reynold Xin44:00

Yeah, sell clothes, but-

Matei Zaharia44:00

... be down because like some weird like nerd like couldn't get streaming pipelines working." That is not what I'm doing.

Reynold Xin44:07

Yeah.

Matei Zaharia44:07

So-

Reynold Xin44:07

Yeah, this makes them great customers, to be honest.

Matei Zaharia44:09

Yeah. Yeah

Reynold Xin44:09

Right?

Matei Zaharia44:09

But you have to understand that it's hard without having worked there and stuff, like you may not appreciate.

Reynold Xin44:15

Look, I think they're all great. Uh, don't get me wrong, they have different challenges, but the, uh, uh, many of the tech companies, uh, for sure there's a lot, far more DIY.

Matei Zaharia44:24

On the flip side, you have people who are... They're very much experts in their domain, like they're building airplanes, they're, you know, designing medicines, whatever, and they just wanna bridge the technology, right? Like they don't wanna learn, you know, databases or whatever.

As cool as we think it is, even as, as interesting as the average software engineer might think it is to read a little bit, like they just never wanna know. They just say, "I have a, you know, giant like, you know, matrix or whatever with my, uh, clinical data.

Like how do I, you know, how do I like cluster it or whatever?" So yeah.

Reynold Xin44:54

Yeah. Yeah. That's true. Okay, so and then I wanted to actually build out the, the sort of dream engine, uh, vision. Um, where does this all lead? So one of the thing we, uh, realized maybe a couple years back is that actually every single database engine out there, especially on the analytics side, are kind of a decade old.

Um, pretty much everything that have reasonable traction are about a decade old, and they all started targeting some very specific narrow use cases, and then over time it's become more and more successful. They've grown in their ambition, and then they try to support more and more use cases.

But the fastest way to support those use cases tend to be hacked around the abstractions that were initially created that were not for those use cases.

Matei Zaharia45:37

Yeah.

Reynold Xin45:37

And then, but you can kind of support them more or less okay. And before you know it, after 10 years of organic evolution that way, it becomes a gigantic pile of shit. Um, the... And, uh, but that includes Databricks, and very, very few company or very few systems, I think, have the, uh, gut to say, let's go sh- start from scratch.

Let's go back to the drawing board and design knowing everything we know today after a decade of workloads and probably billions in revenue, let's attempt to rewrite it from scratch and actually make sure it will work and it can support all of these use cases.

So we started doing that, but it's a very ambitious project. Uh, by the way, you can search on Wikipedia, there's this thing called second system syndrome.

Matei Zaharia46:22

Yeah, I know that. Yes.

Reynold Xin46:23

Or second system effect.

Matei Zaharia46:25

Every developer must know-

Reynold Xin46:26

Mm-hmm

Matei Zaharia46:26

... what a second syndrome is.

Reynold Xin46:26

It's basically you built your first thing and it works out great, and the second one's bound to fail because you become too ambitious. And then you ask so many requirements.

Matei Zaharia46:34

Or like, you know, you think you know everything-

Reynold Xin46:35

Yeah

Matei Zaharia46:36

... and then you're like-

Reynold Xin46:36

You just-

Matei Zaharia46:37

... "I'm gonna design a perfect system this time."

Reynold Xin46:38

Yeah, and it turned out it's not perfect, and then it start failing and you're too ambitious, never launch, um, and you get killed. The, and the engineering team that actually started this, they were brilliant. I think we hired some of the best database engineers, um, on the planet into Databricks, and they were brilliant.

Thank God it's not their second system. Many of them- ... have built more than two in the past.

Matei Zaharia46:58

Ah, nice.

Reynold Xin46:59

But they were still worried about this, hey, building a database engine from scratch, I think the conventional wisdom is gonna take like five years to mature. This would be a very long-term project. It could fail. Um, I think one of the engineers kind of jokingly said, "Hey, maybe we just call it Reynold Stream Engine."

Matei Zaharia47:13

Mm-hmm.

Reynold Xin47:13

If we name after a co-founder, maybe we then may get canceled or killed. But I think they built something pretty remarkable. Um, they went back to... They, they kind of changed the way the database engines were built from a paradigm point of view.

Usually when you build a database engine, you read a lot of academic papers, you try to understand what are the latest algorithms and data structures, and you put them together and see if they work or not. And there's a high risk of failure there also because whatever that looks really good on paper might work out, might, might actually look really good in 70% of the workloads, but then it backfires on the other 30%.

Um, they actually went build a more of a factory for building the database. So they spent more time building this factory, and the factory takes the decade of traces we have. I think they count as like quadrillion data points in the trace table.

Matei Zaharia48:01

You don't drop anything? Or you see sample?

Reynold Xin48:03

We for sure sample, but-

Matei Zaharia48:05

Yeah

Reynold Xin48:05

... the, there's like massive amount of things. And the, uh, and they use that to build a model, um, like a machine learning model. Not an AL, a machine learning model. Machine learning model basically it can very, very quickly tell us how any algorithm and how any implementation would perform for any specific type of queries with very, very high fidelity.

And based on that, they can, uh, pick the most likely algorithm and data structure that will actually help with the different kinds of workloads.

Matei Zaharia48:35

Mm-hmm.

Reynold Xin48:35

Both at runtime as well as at implementation time.

Matei Zaharia48:38

Mm-hmm.

Reynold Xin48:39

Because there's like unlimited number of-

Matei Zaharia48:41

I mean, it sounds like you want to-

Reynold Xin48:42

... papers

Matei Zaharia48:42

... like route to different data structures.

Reynold Xin48:45

Mm-hmm. Yeah, I mean, if you think about it-

Matei Zaharia48:46

This is not one database

Reynold Xin48:47

... a single database has many things implemented-

Matei Zaharia48:50

Yeah

Reynold Xin48:50

... together. But you want to make sure they all work well-

Swyx48:53

Yeah

Reynold Xin48:53

... with each other, and then for any given operation, there might be more than one implementation, so we make it actually run really, really... Uh, reality is things, algorithms that work super well, for example, for very, very low latency might not work very well for, say, scanning through petabytes of data.

Swyx49:08

Yeah.

Reynold Xin49:08

Right? Actually, most often there's a trade-off there between throughput and latency.

Swyx49:12

What are the key dimensions like scale, throughput, latency? What, what, what-

Reynold Xin49:16

Yeah, scale-

Swyx49:16

Anything else?

Reynold Xin49:16

Uh, and, and the distribution of data.

Swyx49:19

Yeah.

Reynold Xin49:19

Right? How sparse the data is.

Swyx49:20

How hard-

Reynold Xin49:21

That matters-

Swyx49:21

Yeah

Reynold Xin49:21

... very a lot. Um, how frequently do you hit the same data?

Matei Zaharia49:24

Yeah, how many distinct values-

Reynold Xin49:26

Yeah

Matei Zaharia49:26

... and stuff like that.

Reynold Xin49:27

Those things matter a lot.

Matei Zaharia49:28

Yeah.

Reynold Xin49:28

Like number of distinct value basically impacts the memory consumption of your aggregation, your hash... Like at some point there's a hash table.

Swyx49:35

Somebody... I, I'm gonna, in my write-up, I'm gonna try to list all this out because I, I really want a taxonomy. To me, taxonomies are-

Reynold Xin49:40

Mm-hmm

Swyx49:40

... so helpful because it covers everything that you should think about.

Matei Zaharia49:43

Mm-hmm.

Reynold Xin49:43

I think if you actually try to list it out, try to be like a million different features.

Swyx49:48

I always want like, okay-

Reynold Xin49:49

It's not a trillion-

Swyx49:49

Give me like 12. Give me, you know.

Matei Zaharia49:51

Out there.

Reynold Xin49:52

Um-

Swyx49:52

You know, you know like a... Someone did, like, I think an Oracle paper in like 40 years ago, did like the, these are the eight fallacies of distributed systems.

Matei Zaharia49:58

Mm-hmm.

Reynold Xin49:59

Yeah.

Swyx49:59

Right? That kind of thing is super useful.

Matei Zaharia50:00

Yeah, that's-

Swyx50:01

It's like, okay-

Matei Zaharia50:01

Yeah

Swyx50:01

... think through these eight.

Reynold Xin50:02

But let me give you a very, uh, weird example, uh, but it actually has profound implication on performance, which is, like is your string just ASCII or does it have Unicode in it? How should you encode it?

Swyx50:13

Strings, I mean, strings are the most complex data types.

Reynold Xin50:16

Yeah. So the... And that, like for example, if string is super dense, you could actually convert every string into a... Like imagine you have to do a aggregation. Instead of having a hash table, you could actually have an array.

Because if your string is dense enough, if you only have 256 options, you don't need a hash table. You can just do array-

Swyx50:35

Yeah

Reynold Xin50:35

... lookup.

Swyx50:35

Yeah.

Reynold Xin50:36

Um, and that, that'll be far fast.

Matei Zaharia50:37

Yeah, because the string is like a country code or something.

Reynold Xin50:39

Yeah.

Matei Zaharia50:39

Yeah.

Reynold Xin50:40

So it's actually like probably millions of, uh, features in that model. But using that, they can, one, basically prioritize the different algorithms that might actually impact in practice. And many of them are very counterintuitive. There's actually things that you think, hey, might work super well, actually don't work that well in practice.

But also more importantly, at runtime, you can dispatch the right algorithm and structure.

Swyx51:01

I'm listening to, uh, to the, the dream. I feel like Databricks is, is doing a really good job of the incremental evolution. Do you have to hard cut to a new system at any point? Or like, you know-

Reynold Xin51:12

We designed it in a way that it can be incremental.

Swyx51:15

Yeah.

Reynold Xin51:15

So first we're releasing a new endpoint. Uh, but, but this goes to the broader ocean versus... W- what we wanted to do is wanted, by design, this new engine should be able to do everything we're able to do before and better, right?

It's been particular, the better part refers to very low latency, low latency workloads that can finish in 10 submarines. But we want to roll it out incrementally with incremental capabilities so it doesn't take like five years to actually, uh, see the light at the end of the tunnel.

Swyx51:44

I think that's a heroic task. I don't know what, what, what other way to say it. I am really interested in any, any sort of, uh, new workload and new databases. I mean, obviously, I think, uh, if, uh, I've maybe established that I'm a little bit-

Reynold Xin51:57

Mm

Swyx51:57

... of a database nerd. The transactional databases, sorry, the accounting databases, like the, the Tiger Beetles-

Reynold Xin52:02

Mm-hmm

Swyx52:02

... I don't know if you've, uh, seen those.

Reynold Xin52:04

What do they do?

Swyx52:05

Dual entry accounting database. Like it's just meant to really model like financial accounts and credit systems and-

Reynold Xin52:10

Oh, I see.

Matei Zaharia52:10

Mm-hmm

Reynold Xin52:11

... it's like a very specific-

Swyx52:12

Very, very high throughput.

Reynold Xin52:13

Yeah.

Swyx52:13

Yeah.

Reynold Xin52:13

Mm-hmm.

Swyx52:14

No, so what you're talking about how everyone like starts with-

Matei Zaharia52:16

Yeah

Swyx52:16

... a thing and then like-

Reynold Xin52:17

Oh, I see

Swyx52:17

... they scale up and then they tack on other things. It's exactly that.

Reynold Xin52:20

Mm-hmm.

Swyx52:20

And then, uh, I recently interviewed Simon from TurboPuffer.

Reynold Xin52:22

Yeah.

Swyx52:22

Same, same thing.

Matei Zaharia52:23

Yeah.

Swyx52:23

Like, well, and, and Chroma as well, like the, the, all the vector database companies of 2023-

Reynold Xin52:28

Yeah

Swyx52:28

... all are suddenly now just, we're just generalist, general storage block storage.

Matei Zaharia52:32

Vector database should have never been a separate category.

Swyx52:35

I, I think that used to be a hot take, now, now it's like the, the conventional wisdom nowadays. What i- what should be a separate category? You know, if everything becomes LTAP, like what's...

Reynold Xin52:45

I think the thesis of LTAP is we're not collapsing the databases at the actual query layer. We're just collapsing-

Swyx52:52

Indexing layer

Reynold Xin52:52

... the storage layer.

Swyx52:52

Yeah.

Matei Zaharia52:53

Mm.

Reynold Xin52:53

Um, and that's, uh, I think, a very important part. And we actually don't think it makes sense to collapse the query layer into a single, like HTAP style database. And part of it... By the way, the other thing I think a lot of people had is, hey, it would be nice if there's only one query language I have to worry about.

Instead of worrying about PostgreSQL and maybe Spark SQL, why not just one? But I don't think that's an issue for agents. Agents are very-

Swyx53:18

Mm

Reynold Xin53:18

... eloquent in PostgreSQL or Spark SQL. It's never gonna get confused. As long as the data is there and it's-

Swyx53:24

Yeah

Reynold Xin53:24

... accessible, um, agents will do fine. That, that might have been, uh, so-

Matei Zaharia53:28

Yeah, and then-

Reynold Xin53:29

Five years ago might have been a problem-

Matei Zaharia53:30

Mm-hmm

Reynold Xin53:30

... for humans.

Matei Zaharia53:31

That could arise over time also, but it, it should... And this is, leads to how to do things incrementally, right? Like we, we realize you don't need it right now. We don't need to solve that problem to have a lot of value, uh, from, from the current LTAP.

Swyx53:45

Yeah. Okay. I'm gonna end the pod with a little bit of more sort of spicier things. Um, everyone has like had to receive within a separation of storage and compute and try to build, uh, you know, the, the clouds.

Snowflake53:58

Swyx53:58

I had the same pitches from Snowflake.

Reynold Xin54:01

Mm-hmm.

Swyx54:01

How have you succeeded where they failed? That's rough.

Reynold Xin54:06

Well, I mean-

Swyx54:07

I mean, uh, respecting that they are a competitor-

Reynold Xin54:08

Yeah, yeah

Swyx54:09

... objectively, you have outpaced them. What is the, the core insight from your point of view that you, you guys just-

Matei Zaharia54:15

Mm

Swyx54:15

... went different directions?

Reynold Xin54:17

Probably the biggest fundamental difference, both companies started around the same time, both went to the cloud, both focused on storage from compute architecture. But the biggest difference, one is, uh, open. Like Databricks had never had the proprietary format, right?

We started with the open sort of, uh, ecosystem-

Matei Zaharia54:33

Mm

Reynold Xin54:33

... started with Parquet and then evolved into Delta and Iceberg and all that. It's like one big thing. I think it matters a lot. The other one is AI. Um, I mean, before 2022, October 2022, uh, when ChatGPT came out, we had always pitched Databricks as a machine learning plus data- And a lot of the platform were built with machine learning use cases in mind, and obviously AI is a little bit different, and Matei's like-

Swyx54:59

Mm-hmm

Reynold Xin54:59

... spent far more time there than I do. But, uh, the whole platform was... We, we never felt, "Hey, we're just a data infrastructure platform."

Swyx55:07

Like Databricks will be.

Reynold Xin55:08

Yeah.

Swyx55:09

We all-

Matei Zaharia55:09

I think they, they started with, like, they thought, "Okay, we'll just manage the most valuable data and try to make it really fast. For that, we'll have our own storage, you know, which is optimized with the engine, and then we'll just start at, like, the small amount of data that, like, the managers and whatever, you know, finance people and so on look at and make that super fast to serve."

And, um, you know, it was a, a different space. Whereas we started with, like, we'll do the bulk processing and ingest. Like, you've got a bunch of, you know, JSON log files, you've got whatever. We do that very large scale stuff, 'cause that's what Spark was, was for, the large scale MapReduce-like stuff.

And then we'll keep the, the, the data in a open format. Might be slower, but, like, it's already out there. You can consume it downstream. And, uh, it turned out that, you know, it's easier to, to go from that bulk thing that's really good at the scale and ingesting and, and super low cost, and create versions in it that have the speed and features of the, you know, super easy to use, like, smaller data for, uh, uh, business users thing.

So start open-

Reynold Xin56:17

And there was a lot-

Swyx56:17

Then optimize.

Matei Zaharia56:18

Yeah, start open and start large. Like, in some sense, we started upstream of them. And there was a time, actually, when we both, like, sort of listed each other as partners because we said if you used both solutions together, use Databricks for, like, your ingest and compute, and then serve the tables out of Snowflake, you get all the visualization, all the very fast stuff, like, that's great.

And then, you know, we both realized, like, customers were telling us, like, "Why do I need this other thing? Why can't I just query your tables?" And we said, "No, we're horrible at that. Like, please use our partner for the SQL warehouse stuff."

And then they realized that, like, wait a minute, so much of the compute is moving upstream into this, this other thing. We better stop that.

Swyx56:57

You have to go into each other's territory, yeah.

Matei Zaharia56:59

But I think we did start with, like, the bigger scope, uh, and with the open thing and, and that's important architecture. Like, as a... Again, it goes to enterprises. Like, if your company's existed for, like, 30 years, you've experienced, you know, being locked into Oracle and, like, all kinds of, like, crazy things, and if you're the CTO there and you're setting up the architecture for the future for your company, you're gonna wanna pick a, a foundation that's open.

And you only want, like, one way to manage data in your company, ideally. You don't want, like, seven different systems. Um-

Reynold Xin57:31

But, like, the open data format have won. Like, I think now every enterprise wants to put data in open data format, but, uh, it was actually very controversial, like, back then. I think five, six... When exactly as... One of the Snowflake co-founders actually wrote a blog called-

Matei Zaharia57:45

Yeah

Reynold Xin57:45

... Choosing Open Wisely, which basically argued against ...

Matei Zaharia57:49

Yeah, yeah.

Reynold Xin57:49

I think they might have taken it down. You have to find it on archive now.

Swyx57:52

Oh, I mean, it's, it's, it's never going away now. Uh, no, no, it's still there. I love the, the sort of perspective o- that only you guys will have because obviously you, you, you run the company. Uh, and I, I...

Thank you for adopting this. It's a incredible, uh, perspective. Would love to-

Reynold Xin58:06

Maybe one last one.

Swyx58:08

Yeah.

Reynold Xin58:08

Um, as you were talking-

Matei Zaharia58:10

Mm-hmm

Reynold Xin58:10

... I think... I have to give Ali a lot of credit.

Matei Zaharia58:12

Mm-hmm.

Swyx58:13

Yes.

Reynold Xin58:13

He's an incredible CEO. I think he is the perfect combination of IQ, EQ, technology obsession, execution, business acumen.

Matei Zaharia58:20

Mm-hmm.

Reynold Xin58:21

Um, and, uh, and he's also a founder, which makes a lot, make him, a lot easier for him-

Matei Zaharia58:26

Yeah.

Swyx58:26

Yeah

Reynold Xin58:26

... to, uh, mobilize and execute. Um, I think that's, uh-

Swyx58:29

Oh, that was, that was it? Oh , so j- uh, you have Ali, and he, they don't, like, okay .

Reynold Xin58:34

Well, they... Probably a lot of things, but I think Ali play a pretty big role in the, uh-

Swyx58:37

I was-

Matei Zaharia58:37

Yeah.

Swyx58:38

I was, I was, I thought he wa- there was, like, gonna be some technical, uh, choice that he, he contributed to.

Reynold Xin58:42

Oh, no, no, no. I, I, well, I mean he-

Matei Zaharia58:42

Well, he, he did for a lot of these. Like, there were sort of forks in the road where he pushed for, like, one way, and then it, it became clear that, like, that was the right way. Um, yeah.

Swyx58:51

I, I mean, there's a whole book that needs to be written about how, like, the eight of you, like, you know, work together-

Matei Zaharia58:56

Mm-hmm

Swyx58:56

... and all that. I think there's been profiles that people have done. Second one, not a cleared, uh, question a- again. Uh, Mosaic.

Reynold Xin59:03

Things that are there. Oh.

Mosaic Models59:03

Swyx59:04

Mosaic.

Reynold Xin59:04

Yeah.

Swyx59:05

A lot of people in our community are in, are curious on, like, what's the sort of the model story of Databricks, right?

Matei Zaharia59:10

Mm-hmm.

Swyx59:10

Like, when you guys bought Mosaic, like the, the thing was like, okay, well, we're gonna do fine-tuning, we're gonna do-

Matei Zaharia59:15

Mm-hmm

Swyx59:16

... uh, in-house model 'cause they had, uh, the Mosaic models, and it seems like you're, you're not doing that, and it seems like you're going towards more of the, uh, LTAP and, and, uh, the harness stuff. What's the story there?

You know, just-

Matei Zaharia59:28

Yeah. I guess when, when Mosaic started, I think it was, it was well-known or became most well-known for releasing open source LLMs early on, and, and they were general models. Actually, before that, they were doing other things. They were about optimizing, uh, training systems, basically.

So they had the fastest, like, image model training stack in, in the world and stuff like that. And then they decided to do LLMs, which, which was smart. They, they moved into it before ChatGPT, so they had some of the first open source LLMs.

Swyx59:57

Yeah.

Matei Zaharia59:57

Um-

Swyx59:57

We interviewed Gianfranco and-

Matei Zaharia59:59

Oh, yeah

Swyx59:59

... Abi, Abi for MPC seven B.

Matei Zaharia1:00:00

Yeah, exactly, yeah. Oh, yeah, very cool, yeah. Yeah. So we, uh, decided, you know, even though we, we did launch a open source model DBRX and, you know, we, we went up to, like, sort of above the Llama 3 scale, we decided that we really wanna focus on...

There'll be so many people releasing models, and w- instead of doing the general model where, like, you know, a big part of the recipe is just throw in a lot of compute and, and just scale, uh, we wanna focus on, like, the next step also of, uh, let's say you have the very smart model, how do you make it, you know, useful?

Uh, for us, it was a lot about automating, like, how... Like, like, making it very good at querying data. That's the, the first party agents we have called Genie. Uh, so it's like a virtual data scientist. Imagine, you know, there's someone who already knows all the stuff in your company inside out and knows all the machine learning libraries, all the data libraries, you know, all the stuff on the web, and you can ask them questions, you know?

That's, um, that's what we wanted to do first. So that meant, like, let's not focus as much on, like, let's just train some kind of frontier model, but let's build a system using either external models or, or, um, fine-tuned, uh, uh, customized components.

Um, we're still doing quite a bit of model training though, and in fact, we're always... You know, we're procuring, like, lots of GPUs and stuff all the time to do it. Um, and there's a few places where we're doing it.

One is, uh, there are many high volume use cases where if you have a specialized model, it's just so much better than any of the, of the general models you get. A nice example of that is understanding, like, documents, like PDF, Word documents, stuff like that, parsing them.

If you've ever tried to do that, it's frustrating 'cause you send it to, like, you know, like Claude, Fable, or whatever, it, like almost gets it, but it gets some things wrong, and it's super expensive. You just burnt a huge amount of tokens plopping in an image into there.

So our team, uh, built this, uh, document, uh, sort of vision model that takes a page and gives you back a nice JSON with all the components, and it's very competitive. It's, like, probably, like, 100X cheaper than those, uh, frontier models and still better.

Swyx1:02:11

Yeah.

Matei Zaharia1:02:11

And that's actually done by one of the researchers who came from DeepMind, was a co-founder of Adept, like very early LLM-

Swyx1:02:19

Mm-hmm

Matei Zaharia1:02:19

... scaling person, but, um, but focused on, on this. Um, likewise, we have, um, uh, we're doing specialized sub-agents for part of what the coding agent does. And if you've seen the stuff on advisor models, uh, from Harvey-

Swyx1:02:32

Yes

Matei Zaharia1:02:32

... um, also from-

Swyx1:02:33

Anthropic has been putting on

Matei Zaharia1:02:34

... and Anthropic and-

Swyx1:02:34

Commission also.

Matei Zaharia1:02:35

Yeah.

Swyx1:02:35

Yeah.

Matei Zaharia1:02:36

And UC Berkeley, actually one of my grad students there, uh, wrote a paper called Advisor Models, I think before those came out. I mean, I'm sure others had the idea at the same time.

Swyx1:02:45

Yeah.

Matei Zaharia1:02:45

But that's, uh, something that helps a ton. So yeah, we actually showed some, some stuff just today at the keynote on, uh-

Swyx1:02:52

Is it Parth? Oh, you know Parth?

Matei Zaharia1:02:53

Parth, yeah, yeah. Parth is-

Swyx1:02:53

Oh, he's speaking at my thing. Uh, he's doing-

Matei Zaharia1:02:55

Oh, nice

Swyx1:02:55

... continual learning bench.

Matei Zaharia1:02:56

Yes, yes.

Swyx1:02:57

Uh-

Matei Zaharia1:02:57

Yeah, I'm one of his advisors, uh, at Berkeley.

Swyx1:02:59

Oh, yeah.

Matei Zaharia1:02:59

Yeah.

Swyx1:02:59

We interviewed his brother Chai.

Matei Zaharia1:03:01

Oh, okay.

Swyx1:03:01

'Cause he's also at Abridge.

Matei Zaharia1:03:02

Yeah, yeah, yeah. Cool.

Swyx1:03:03

Uh, that, that family's very smart.

Matei Zaharia1:03:05

Yeah, yeah. Yeah. They're, they're awesome, yeah. So yeah, so we're doing some of that and, and as we get experience with these in the first party agents, we're also doing them with customers. So my, my feeling is, like, um, customizing models is actually gonna get way easier over time.

That's what we're finding, 'cause the base models are smarter, so they generate better traces in RL already, and then RL is about learning from your own past traces. And then synthetic data generation is way better, way easier now.

Uh, we have pipelines just using open source models. Like, the same model generates training environments and trains itself and beats, like, Opus and GPD 5.5 and stuff at a task. So I do think it's gonna pick up. Like, m- m- you know, the thing is, the ease of training the algorithms is only gonna go up over time.

There's a question of when it crosses into mainstream. Like, instead of this like, you know, specialized document parsing thing we did where, like, you need a hardcore LLM researcher, when does it get easy enough that anyone can, like, sort of plop in some stuff and describe a task?

Swyx1:04:07

Yeah.

Matei Zaharia1:04:07

Yeah.

Swyx1:04:08

Well, you know what makes it easy? Interfaces.

Matei Zaharia1:04:10

Yeah, yeah.

Swyx1:04:10

And, uh, unified APIs.

Matei Zaharia1:04:12

Exactly.

Swyx1:04:12

'Cause obviously if it's not interoperable, then you cannot switch.

Matei Zaharia1:04:14

That's what we're seeing with these, like, with, with Omnigent and the-

Swyx1:04:18

Yeah, yeah

Matei Zaharia1:04:18

... composable agents, like you can have sub-agents or with specialized models, and then you can train the whole thing. I think that'll help a lot, too.

Data Thesis1:04:25

Swyx1:04:25

Yeah. The last thing I was gonna leave, actually this is, I'm sequencing this, so I'm actually kind of proud of myself. Satya, uh, is, uh, is, uh, you know, talking about this. Uh, I, I interviewed him at, uh, Microsoft Build-

Matei Zaharia1:04:36

Mm-hmm, yeah

Swyx1:04:36

... a couple weeks ago, and then he wrote this essay, which I'm sure you've seen-

Matei Zaharia1:04:39

Mm-hmm. Yes

Swyx1:04:40

... uh, which is, uh, talking about building frontier ecosystem. He sounded, uh, when I was talking to him, more like a Databricks CEO than I've ever

Matei Zaharia1:04:49

Uh-huh.

Swyx1:04:49

Um, uh, uh, what... Is there, is there a cur- I mean, uh, this thing presumably went viral in my circles. I don't know if it did in your circles.

Matei Zaharia1:04:55

Mm-hmm.

Swyx1:04:55

What, what's the sort of theory of like, uh, you know, uh, I guess tokens as IP, building up the context, you know? He, he basically said everything but data is the new oil or context is the new oil.

Some, some version of that-

Matei Zaharia1:05:06

Mm-hmm

Swyx1:05:06

... you know, that, that you guys have heard before.

Matei Zaharia1:05:08

Yeah, I agree. I think the, the data you have, as you get better technology around it, like you can just do more in your domain with it. It's not even just about AI. Even when people, uh, you know, started collecting stuff in real time, like I remember all the power companies put like the smart meters and stuff, and all the car manufacturers started putting like sensors and cameras and stuff.

Any technology like makes data more valuable and can give you some advantage. B- anything that helps you do something with it and make some decisions, and AI is the same way. Like you had all this stuff that's just sitting there, now you can have an agent automatically tell you.

Like for example, you know, instead of I discover there's a, well, a feature in my product is broken 'cause a customer complained, the agent tells me, "I notice no one is like uploading files anymore 'cause they, they get errors or whatever."

And as you saw with like Raydian, like as a database company, because we have all these, the history of all the queries and all the table layouts and like how they worked, we can build a new engine very quickly that, uh, actually is, is good, and w- we're confident that it's gonna be good.

So I think, I think this is right. I think the, the question is, is exactly how it will, uh, land. But I do think like custom, model customization, which Satya talked about, is gonna get easier over time.

Swyx1:06:24

Yeah.

Matei Zaharia1:06:24

Um, and-

Swyx1:06:24

Which is why, by the way, I brought up the model thing, 'cause they have their MEI things and you guys don't. That's the, that was the, to me the mental question.

Matei Zaharia1:06:31

Yeah. We, we do have, um... We're doing like RL fine-tuning as a service a- and, uh, with, uh, with a bunch of customers. We don't have like... Basically, you know, we, we have like preview customers, and we have a general something called AI Runtime that's like we get you GPU clusters on demand with a software stack in there that makes it easy to do training.

So we didn't like sort of launch-

Swyx1:06:53

Do offensive aid, yeah

Matei Zaharia1:06:53

... but that's existed for a while. We've had like GPU compute for a while, and that's where a lot of the Mosaic, uh, stack went to-

Swyx1:07:00

Yeah

Matei Zaharia1:07:00

... to help scale that. But yeah, we found that the engagements, like some of the... There's two types of customers. There's some who just want GPUs and libraries to like get data in and out and monitor, so that's what AI Runtime is.

And then there's some that say, "Hey, you know, can you actually work with me, build evals, build synthetic data-

Swyx1:07:19

Yeah

Matei Zaharia1:07:19

... and create-"

Swyx1:07:19

The more forward deploys-

Matei Zaharia1:07:20

Yeah

Swyx1:07:20

... solutions architects.

Matei Zaharia1:07:21

And then that's what we're doing. And, and as... and more things will transition from like being custom to, to not. But, um, th- that's sort of how it is today.

Swyx1:07:29

Going back to your original question, I think one of the thesis we have is actually the, uh, once you can get the data in the right place, the AI models are becoming pretty good. The generic agents are fairly...

I mean, Ali talked-

Matei Zaharia1:07:41

Yeah

Swyx1:07:41

... about AGI is already here. They have pretty good reasoning capabilities. Actually, I think many of the traditional software will be sort of, uh, rewritten, uh, with this new paradigm, which is just get the data to be there, and then just slap some agent on top.

Magic will come out.

Matei Zaharia1:07:56

Yeah.

Swyx1:07:56

Um, but without the right data, you can't really do that. And it's actually our approach going to security and our approach going to the, uh, s- customer data platform space-

Matei Zaharia1:08:05

Yeah

Swyx1:08:05

... is, uh, like we launched two products-

Matei Zaharia1:08:08

Yeah

Swyx1:08:08

... at Data and AI Summit, one targeting sort of security teams and the other one targeting marketing teams. And those all are, have a lot of existing technologies out there, and our, I think our approach is just, hey, once you get the data in, everything is a lot easier with agents on top.

Closing1:08:23

Matei Zaharia1:08:23

Yeah, yeah. Well, and you guys have been fantastic guests. I, I just love this discussion. I, I just love the ability to dive in on the tech side, but also culture and strategy. I hope this isn't the last time we chat.

Like, I mean, congrats on all the success so far.

Reynold Xin1:08:37

Thank you.

Matei Zaharia1:08:38

Yeah.

Reynold Xin1:08:38

Congrats on-

Swyx1:08:39

Thanks

Reynold Xin1:08:39

... your success also.

Matei Zaharia1:08:41

Yeah.

Swyx1:08:41

Yeah, yeah. I mean, uh, Databricks is actually supporting my, uh, event, which is, uh, it's a... So I run-

Matei Zaharia1:08:45

Mm-hmm, yeah

Swyx1:08:46

... annual con- conference. And it is actually... I was, I, I've been attendee of Data + AI Summit-

Matei Zaharia1:08:51

Mm-hmm

Swyx1:08:51

... for a long time, and I noticed that it was like kind of... Th- this was back in 2022. It was like 90% data and then 10% AI.

Matei Zaharia1:08:58

Yeah.

Swyx1:08:58

And I was just like, "Well, okay, like we need a, we need the community thing that is like just 90% AI."

Reynold Xin1:09:03

Yeah.

Swyx1:09:04

Which like now everybody is.

Matei Zaharia1:09:05

Yeah, yeah. No, we're excited to support.

Swyx1:09:06

Uh, so yeah, yeah. So Databricks will be, will be at the conference. Uh, and I, you know, I, I just, it's just amazing to see you guys, uh, build out the most like interesting like cloud that I have s- I've seen outside of like the, you know, the, the, the, the big three.

And like it's amazing how far you've grown. Like, uh-

Matei Zaharia1:09:21

Thank you

Swyx1:09:21

... one of the, one of the most, uh, insightful, like, you know, I don't... I'm not a VC, but I play one on TV. Um, like Ben Horowitz, like-

Matei Zaharia1:09:28

Mm-hmm

Swyx1:09:28

... when he was talking to you guys, advising you on just like where is this company going, he was like, "Don't sell it to 100 billion," or some-

Matei Zaharia1:09:35

Mm-hmm

Swyx1:09:35

... some version of that story, right?

Reynold Xin1:09:36

Yeah, it was like the, the company should be worth a trillion dollars. You're underselling it for 10 billion.

Swyx1:09:40

And like he doesn't do that for everyone, you know? Like like, like for some, some reason, like, you know, I, I think he saw the vision, but also, uh, the infinite runway that you have.

Reynold Xin1:09:50

We're lucky to have Ben. Yeah.

Swyx1:09:51

Yeah.

Reynold Xin1:09:51

He's a big supporter.

Swyx1:09:53

Yeah, amazing. Okay, well thank you so much.

Reynold Xin1:09:55

All right. Thank you so much, swyx.