Origins0:00
Hey, everyone. Welcome to the "Latent Space Podcast." This is Alessio, partner and CTO at Decibel Partners, and I'm joined by my co-host, Swyx, founder of Smol.ai.
Hey, and today we're in the studio with Stan Polu. Welcome.
Thank you very much for having me.
Visiting from Paris.
Paris.
And, uh, you have had a very s- uh, distinguished career. I-- It's very hard to summarize, but, uh, you went to college in, in both, uh, École Polytechnique and Stanford, and then you worked in a number of places, Oracle, Totems, Stripe, and then OpenAI, pre-ChatGPT.
Uh, we'll talk... we'll spend a little bit of time about that. About two years ago, you left OpenAI to start Dust. I think you were one of the first OpenAI alum founders.
Yeah, I think it was about at the same time as, uh, the Adept guys. So-
Yeah
...was, uh, that first wave.
Yeah. And, uh, people really loved our David episode. Uh, we love a few, like, sort of OpenAI stories, uh, you know, for back in the day, like, like we're talking about pre-recording. Probably the statute of limitations on, on th- some of those stories have-- has expired, so it's-- you can talk a little bit more freely-
Yeah
...without, without, uh, them coming after you. But maybe we'll just talk about, like, what was your journey into AI? Um, you know, you were at Stripe for, for almost five years. There are a lot of Stripe alums going into OpenAI.
I think the Stripe culture has come into OpenAI quite a bit.
Yeah. So I think the, the, the buses of Stripe people-
Yeah
...uh, really, uh, started, uh, uh, flowing in, I guess, after ChatGPT. But, uh, yeah, my journey into AI is, uh, is, uh-
Meaning Greg Brockman.
Yeah. From, from, from Greg, of course. Um, and, uh, and Daniela, actually, uh, back in the days. Daniela Moses.
Yeah.
Yeah.
She was, uh, COO? I mean, she is COO. Yeah.
She had a pretty high, uh, high job at OpenAI at the time. Yeah, for sure. My journey started, uh, as anybody else. Uh, you're fascinated with computer science, and you wanna make them think. It's awesome, but it doesn't work.
Well, I mean, it was a long time ago. It was, uh, like maybe 16, so it was 25 years ago. Then the first big exposure to AI would be at Stanford, and I'm gonna divu-- uh, like disclose a whole lie because, uh, at the time, it was a class taught by, uh, Andrew Ng.
Mm.
And it was-- there was no deep learning. It was half features for vision and A* algorithm, so it was fun. But it was the early days of, uh, of deep learning at, at the time. I think a few years after, it was the first project at Google, uh, but you know that, that cat face or the human face, uh, trained from many images.
Went to, uh, is dated doing a PhD, more in systems. Eventually decided to do, uh, to go into, uh, into, ha- getting a job. Uh, went at Oracle, started a company, did a gazillion mistakes, got acquired by Stripe, worked with Greg Brockman there.
And at the end of, uh, of Stripe, I started interesting myself in AI again. Felt like it was the time. You had the Atari games. You had, uh, self-driving, uh, craziness at the time. And I started exploring project.
It felt like the Atari games were incredible, but they were still games, and I was looking into exploring project that would have an impact on the world. And so I decided to explore three things: uh, self-driving cars, cybersecurity and AI, and math and AI.
It's like I think by a decreasing order of impact on the world, I guess.
Yeah. Discovering new math would be very foundational.
It is extremely foundational, but it's not as direct as- ...you know, driving people around.
And sorry, you, you're doing this at Stripe. You're, like, thinking about your next move.
No, no. I was at, at Stripe kind of a, uh, a bit of time where I, I started exploring. Did a bunch of, of work with friends on, uh, uh, trying to get RC cars to drive autonomously. Almost started a company in France or Europe about, uh, self-dive driving trucks.
We decided to not go for it because it would, like-- it was, like, probably very operational, and I think the ide- the iden- idea of the company-- of, of the team wasn't there. And also, I realized that if I wake up a day, and because of bug I wrote, I killed a family-
Mm.
Mm
...it would be a bad experience.
Mm. Yeah.
And so just decided like, no, that's, that's just too crazy. Then I explored, uh, cybersecurity with a friend. We're trying to apply transformers to cut fuzzing. So cut fuzzing, uh, you have kind of a, sorry, an algorithm that goes really fast and tries to mutate, uh, the inputs of a library to find bugs.
And we try to apply a, a transformer to that and do, uh, reinforcement learning with the signal of how much you propagate within the, the, the, the binary. Uh, it didn't work at all because the transformers are so, so slow compared to, uh, evolutionary algorithms that it kind of didn't work.
And then started interesting myself in, uh, math and AI, and, uh, started working on SAT solving with AI. And at the same time, OpenAI was kind of starting the reasoning team, uh, that were tackling that project as well.
I was in chat, uh, in, in touch with Greg, and eventually got in touch with Ilya, and finally found my way to OpenAI. Yeah. I don't know how much you want to dig into that. The way to find your way to OpenAI when you're in Paris was kind of an interesting adventure as well.
Joining4:33
Please. Uh, and all-- I wanna know, this was a two-month journey. You did all this in two months, this search.
The search for what, sorry?
Your search for your next thing, 'cause you left in July-
Yeah
...twenty nineteen, and then you joined OpenAI September.
I'm gonna be ashamed to say that, uh-
You were searching before. Yeah.
I was searching before.
Yeah. I mean, it's, it's, it's normal. It's normal.
No, the truth is that I moved back to Paris through Stripe.
Yeah, yeah.
And I just felt the hardship of being remote from your team, uh, nine hours away. And so it kind of freed a bit of time for me to start the exploration before. Sorry, Patrick. Sorry, John.
Well-
Hopefully they're listening.
Um, joining OpenAI from Paris and from, like... Ob- obviously, you, you had worked with Greg, but not every-- not anyone else.
No. Yeah. So I hadn't worked with-- I had worked with Greg, but not Ilya. But I had started chatting with Ilya, and Ilya was kind of excited because he knew that I was a good engineer through Greg, I presume, but I was not a trained researcher, didn't do a PhD, never did research.
And I was still chatting, and he was excited all the way to the point where he was like, "Oh, come pass interviews. It's gonna be fun." I think he didn't care where I was. He just wanted, wanted to try working together.
So I go to SF, uh, go through the interview process, get an offer, and so I get Bob McGrew on the phone, uh, for the first time. He's like, "Hey, Stan, it's awesome. You've got an offer. When are you coming to SF?"
I'm like, "Hey, it's awesome. I'm not coming to the SF, to SF." "I'm based in Paris, and we just moved." He's like, "Hey, it's awesome. Well, you don't have an offer anymore."
Oh, my God.
No, it wasn't as harsh as that-
Okay
...but that's basically the idea. Uh, and it took me, like, maybe a couple more time to keep chatting, and they eventually decided to try a contractor setup. And, uh, that's how I kind of started working at OpenAI officially as a contractor, but, um, but in practice really felt like being an employee.
What did you work on?
So it was a story focused on maths and AI-
Okay
... and in particular in the application. So the study of the large language models' mathematical reasoning capabilities, and in particular in the context of formal mathematics.
Okay.
The motivation was simple. Transformers are very creative, but yet they do mistakes. And, uh, formal math systems have the ability to verify a proof, and, uh, the tactics they can use to solve problems are very mechanical, so you miss the creativity.
And so the idea was to try to explore b- both together. You would get the creativity of the LLMs and the, uh, kind of verification capabilities of the formal system. A formal system, just to give you a bit of context, is a, is a system in which a proof is a program, and the formal system is a type system, a type system that is so evolved that it can verify the program.
If the type checks, it means that the program is correct.
Is the verification much faster than the in-
Yeah, yeah
... than actually executing program? It is, right?
Verifica- verification is instantaneous, basically.
Yeah.
So the truth is, is that what you code in involves tactics that may involve computation to search, uh, for solutions. So it's not instantaneous. You do have to do the computation to expand the tactics into the actual proof.
The, uh, verification of the, the proof at the very low level is instantaneous.
Yeah. How quickly do you run into like, you know, halting problem P and NP type things-
No
... like impossibilities where you're just like that?
I mean, you don't run into it at the time. It was really trying to solve very easy problems.
Okay.
So I think the-
Can you-- eg- what example of easy?
Yeah, it's a-- so that, that's the math benchmark that everybody knows today.
I see, the Dan Hendrycks one.
Uh, the Dan Hendrycks one, yeah.
Yeah.
And I think it was the, uh, low-end part of the math benchmark at the time, uh, because that math benchmark includes AMC problems, AMC 8 times 10, 12, so these are the easy ones.
Yeah.
Then AIME problems, somewhat harder-
Yeah
... and some IMO problems like Crazy Elle.
For, for listeners, we covered this in, in our Benchmarks 101 episode.
Mm-hmm. Awesome.
AMC is literally the, the grade of like high school grade 8-
Exactly
... grade 10, grade 12.
Exactly.
So you can solve this.
Yeah.
Um, just briefly to mention this because w- I don't think we'll touch on this again, there's a bit of work with like Lean and then with, uh, you know, more recently with, uh, with DeepMind doing like scoring like silver on the IMO.
Um, any commentary on like how math has evolved from your early work to today?
I mean, that result is mind-blowing. I mean, from my perspective, spent three years on that. At the same time, Guillaume Lample in Paris, uh, we were both in Paris actually, uh, he was at FAIR, was working on the same problems.
We were pushing the boundaries, and the goal was the IMO, and we cracked a few problems here and there. But the idea of getting a medal at an IMO was like just remote.
Yeah.
So this is an impressive result. And we can-- I think the, uh, DeepMind team just did a good job of scaling. I think there's nothing too magical in their approach. Even if it's-- it hasn't been published. There's a, a Dan Silver talk from seven days ago where it goes a little bit into more details.
It feels like there's nothing magical there. It's really applying, uh, reinforcement learning and scaling up the amount of data they can generate through autoformization. So we can dig into what autoformization means if you want.
Let's talk about the, the tail end maybe of the OpenAI.
Research9:27
Yeah.
So you join, and you're like, "I'm gonna work on math and do all, all of these things." I saw in one of your blog posts you mentioned you fine-tune over ten thousand models at OpenAI using ten million A100 hours.
How did the research evolve from, you know, the GPT-2 and then getting closer to like the DaVinci-003, and then you left just before ChatGPT was released. But like tell people a bit more about the research path that took you there.
Yeah. I can give you my, uh, my perspective of it. I think, uh, at OpenAI, there's always been like a large chunk of the compute that was reserved to train the GPTs, which makes sense. Uh, so it was pre-Anthropic splits.
Most of the compute was, uh, going to a product called Nest, which was basically GPT-3. And then you had a, a bunch of, let's say, remote, not core, uh, research teams that were trying to explore, uh, maybe more specific problems or maybe the algorithm part of it.
The interesting part, I don't know if it, if, if it was where your question was going, is that in those labs, you're managing researchers, so by definition, you shouldn't be managing them.
Mm-hmm.
But in that space, there's a managing tool that is great, which is compute allocation. Basically, by managing the compute allocation, you can message the team of where you think the priority should go. And so it was really a question of you were free as an, as a researcher to work on whatever you, you, you wanted, but if it was not aligned with OpenAI mission, and that's fair, you wouldn't get the, the compute allocation.
Okay.
As it happens, solving math was very much aligned with with the compute-- uh, with the, the direction of OpenAI, and so I was lucky to generally get the compute I needed to make, uh, to make good progress. Yeah.
What do you need to show as incremental results to get funded for further results?
It's an imperfect process. If you're working on maths and AI, obviously, there's kind of a prior that it's gonna be aligned with the company, so it's much easier than to go into something much riskier, I guess. You have to show incremental progress, I guess.
It's like you ask for a certain amount of compute, and you d- you deliver a few weeks after, and you-- so you demonstrate that you have a progress. Progress might be a positive result. Pro- progress might be a, a strong negative result.
And a strong negative result is actually often much harder to get or much, uh, much more interesting than a positive result. And then it's, uh, it, it generally goes into, as any organization, you would have kind of a people finding your project or any other project kind of, uh, cool and fancy.
And so you would have that kind of phase of growing up compute allocation for it all the way to a point, and, and then maybe you reach a, an a- an apex, and then maybe you go back to br- mostly to zero and restart the process because you're going in a different direction or something else.
That's how I felt, uh-
Explore, exploit.
Yeah, yeah, exactly. Exactly. Exactly. It's, it's a reinforcement AI-
Classic
... approach
... PhD student, uh, search process. And you were reporting to Ilya? Like, uh, the results you were kind of bringing back to him or like what's the structure? It's almost like when you're doing such cutting-edge research, you need to report to somebody who's actually really smart to understand if the direction is right.
So we had a reasoning team which was, uh, working on, uh, on reasoning obviously and, uh, and some maths in general. So and that team had a manager, but Ilya was extremely involved in the team as an advisor, I guess.
Yeah.
Since he brought me in OpenAI, I was lucky to, mostly for during the first years, to have, uh, kind of a, uh, direct access to him. He would really coach me as a- ... trainee researcher, I guess, with good engineering skills.
And Ilya, I think at OpenAI, he was the one showing the, showing the North Star, right? He was-- His job, and I think he really enjoyed that, and he did it super well, was going through the teams and saying, "This is where we should be going," and trying to, you know, flock the different teams together towards a, towards an objective.
Yeah.
I would say, like the public perception of him is that he was the strongest believer in scaling.
Oh, yeah.
He was- he's always pursued like the compression thesis.
Yep.
You have worked with him personally. What, what does, what does the public not know about how he works?
I think he's really focused on building the vision and communicating the vision within the company, which was extremely useful. I was personally surprised that he spent so much time, you know, working on communicating that vision and getting the teams to work together versus-
So, and to be specific, vision is AGI?
Oh, yeah, yeah. The vision is like, uh, yeah, it's, uh, it's the belief in compression and-
Okay
... and scaling compute. I remember when I started working on the reasoning team, he-- it, it was the excitement was really about scaling the compute around reasoning, and that was really the, the belief he w- he wanted to ingrain in the team, and that's what has been, uh, useful to the team and, and with the DeepMind results shows that it was the right approach.
With the, the, the success of GPT-4 and stuff shows that it was the right approach.
It-- Was it according to the neural scaling laws, the Kaplan paper, uh, that was published? Or-
I think it was before that because-
Yeah
... those ones came with-
Twenty twenty-ish
... GPT-3.
Yeah.
Basically at the time of GPT-3 being released or being ready internally. But before that, there's really was a strong belief in, in scale. I think it was just the belief that the transformer was a generic enough architecture that you could learn anything, and that it was just a question of scaling.
Any other fun stories you wanna tell where, you know, the-
David, Sam Altman, Greg, you know, any-
I didn't w-- Weirdly, I didn't work that much with Greg, uh, when I was at OpenAI. Uh, he was, uh, he's always been, uh, mostly focused on, on training the GPTs, and rightfully so. One thing about Sam Altman, he really impressed me because when I joined, he had joined not that long ago, and it felt like he was a kind of a, a very high-level CEO.
And I was mind-blown by how deep he was able to go into the subjects within a year or something, all the way to, to a situation where when I was having lunch, uh, by year two I was at OpenAI with him, he would just quite know deeply what it was doing and, uh, and what was important.
With, with no ML background. Like that's, you know.
Yeah, with no ML b- But I didn't, I didn't have any either so I guess that explains why. But I think you can-- you know, it's, it's a question about really, uh, you, you don't necessarily need to, to understand the very technicalities of how things are done, but you, you need to understand what's the goal and what's being done and what are the recent results and all of that in you, and we could have kind of a very productive, uh, discussion, and that really impressed me given the size at the time of OpenAI, which was not negligible.
Yeah. I mean, you've been a-- You were a founder before.
Yep.
You're a founder now.
Yep.
And you've seen Sam as a founder. How has he affected you as a, as a founder?
I think having that capability of, uh, changing the, the scale of, of your attention in the company.
Mm.
Because most of the time you operate at a very high level, but being able to go deep down and being in the known of what's happening at, on the ground is something that I feel is really enlightening. Uh, that's, that's not a place in which I ever was as a founder because first company we went all the way to ten people.
Current company, there's twenty-five of us. So the high level, the, the, the sky and the ground are pretty much at the same place.
No, yeah, you're being too humble. I mean, Stripe was also like a huge rocket ship.
And but Stripe I wasn't a found- Stripe, Stripe I wasn't a founder, so I was like at OpenAI, I was, uh, really happy being, uh, being on the ground, pushing the machine-
Yeah. Yeah, yeah
... making it work.
Yeah. Uh, last OpenAI question.
Yep.
The Anthropic split you mentioned.
Yep.
Uh, you were around for that. Very dramatic. Uh, David also left in, in around that time.
Yep.
You, you left. This year, we've also had a similar management, uh, shakeup, let's just call it. Can you compare what it was like going through that split during that time? And then like does that have any similarities now?
Like are we gonna see a new Anthropic e-emerge from these folks that have just left?
That I really, really don't know. At the time, the split was pretty surprising because they had been training GPT-3, it was a success, and to be completely transparent, I wasn't in the weeds of the split. What I understood of it is that there was a, a disagreement of the, uh, commercialization of that technology.
I think the, the focal point of that disagreement was the fact that we started working on the API and wanted to make those models available through an API. Is that really the core-
Yeah
... disagreement? That I don't know.
Or was it safety? Was it commercialization?
Yeah, yeah. Exactly.
Or did they just wanna start a company?
Exactly.
Yeah.
Exactly. That I don't know. But I think it-- the-- what I was surprised of is how quickly OpenAI recovered at the time, and I think it's just because the-- we were mostly a research org and the mission was so clear that some div-divergence in some teams, some people leave, uh, the mission is still there.
We have the compute, we have a, we have a shot at it, so let's just keep going.
Yeah. Very deep bench, like just a lot of talent.
Yeah.
Yeah.
So that was the OpenAI-
Dust Origins17:50
Intro
... part of the history. Exactly. So then you leave OpenAI in September 2022, and I would say in Silicon Valley, the two hottest companies at the time were you and LangChain.
Nah.
What was that start like, and why did you decide to start with a more developer focus, kind of like a AI engineer-
Yeah
... tool, rather than going back into doing some more research on something else?
Yeah. First, I'm not a trained researcher, so going through OpenAI was really kind of a, the PhD I always wanted to do. But research is hard. You're digging into a field all day long for weeks and weeks and weeks, and you find something, you get super excited for 12 seconds, and at the 13th second you're like, "Oh, yeah, that was obvious," and you go back to digging, to digging.
I'm not a trained, like formally trained researcher, and it wasn't kind of a necessarily a, an ambition of me of creating a, of having a, a research career, and I felt the hardness of it. I enjoyed a lot of like that a ton, but at the time I, I decided that I, I wanted to, to go back to, uh, something more productive.
And the other fun motivation was like, uh, I mean, if we believe in AGI, and if we believe the timelines might not be too long, it's actually the last train leaving the station to start a company.
Mm-hmm.
After that, it's gonna be computers all the way down.
Right.
And so that was kind of the true motivation for like, uh, trying to go, uh, to go there. So that's kind of the core motivation at the be- beginning personally, and the, uh, the motivation for starting a company was, uh, pretty simple.
I had seen GPT-4 internally. At the time, it was September 2022, so it was pre-ChatGPT, but GPT-4 was ready since a few... I mean, had been ready for a few months internally. I was like, "Okay, that's, that's obvious the capabilities are there to create an insane amount of value to the world."
Mm-hmm.
And yet the deployment is not there yet. The revenue of OpenAI at the time were ridiculously sl- small compared to what it is today. And so the thesis was there's probably a lot to be done at the product level to unlock the usage.
Yep. Let's talk a bit more about the form factor, maybe. I think one of the first successes, uh, you had was kind of like the WebGPT like thing, like using the models to traverse the web and like summarize thing, and the browser was really the, the interface.
Why did you start with the browser? Like, what... Was it import- And then you built XP1, which was kind of like-
Yeah
... the, the browser extension.
So the, uh, the starting point at the time was, uh, so if you wanted to, to talk about, uh, LLMs, it was still a s- a rather small community, a sm- a community of mostly researchers and some extent, uh, very early adopters, very early engineers.
It was almost inconceivable to just build a product and go sell it to the enterprise, though at the time there was a few companies doing that, the one on, uh, on marketing, I don't remember its name. Jasper. But so the, uh, the natural first intention, uh, first, first, first intention was to go to the developers and try to, to create tooling for them to create product on top of those models.
And so that's what Dust was originally. It was quite different than LangChain, and LangChain just, uh, beated the shit out of us- ... uh, which is great. Uh-
It, it's, it's a choice. You were cloud in closed source.
Yeah.
They were open source.
Yeah. So technically we were open source, and we still are open source, but I think that doesn't really matter. I had the strong belief from my research time that you cannot create an LLM-based workflow on just one example.
Basically, if you just have one example, you have a fit. So as you develop your interaction, your orchestration around the LLM, you need a dozen example. Obviously, if you're running a dozen example on a multi-step workflow, you start parallelizing stuff.
And if you do that in the console, you just have, like, a messy stream of tokens going out, and it's very hard to observe what's going there. And so the idea was to go with a new UI so that you could kind of introspect easily the output of each interaction with the model-
Right
... and dig into there through a new UI, which is-
Was that open source? I actually didn't come across it.
Oh, yeah, it was, I mean, uh, Dust is entirely open source even today. We're not going for an open source-
If it matters, I didn't know that.
Yeah, yeah. No, no, no. The, the reason why is because we're not open source because we're not doing an open source strategy.
Yeah.
It's not an open source go to market at all.
Yeah.
We're open source because we can, and it's fun.
Open source is marketing. You have all the downsides of open source, which is like people can clone you. Um, but then-
But I think, I think that downside is a, is a big fallacy.
Okay.
Yes, anybody can clone Dust today, but the value of Dust is not the current state. The value of Dust is number of a- eyeballs and, and hands of developers that are creating to it in the future. And so, yes, anybody can clone us today, but that wouldn't change anything.
There is some value in being, uh, open source. In a discussion with the security team, you can be extremely transparent-
Yeah
... and just show the code.
Mm-hmm.
When you have discussion with users and there's a bug or a feature missing, you can just point to the issue, show the pull request-
PR welcome
... show the, show the... Exactly. Or PR welcome. That doesn't happen that much. But, but you can show the progress. If the inter- if the person that you're chatting with is a bit technical, they really enjoy seeing the pull request, uh, advancing and seeing all the way to deploy.
And then the, the downsides are mostly around security. You never want to do security by obfuscation. But the truth is, is that your vector of attack is facilitated by you being open source. But at the same time, it's a good thing because if you're doing anything like bug bountying or stuff like that, you just give much more tools to the bug bountiers so that their, their output is much better.
So there's a, there's many, many, many trade-off. I don't believe in the value of the code base per se.
Wow.
I think it's really the people that are on the code base that, uh, that have the value, and the go to market, and the product, uh, and all of the things that are around the code base. Obviously, that's not true for every co- every code base.
Mm-hmm.
If you're working on a very secret kernel to accelerate, uh, the, uh, inference of LLMs, I would buy that you don't wanna be open source. But for product stuff, I really think there's no-- there's very little risk.
Yep. I signed up for XP1, I was looking, January 2023. I think at the time you were on DaVinci double oh three. Given that you were seeing GPT-4, how did you feel having to push a product out that was like using this model that was like so inferior, and you're like, "Please just use it today.
I promise it's gonna get better." It's like, just overall as a founder, like how do you build something that maybe doesn't quite work with the model today, but you're just expecting the new model to be better?
Yeah. So actually, XP1 was even on a, on a smaller one. That was the, uh, post-ChatGPT release small version, so it was-
Ada, Babbage.
No, no, no, no, not that far away, but it was, uh, the, uh, the small version of, uh, of ChatGPT basically. I don't remember its name. Yes, you have a frustration there, but at the same time, I think XP1 was designed...
was an experiment, but was designed as a way to be useful at the current capability of the model. If you just want to, uh, extract data from a LinkedIn page, that model was just fine. If you want to summarize an article on, on a newspaper, that model was just fine.
And so it was really a question of trying to find a product that works with the current capability, knowing that you will always have tailwinds as models get better and faster and cheaper. So that was kind of a...
There's a bit of a frustration because you know what's out there, and you know that you don't have access to it yet, but you-- it's also interesting to try to find a product that works with the current capability.
Platform24:51
And we highlighted XP1 in our anatomy of autonomy post in April- ... of last year, which was, you know, where are all the agents, right? So now we spent 30 minutes getting to what you're building now.
Yeah.
So you basically had a developer framework, then you had a browser extension, then you had all these things, and then you kind of got to where Dust is today. So maybe just give people an overview of- What Dust is today-
Yeah, yeah
... and the core thesis behind it.
Yeah, of course. So, uh, Dust really want-- We really wanna build the infrastructure so that companies can deploy agents within their teams. We are horizontal by nature because we strongly believe in the emergence of use cases from the people having access to creating an agent that don't need to be developers.
They have to be tinkerers, they have to be curious, but they can-- like, anybody can create an agent that will solve an operational thing that they're doing in their day-to-day job. And to make those agents useful, there's two focus, which is interesting.
The first one is an infrastructure focus. You have to build the pipes so that the agent has access to the data. You have to build the pipes such that the agents can take action, can access the web, et cetera.
So that's really an infrastructure play. Maintaining connections to Notion, Slack, GitHub, uh, all of them is a lot of work. It is boring work, boring infrastructure work, but that's something that we know is extremely valuable in the same way that Stripe is extremely valuable because it maintains the pipes.
And we have that dual focus because we're also building the product for people to use it. And there, it's fascinating because everything started from the conversational interface, obviously, which is a great starting point, but we're only scratching the surface, right?
I think we're at the pong level of LLM productization, and we haven't invented the Civ III, we haven't invented Counter-Strike, we haven't invented Cyberpunk 2077.
Yeah.
So this is, uh, really our, our mission is to, uh, to really create the product that let people equip themself to just get away all the work that can be automated or assisted by LLMs.
And can you just comment on different takes that people had? So maybe at the most open is, like, AutoGPT. It's just kinda like just try and do anything. It's like it's all magic. There's no way for you to do anything.
Then you had the Adept. Uh, you know, we had David on the podcast. They're very, like, super hands-on with each individual customer to build super tailor. How do you decide where to draw the line between this is magic, this is exposed to you, especially in a market where most people don't know how to build with AI at all?
Yep.
So if you expect them to do the thing, they're, they're probably not gonna do it.
Yeah, exactly. So the AutoGPT approach obviously is extremely exciting, but we know that the agentic capability of models are not quite there yet. It just gets lost. So we're starting, uh, we're starting where it works, same with, uh, Experiment.
And, uh, where it works is pretty simple. It's like, uh, simple workflows that involve a couple tools where you don't even need to have the model decide which tools it use in the sense of you just want people to put it in the instructions.
Mm-hmm.
It's like-
Right
...take that page, do that search, pick up that document, do the work that I want in the format I want, and give me the results. There's no smartness there, right, in terms of orchestrating the tools. It's, it's mostly using English for people to program a workflow where you don't have the constraint of having compatible API between the two.
That kind of personal automation, y- would you say it's kind of like, um, LLM Zapier type of thing? Like if this, then that, and then, you know, do this, then this, this. Is-- So it's very you're programming with English?
So you're programming with English, so you're just saying, "Oh, do this, and then that." You can even create some, some, some form of APIs. You say, "When I give you the command X, do this. When I give you the command Y, do this," and you describe the workflow.
But you don't have to create boxes and create the workflow explicitly. It just need to describe what are the tasks supposed to do, uh, to be, and make the tool available to the agent. Tool can be a semantic search.
The tool can be querying into a, a structured database. The tool can be, uh, searching on the web. Um, and obviously, the interesting tools that we only starting to scratch are actually creating external actions like reimbursing something on Stripe, uh, sending an email, clicking on a button in the admin or something like that.
Do you maintain all these integrations?
Today, we maintain most, most of the ins-- integrations. We do always have an escape hatch for people to-
Custom
... kind of custom integrate.
Open API spec, yeah.
But the reality is that-- The reality of the market today is that people just want it to work, right? And so it's mostly us maintaining the integration. As an example, a very good source of information that is tricky to productize is Salesforce because Salesforce is basically a database and a UI, and they do the fuck they want with it.
And so every company has different models and stuff like that. So right now we haven't-- we don't support it natively, and the type of support or real native support is-- will be, uh, slightly more complex than just auth-ing into it, like is the case with Slack as an example.
Because it's probably gonna be, "Oh, you want to connect your Salesforce to us? Give us the SOQL," that's the Salesforce QL language. "Give us the queries you want us to run on it and inject in the context of Dust."
So that's interest- interesting how not all int-integrations are equal, and some of them require a bit of work on the user. And for some of them that are really valuable to our users, but we don't support yet, they can just build them internally and push the data to us.
I think I understand the Salesforce thing, but let me just clarify. Are you using browser, browser automation because there's no API for something?
No, no, no, no, no. In, in that case-- So we do have browser autom-automation for all the use cases that imply the pub- the public web. But for most of the integration with the internal system of a company, it's really g-- runs through API.
Haven't you felt the pull to RPA, browser automation, that kind of stuff?
I mean, what I've been saying for a long time, maybe I'm wrong, is that if the future is that you're gonna stand in front of a computer and looking at an agent clicking on stuff, then I'll eat my computer.
And my computer is a big Lenovo. It's black. Doesn't sound good at all compared to a Mac. And if the APIs are there, y- we should use them. There, there is gonna be a long tail of stuff that don't have APIs, but as the m- the, the world is moving forward, that's, that's disappearing.
So the core API value, the, uh, in the past has really been, "Oh, this, this old '90s product doesn't have an API, so I need to use the UI to automate." I think for most of the ICP companies, uh, the companies that ICP for us, the scale-ups that are between five hundred and five thousand people, tech companies, most of the SaaS they use have APIs.
Not as an interesting question for the open web because there are stuff that you wanna do that involve websites that don't necessarily have APIs, and the current state of web integration from, which is us and OpenAI and Anthropic, I don't even know if they have web navigation, but I don't think so.
The current state of affair is really, really broken because you have what? You have basically search and headless browsing. But headless browsing, I think everybody's doing basically body.innertext- ... and, and fill that into the model, right?
There's parsers into Markdown and stuff.
We're super excited by the companies that are exploring the, the capability of, of rendering a webpage into a way that is compatible for a model, being able to maintain the selectors, so that's the-- basically the, the place where to click in the page through that process, expose the actions to the, to the model, have the model select an action in a way that is compatible with, with model, which is not a big page of a, a full DOM that is very, uh, noisy, and then being able to, uh, decompress that back to the original page and take the action.
Mm.
And that's something that is really exciting and what-- that will kind of change the, uh, the level of things that a model-- that, that agents can do on the web. That I feel exciting, but I also feel that the bulk of the useful stuff they can do within the company can be done through API.
The data can be retrieved by API. The action's gonna be taken through API.
Yeah. For, for listeners, I'll note that you're basically completely disagreeing with David on -
Exactly. Exactly
... on that.
I've seen the same-
And, you know, I mean-
... good summary
... Adept is where it is and, you know, and Dust is where it is. So Dust is still standing.
Can we just quickly comment on function calling?
Function Calling32:52
Yeah.
You mentioned you don't need the models to be that smart to actually pick the tools. Have you seen the models not be good enough, or is it just, like, you just don't wanna put the complexity in there? Like, is there any room for improvement left in function calling, or do you feel usually consistently get always the right response, the right parameters and all that?
So that's the tricky product question because if you-- if the instructions are good and precise, then you don't have an issue because it's scripted for you, and the model just look at the scripts and just follow and say, "Oh, he's probably talking about that action, and I'm gonna use it, and the parameters are kind of abused from the state of the conversation.
I'll just go with it." If you provide a very high level, kind of a Auto-GPT-esque level at-- in the instructions and provide 16 different tools to your model, yes, we're seeing the models in that state making mistakes.
Mm-hmm.
And there is, um, obviously some progress, some, some progress can be n-made on the capabilities. But the intrest-interesting part is that there is already so much work then-- and that can assist, augment, accelerate by just going with pretty simply screw two four actions, uh, agents.
What I'm excited about by starting in, like, pushing our users to create rather simple agents, is that once you have those working really well, you can create meta agents that use the agents as actions.
Yeah.
And all of a sudden, you can kind of have a hierarch-hierarchy of, of responsibility that will probably get you almost to the point of the Auto-GPT value. It required the construction of intermediary artifacts, but you, you, you're probably gonna be able to achieve, uh, something great.
I'll give you some example. We have, uh, our incidents are shared in Slack in a specific channel, or Shipt are shared in Slack. We have a weekly meeting where we have a table about incidents and, uh, Shipt stuff.
We're not writing that weekly meeting table anymore. We have an assistant that just go find the right data on Slack and create the table for us, and that assistant works perfectly.
Mm-hmm.
It's trivially simple, right? Take one week of data from that channel and just create the table. And then we have in that weekly meeting, uh, some, uh... obviously some graphs and, and, and reporting about our financials and our progress and our ARR, and we've created assistants to generate those graph directly, and those assistants works great.
By creating those assistants that cover those small parts of that weekly meeting, slowly we're getting to in a world where we'll have a weekly meeting assistants. We'll just call it. You don't need to prompt it. You don't need to say anything.
It's gonna run those different assistants and get that Notion page just ready. And by doing that, if you get there, and that's an objective for us to, us using Dust, get there, you're saving, I don't know, an hour-
Mm.
-of company time-
Yeah
... every time you run it.
Yeah. That's my pet, pet topic of NPM for agents, is like how do you build dependency graphs of agents and, like, how do you share them? Because why do I have to rebuild some of the smaller levels-
Yeah
... of what you built already?
I have a quick follow-up question-
Yeah, yeah
... on agents managing other agents. It's a topic of a lot of research both from, like, Microsoft and e-even in startups. What you've discovered best practice for, uh, let's say, like, a manager agent controlling a bunch of small agents?
That it's two-way communication. I don't know is there should be a protocol format.
To be com-completely honest, the, the state we are at right now is creating the simple agents.
I see.
So we haven't even explored yet the meta agents.
I see.
We know it's there. We know it's gonna be valuable. We know, we know it's gonna be awesome. But we're starting there because it's the simplest place to start, and it's also what the market understands. If you go to a company, random com-- SaaS B2B company, not necessarily specialized in AI, and you take the-- an operational team-
Mm-hmm
... and you tell them, "Build some tooling for yourself," they'll understand the small agents. If you tell them, "Build Auto-GPT," they'll be like, "What? Auto what?"
And I noticed that in your language, you're very much focused on non-technical users.
Yep.
You, you don't really mention API here.
Yep.
You mention instruction instead of system prompt.
Yep.
Right? That's very conscious.
Yeah, it's very conscious. It's a mark of our designer, Ed, who kind of pushed us to create a friendly product. I was knee-deep into AI when I started, obviously. And my co-founder Gabriel was, uh, was at Stripe as well.
Uh, we started a company Glazer that got acquired by Stripe ter- 15 years ago. Was at Alan, a healthcare company in, in, in Paris after that. He was a little bit, uh, less so knee-deep in AI, uh, but, uh, really focused on product.
And I didn't realize how important it is to make that technology not scary to end users. It didn't feel scary to me, but it was really s-seen by Ed, our designer, that it was f-feeling scary to the users.
And so we were very proactive and very deliberate about creating a brand that feels not too scary and creating a wording and a, a, a language, as you say, that's, that's really- Try to communicate the fact that it's, it's gonna be fine, it's gonna be easy, you're gonna make it.
Tech Stack37:31
And another big point that David had about Adept is, like, we need to build an environment for, like, the agents to act, and then if you have the environment, you can simulate what they do. How's that different when you're interacting with APIs and you're kind of touching systems that you cannot really simulate?
Like, you know, if you call the Salesforce API, you're just calling it, you know?
Yep. So, uh, I think that goes back to the DNA of the companies that are very different. Adept, I think, was a product company with a very strong research DNA, and they were still doing research. One of their goal was building a model, and that's why they raised a large amount of money to, et cetera.
We are 100% deliberately product company. We don't do research. We don't train models. We don't even run GPUs. We're using the models that exist, and we try to push the product boundary as far as possible with the existing models.
So that creates an issue indeed. So to answer your question, when you're interacting in the real world, well, you cannot simulate, so you cannot improve the models. Even interacting your-- i-improving your instructions is complicated for a builder. The hope is that you can use models to evaluate the conversations so that you can get at least feedback, and you could get quantitative information about the performance of the assistants.
But if you take actual trace of interaction of humans with those agents, it is, even for us human, extremely hard to decide whether it was a productive interaction or a really bad interaction. You don't know why the person left.
You don't know if they left happy or not. So being extremely, extremely, extremely pragmatic here, it becomes a product issue. We have to build a product that incentivizes user-- the end users to provide feedback, so that as a first step, person that is building the agent can iterate on it.
As a second step, maybe later when we start training model and post-training them, et cetera, we can optimize around that for each of those companies.
Yeah. Do you see in the future products offering kind of like a simulation environments? The same way all SaaS now kind of offers APIs to build programmatically. Like, in cybersecurity, there are a lot of companies working on building simulative environments so that then you can use agents to, like, red team.
But I haven't really seen that.
Yeah, no. Me neither. Uh, that's a super interesting question. I think it really gonna depend on how much, uh-- Because you need to simulate to generate data. You need to train data to train models.
Mm-hmm.
And the, the question is at the end is, are we gonna be training models or are we just gonna be using frontier models as they are? On that question, I don't have a strong opinion.
Mm-hmm.
It might be the case that we'll be training models because in all of those AI-first products, the model is so close to the surf- to the product surface that as you get big and you wanna really own your product, you're gonna have to own the model as well.
Owning the model doesn't mean doing the, uh, pre-training. That would be crazy. But at least having, uh, an internal post-training realignment loop makes a lot of sense. And so if we see many companies going towards that over time, then there might be incentives for the s- the, the SaaSes of the world to provide, uh, assistance in getting there.
But at the same time, there's a tension because those SaaS, they don't wanna be interacted on-
Right. Yeah, yeah, yeah. Exactly
... by, by assistants. They want it there by agents. They want, they want the human to click on the button. So that's an interesting-
Yeah, they gotta sell seats.
Yeah, exactly.
So...
Exactly. Just a quick question on models. I'm sure you've used many, probably not just OpenAI. Would you characterize some models as better than others? Do you use any open source models? What have been the trends in models over the last two years?
We've seen over the past two years kind of a, a bit of a race, uh, in between models, and at, uh, at, at times it's the OpenAI model that is, uh, the best. At times it's the Anthropic models that is the best.
Our take on that is that we are agnostic, and we let our users pick their model.
Oh, they choose?
Yeah. So when you create an assistant, you know, an agent, you can, uh, you can just say, "Oh, I'm gonna run it on GPT-4 or GPT-4 Turbo or..."
Don't you think for the non-technical user, that is actually an abstraction that you should, you should take away from them?
We have a sane default.
Yeah.
So we take, uh... We move the default to the latest model that is cool.
Mm.
And we have a sane default, and it's actually not very vi-visible in our flow to create an agent. You, you would have to go in advanced and go pick your model.
Mm-hmm.
So this is something that the technical person will, will care about, but that's something that obviously, uh, is a bit, uh, more-- too complicated for the, um-
And do you care most about function calling or instruction following or something else?
I think we care most for function calling.
Right.
Because you wanna-- There's nothing worse than a function call including incorrect parameters or being a bit off because it just, uh, drives the whole interaction off.
Yeah. So got the Berkeley function calling leaderboard.
Yeah. These days, it's funny how the comparison between GPT-4o and GPT-4 Turbo is still up in the air on function calling.
Mm.
I personally don't have proof, but I know many people, and I'm probably part of them, to think that GPT-4 Turbo is still better than GPT-4o on function calling.
Wow.
We'll see what comes out of, uh, uh, the o1, uh, class if it ever gets, uh, function calling. And Claude 3.5 Sonnet is great as well. Uh, they kind of innovated in an interesting way, which was never quite publicized, but it's that they have that kind of chain of thoughts step whenever you use a Claude model or Sonnet model with function calling.
That chain of thoughts step doesn't exist when you just interact with it-
Yeah
... just for, for answering questions. But when you use function calling, you get that step, and it, it really helps getting better function calling.
Yeah. We actually just recorded a podcast with the Berkeley team that runs that leaderboard this week. So they just released V3.
Yeah.
Uh, it was V1 like two months ago, and then they V2, V3. Turbo is on top.
Turbo is on top.
Turbo is over 4o, and then the third place is xLam from Salesforce, which is a large action model they've been trying to popularize.
Yep.
o1-mini is actually on here, I think. o1-mini is, uh, number 11.
But arguably-
Yeah.
Yeah, yeah
... o1-mini has a middle line for that, so.
It wasn't two, I mean.
Yeah.
Do you use leaderboards? Do you have your own evals? I mean, this is counterintuitive, right? Like, using the older model is better. I think most people just upgrade.
Yeah.
Yeah. What's the, what's the eval process like?
It's funny because I, I've been doing research for three years, and we have bigger stuff to cook.
Yeah.
When you're deploying in a company, one thing where we really spike is that when we manage to activate a company, we have a crazy penetration. The highest penetration we have is 88% daily active users within the entire employee of the company.
The kind of average, uh, penetration and activation we have in our current enterprise customers is something like more like 60 to 70% weekly active. So we basically have the entire company interacting with us. And when you're there, there is so many stuff that matters most than getting evals-
Damn
... getting their best model. Because there is so many s- places where you can create products or do stuff that will give you the 80% with the work you do, whereas deciding if it's GPT-4 or GPT-4 Turbo or et cetera, you know, it'll just give you the, the 5% improvement.
Mm-hmm. Yeah, yeah, yeah.
But the reality is that you wanna focus on the places where you can really change the direction or change the, uh, the interaction, uh, uh, more drastically.
Yeah.
But that's something that we'll have to do eventually because-
Yeah
... we still want to be service people.
It, it's funny 'cause i- in some ways the, the model labs are competing for you, right? You don't have to do any effort. You just switch model, and then it, it'll, it'll grow. But what are you rate limited by?
Is it additional sources? It's not models, right? You're not really rate limited by quality of model.
Right now we, right now we are limited by, yes, the, uh, the infrastructure part, which is a ability of... I mean, ability to connect easily-
Mm-hmm
... for our users to all the data they need to do-
Okay
... the, the job they wanna do. Um-
Because you maintain all your own stuff.
Exactly.
You know, there, there are companies out there that are starting to provide integrations as a service, right? I used to work at an integrations company.
Yeah, yeah, yeah. No, no, no. It's just that there is some intricacies about how you chunk stuff and how you process information from one platform to the other. If you look at the, uh, end of the spectrum, you could think of a, you could say, "Oh, I'm gonna support Airbyte, and Airbyte can, can source the-"
I used to work at Airbyte, yeah.
Oh, really?
Yeah.
Makes sense.
They have French founders as well.
French. Yeah, I was gonna say.
I know, I know Jean very well. Uh, I'm seeing him today. And the reality is that if you look at Notion, Airbyte does the job of taking Notion and putting it into in a structured way, but it's a way that is not really usable to actually make it available to, to models in a useful way-
Yeah
... because you get all the blocks details, et cetera, which is useful for many use cases.
But it's also meant for data scientists and not-
Exactly
... not for AI. Yeah.
Uh, the reality of Notion is that sometime you have a, sometime you have... So when you have a page, there's a lot of structure in it, and you wanna ch- you, you wanna capture the structure and chunk the information in a way that respects that structure.
In Notion, you have databases. Sometimes those databases are real tabular data. Sometimes those databases are full of text. You want to get the distinction and understand that this database should be considered like text information, whereas this other one is actually quantitative information.
And to really get a very high quality interaction with that piece of information, I haven't find a solution that will-
Yeah
... work without us owning the connection end-to-end.
That's why I don't invest in these... There's Composio. There's, um, All Hands from, from, uh, Graham Neubig. There's all these other companies that are tr- Like, we will do the integrations for you. You just-- We have the open source community.
We'll, we'll do off the shelf. But then you are so specific in your needs-
Yeah
... that you want to own it.
Yeah, exactly.
You can talk to Michel about that. You know, he wants to put the AI in Airbyte, you know?
Yeah. I will. I will.
Cool.
What are we missing? You know, what are, like, the things that are, like, sneakily hard that you're tackling that maybe pe- people don't even realize they're, like, really hard?
The real hard parts, as we, we kind of touched base, uh, throughout the conversation, is really building the infra that works for those agents because it's a tenuous work. It's, uh, an evergreen piece of work because you always have an extra integration that will be useful to a non-negligible set of your users.
I'm super excited about is that there is so many interactions that shouldn't be conversational interactions, and that could be very useful. Basically, know that we have the firehose of information of those companies, and there's not gonna be that many companies that capture the firehose of information.
When you have the firehose of information-
Yeah
... you can do ton of stuff with those, uh, with models that are just not just accelerating people, but giving them superhuman capability, even with the current model capability, because you can just sift through much more information. An example is documentation repair.
If I have the firehose of Slack messages and new Notion pages, if somebody says, "I own that page, I wanna be updated when there is piece of information that should update that page," this is not possible. You get an email, receiving email saying, "Oh, look at that Slack message.
It says the opposite of what you have in that paragraph. Maybe you wanna update or just ping that person." I think the, uh, there is a lot to be explored on the product layer in terms of what it means to interact productively with those models, and that's, that's, uh, some- a problem that is extremely hard and extremely exciting.
One thing you keep mentioning about infra work, obviously, Dust is building that infra and, and serving that in a very consumer-friendly way. You always talk about infra being additional sources, additional connectors. That is very important, but I'm also interested in, like, the vertical infra.
Orchestration47:57
There is an orchestrator underlying all these things, right? Where you're doing asynchronous work. For example, just the simplest one is a cron job, is do you just schedule things.
Mm-hmm.
But also there, for if this and that, you have to wait for something to be, to executed and, and proceed to the next task. I used to work on an orchestrator as well, Temporal. Um-
We use Temporal.
Oh, you use Temporal?
Yeah.
Oh, how, how was the experience? I need the MPS.
We're doing a self-discovery call now?
No, no, but you can also complain to me- ... 'cause I don't work there anymore, and, uh, you know.
No, we love Temporal. There's, uh, there's some, uh, edges that are a bit, that are a bit rough, uh, surprisingly rough, and you would say, "Why? W- Why is it so complicated?"
It's versioning. It's always versioning.
Yeah, stuff like that. Uh, but we really love it, and we use it for, for exactly what you said, like, uh, managing the, uh, the entire set of, of stuff that needs to happen so that in semi-real time we get all the updates from, uh, Slack or Notion or GitHub into the system.
And whenever we see that piece of information goes through, maybe trigger, uh, workflows because to run agents, uh, because they need to provide alerts to users and stuff like that. And Temporal is great. Love it.
But you haven't evaluated others. You don't wanna build your own. You, you, you're happy with-
Oh, no, we're not, I know, we're not in the business of, of replacing-
Building your own orchestrators
... Temporal. And Temporal is so-- I mean, it is, or any other competitive product, they're, they're very general. If it's there. There's an interesting theory about buy versus build. I think, uh, in that case, when you're a high-gross company, your buy-build trade-off is very much on the side of buy, because if you have the capability, you're just gonna be saving time.
You can focus on your core competency, et cetera. And it's funny because we're seeing, uh, we're starting to see the post high-gross company, post S-curve company going back on that trade-off, interestingly. So that's the Klarna news about removing, uh, Zendesk and Salesforce.
Do you believe that, by the way?
Yeah.
Do you believe-
I just heard a podcast with them.
Oh, yeah?
It's true. Yeah, yeah.
No, no, I, I know. They, of course they say it's true, but, like, also how well is it gonna go?
Oh.
So I'm not, I'm not talking about s- uh, deflecting the customer traffic. I'm talking about building AI on top of Salesforce and Zendesk, basically, if I understand correctly. And all of a sudden, your product surface become much, uh, s- smaller because you're interacting with an AI system that will take some actions.
And so all of the sudden you don't need the product layer anymore, and you realize that, oh, those things are just database that I pay 100 time- ... the price, right? Because you're post-S-curve company and you, you have tech capabilities, you are incentivized to reduce your costs, and you have the capability to do so, and then it makes sense to just scratch the SaaS away.
So it's interesting that we might see kind of a, a bad time for SaaS in post-hypergrowth tech companies. So h- it's still a big market, but it's not that big because if you're not a tech company, you, you don't have the capabilities to reduce those costs.
If you're a high-growth company, always gonna be buying because you go faster with that. But there's interesting new space, uh, new category of companies that, uh, might remove some SaaS.
Yeah. Alessio's firm has a interesting thesis on the future of SaaS and AI.
Predictions50:58
Yeah. Uh, service is a software, we call it. Especially like... Well, the most extreme is, like, why is there any software at all?
Yeah.
You know? Ideally, it's all a labor interface where you're asking somebody to do something for you, whether that's a person-
Asking agents
... an AI agent or whatnot.
Yeah. Yeah. That's interesting.
I have to ask, are you paying for Temporal Cloud or are you, are you h- self-hosting?
Oh, no, no. We're paying. We're paying.
Oh, okay. Interesting. One, one paying user found.
We're paying way too much. It's, it's crazy expensive, but-
That's-
... makes us, makes us, makes, makes, makes-
That's why as a shareholder, I like to hear that, you know?
Makes us go faster, so we're happy to pay.
Other things in the infra stack. I just want a list for-
Yeah
... other founders to think about. Ops, API gateway, evals, you know, anything interesting there that you build or, or, or buy?
I mean, there's always an in- an interesting question. We've been building a lot around the, uh, interface between models and s- because Dust, the, the Dust, the original version was an orchestration platform, and we basically provide a unified interface to every model providers.
That's what I call gateway.
That we add because Dust was that, and so we continued building upon, and we own it. But that's an interesting question, was, and you, you wanna build that or buy it?
Yeah. Like, I would say LiteLLM is the c-
Yep
... current open source consensus.
Exactly.
Yeah.
Yep. There's an interesting question there.
Ops, Datadog, just tracking.
Oh, yeah. So Datadog is an obvious... What are the mistakes that I regret? Uh. I started as pure JavaScript, not TypeScript, and I think you wanna, you wanna... If, if you, if you're wondering, "Oh, I wanna go fast, I'll do a little bit of JavaScript," no, don't.
Just start with TypeScript-
I see
... your life will be better.
Okay.
Uh, that is important.
So interesting. You're, you are a, you know, ML, a research engineer that came out of OpenAI that bet on TypeScript.
Well, the, the reality is that if you're building a product, you're gonna be doing a lot of-
A lot of JavaScript
... JavaScript, right?
Yeah.
And, uh, Next, we're using Next as an example. It's a great, it's a great platform. And our internal service is actually not built in Python either. It's built in Rust.
That's another fascinating choice. The Next.js story is interesting 'cause Next.js is obviously the king of the world in JavaScript land. But recently, ChatGPT just rewrote from, uh, Next.js to Remix.
Mm-hmm.
Uh, we are gonna be having them on to talk about the big rewrite. That is, like, the biggest news in front-end world-
Yeah
... in a while.
All right. Just to wrap, in 2023, you predicted the first billion-dollar company with just one person running it.
Yeah.
And you said that's basically like a sign of AGI-
Yeah
... once we get there.
Yeah.
Uh, and you said it'd already been started.
Yep.
Any 2024 updates on the take?
That quote was, uh, probably independently invented it, but, uh, Sam Altman stole it from me, uh- ... eventually. But anyway, that's, that's the... It's a good quote, so... I hypothesized it was maybe already being started, but if, if it's a uni-person company, it would probably grow really fast, and so we should probably see it already.
I guess we're gonna have to wait for it a little bit, and I think it's because the Dust of the world don't exist, and so you don't have that thing that lets you run those, j- just do anything with models.
But one thing that is exciting is maybe that we're gonna be able to scale a team much further than before. Our generation of company might be the first, uh, billion-dollar companies with engineering teams of 20 people. That would be so exciting as well.
That would be so great, you know? You don't have the management hurdle. You're just 20 focused people with, uh, a lot of assistance from machines to achieve your job. That would be great.
Yeah.
And that, that I believe in, uh, a bit more.
Yeah. I've written a post called Maximum Enterprise Utilization, kinda like you have MFU for GPUs. But it's basically like so many people are focused on, oh, it's gonna like displace jobs and whatnot. But I'm like, there's so much work that people don't do because they don't have the people, and maybe the question is that you just don't scale to that size-
Yeah
... you know, to begin with, and maybe everybody will use Dust, and, uh, Dust is only gonna be 20 people, and then- ... people using Dust will be two people.
Semi-hot take is I, I actually know what vertical they will be in. They'll be content creators and podcasters. Um-
There's already two of us- ... so we're at max.
Max capacity. Most people would regard Jimmy Donaldson, like MrBeast, as a billionaire, but his team is he's, he's got about like 200 people, so he's not a single-person company. The closer one actually is Joe Rogan, where he basically just has like a guy.
Hey, Jamie.
Yeah, exactly.
Put it on the screen.
Exactly.
Jamie, pull it up.
Uh, but Joe, I don't think-- He, he sold his future for $250 million to Spotify, so he's not gonna hit that billionaire status. The non-consensus one, it would be the Hawk Tuah girl who just... Anyway, but like, um, you want creators who are empowered by-
Yeah
... uh, a, a bunch of agents, Dust agents, uh, to do all this stuff because then ultimately it's just the brand, the curation.
That has the value.
Uh, what is the role of the human then? Like what, what is that one person supposed to do if you have all these agents?
That's a good question. Uh, it's, uh... I mean, I think the, uh, it was, uh, um- I think it was Pinterest or Dropbox founder at the time was, uh, when you're CEO, you mostly have an editorial position. You're, you're here to say-
Yes
... yes and no-
Yes
... to things you are supposed to do.
Ah, okay. So I, I make a daily AI newsletter-
Yep
... where I just, I, I-- it's ninety-nine percent AI generated-
Yeah
... but I serve the role as the editor. Like, I write commentary. I choose between four options.
You decide what goes in and goes out.
Yeah, yeah.
And ultimately, as you said, you build up your brand through those many decisions. Uh, and so-
So you should, you should pursue creators. Yeah.
Yeah, so-- And, uh, you've, you've made a-- I think you've made a-- you have an upcoming podcast with NotebookLM, which has, uh, been doing, uh-
Oh, yeah
... crazy stuff there.
Yeah, that is exciting. They were just in here yesterday. I'll tell you one agent that we need. If you want to pursue the creator market, the one agent that we haven't paid for is our video editor agent.
Yeah.
So if you want, you need to, you know, wrap FFmpeg in a, in a GPT and it- ... boom, that's-
Awesome. Well, the, this was great. Anything we missed? Any final kinda like call to action? Hiring is like, obviously people should buy the product and pay you, but-
Closing56:28
No, obviously. And, uh, no, I think we, we didn't dive into the, uh, the vertical versus horizontal-
Oh, yeah
... approach to, uh, to AI agents.
Quick take on that. Yeah.
We mentioned a few things. We spike at penetration, and that's just awesome because we created a tool that, uh, the entire company has and, and use, so we create a ton of value, but it makes our go-to-market much harder.
Vertical solutions have a go-to-market that is much easier because they're like, "Oh, I'm gonna solve the"-
One thing
... "uh, lawyer stuff."
Yeah.
Uh, but the potential within the company after that is, uh, limited. So there's really a nice tension there. We, uh, we're true believers of the, uh, horizontal, horizontal approach, and, uh, we'll see how that plays out. But I think it's a, it's an interesting thing to think about when, as a founder or as a technical person working with agents, what do you wanna solve?
Do you wanna solve something general, or do you wanna solve something specific? And it has a lot of impact on, on eventually what type of company you're gonna build.
Yeah. I'll provide you the-- my response on that. So I've gone the other way. I've gone products over platform.
Yep.
And it's basically your sense on the products drives your platform development. In other words, like, if you're, you're trying to be as many things to as many people as possible. We're just trying to be one thing. We build our brand in one specific niche.
Yep.
And in future, if we want to choose to spin off platforms for other things, we can because we have that, that brand. So for example, Perplexity, we went for products in search, right?
Yep.
But then we al-also have Perplexity Labs that, like-
Yep
... here's the info that we use for search and whatever.
The counterargument to that is that, uh, you always k- have lateral movement within companies, but that's... If you're Zendesk, you're not gonna be-
Zendesk-
... serving engineers
... web services. There are a few. You know, there's success stories on both sides.
Yeah, yeah.
But there's Amazon and Amazon Web Services.
It's-
But it's very few
... and sorry, by platform, I don't really mean the platform as the, uh, platform, platform.
Yeah, yeah.
Yeah.
I mean like the, uh, the product that, uh, is useful to everybody within the company. And our take on that is that there is so many operations within a company. Some of them have been extremely rationalized by the market, like salespeople, like support.
It's been extremely rationalized, and so you can probably product-- uh, create very powerful vertical product around that. But there is so many operations that make up a company that are specific to the company that you need a product to help people a- get assisted on those operations, and that's kind of the bet we have.
Excellent. Awesome. Man, thanks again for the time.
Thank you very much for having me. It was so much fun.
Yeah, great discussion. Thank you.






