Genesis0:00
Hey everyone, welcome to the Latent Space Podcast. This is Alessio, partner and CTO in residence at Decibel Partners, and I'm joined by my co-host, Swyx, founder of Smol AI.
Hey, and today we're in the studio with my good friend and former landlord, uh, Will Bryk.
Yeah.
Roommate. How you doing?
Uh, Will, you're now CEO, co-founder of Exa.ai, used to be Metaphor Systems. What's your background, your story?
Yeah, sure. So yeah, I'm CEO of Exa. I've been doing it for three years. Um, I guess I've always been interested in search, whether I knew it or not. Like, since I was a kid, I've always been interested in, like, high-quality information, and, like, you know, even in high school wanted to improve the way we get information from news, and then in college built a, a mini search engine, and then with Exa, like, you know, it's kinda like fulfilling the dream of actually being able to solve all the information needs I wanted as a kid.
Um, yeah, I guess I would say my entire life has kind of been uh, rotating around this problem, which is pretty cool.
Yeah. What'd you enter YC with?
We e-entered YC with, uh, "We are better than Google," like Google 2.0.
Yeah. What makes you say that? Like, that's so audacious to, to come out the box with.
Yeah. 'Cause, okay, so you might have to remember the time. This was summer 2021, and, uh, GPT-3 had come out. Like, here was this magical thing that you could talk to, you could enter a whole paragraph, and it understands what you mean, understands the subtlety of your language.
And then there was Google, uh, which felt like it hadn't changed in a decade, uh, because it really hadn't, and it like, you would give it a simple query like, I don't know, uh, "shirts without stripes," and it would give you a bunch of results for the shirts with stripes.
And so, like, Google could barely understand you, but GPT-3 could, and the theory was, what if you could make a search engine that actually understood you? What if you could apply the insights from LLMs to a search engine?
And it's really been the same idea ever since, and we're actually a lot closer now, uh, to doing that. Yeah.
Did you have any trouble making people believe obviously there's a Sam Altman YC overlap? Was YC pretty AI-forward, even 2021, or...?
It's nothing like it is today. But, um, uh, there were a few AI companies, but, uh, we were definitely, like, bold, and I think people, VCs generally like boldness, and we definitely had some AI background, and we had a working demo, so there was evidence that we could build something that was gonna work.
But yeah, I think, like, the fundamentals were there. I think people at the time were talking about how, you know, Google was failing in a lot of ways, and so there was a bit of conversation about it, but AI was not a big, big thing at the time.
Yeah.
Yeah.
Before we jump into Exa, any fun background stories? I know you intern at SpaceX. Any Elon, uh, stories? I know you were at Zoox as well, you know, kinda like robotics at Harvard. Any stuff that you saw early that you thought was gonna get solved that maybe it's not solved today?
Oh, yeah. I mean, lots of things like that. Like, uh, I never really learned how to drive because I believed Elon that self-driving cars would happen. It did happen, and I take them every night to get home, but it took, like, ten more years than I thought.
Do you still not know how to drive?
I know how to drive now.
Okay, yeah.
I learned it, like, two years ago.
That would've been great to, like, just nope.
Yeah, yeah.
You know?
Um, I was obsessed with Elon. Yeah. I mean, I worked at SpaceX 'cause I really just wanted to work at one of his companies, and I remember they had a rule, like, interns cannot touch Elon, and, um, that rule actually influenced my actions.
Is it-
But can Elon touch interns? Ooh.
Like ph-like physically or like talk?
Physically, physically.
Okay.
Yeah, yeah.
Interesting.
I-- He's changed a lot, but, um- ... I mean, his companies are amazing. Um, so-
What if you beat him at Diablo II... Diablo IV, you know, like...
Oh, maybe. Yeah.
Yeah. Wanna jump into Exa. I know there's a lot of backstory. It used to be called Metaphor Systems, so, um, and it-- you've always been kinda like a prominent company-
Metaphor Model3:35
Mm-hmm
... maybe at least RAI circles in, uh-
Yeah
... NSF.
I'm actually curious how Metaphor got its initial aura.
Mm-hmm.
You launched with, like, very little.
We launched very little?
Like, there was-
Yeah, yeah
... there was this, like, big splash image of, like, this is Aurora or something.
Yeah, yeah.
Right? And then I was like, "Okay, what this thing?" Like, the vibes are good, but I don't know what it is. And I think, I think it was much more sort of maybe consumer-facing than what you are today.
Mm-hmm.
Would you say that's true?
No. It's always been about building a better search algorithm, like search. Like, just, like, the, the vision has always been perfect search, and if you do that, uh, we will figure out the downstream use cases later. It started on this fundamental belief that you could have perfect search over the web, and we can talk about what that means.
And, like, the initial thing we released was really just, like, our first search engine, like, trying to get it out there. Kinda like, you know, when OpenAI released, uh, ChatGPT. Like, they didn't... I don't know how, how much of a game plan they had.
They kinda just wanted to get something out there and-
Smoky research preview.
Yeah, exactly. And it kinda morphed from a research company to a product company at that point. And I think similarly for us, like, we were a research-- We started as a research endeavor with a f- you know, clear eyes that, like, if we succeed, it will be a massive business to make out of it, and that's kind of basically what happened.
I think there are actually a lot of parallels to, of, w- between Exa and OpenAI. I often say we're the OpenAI of search, um, because we're a research company. We're a research startup that does, like, fundamental research into, uh, making, like, AGI for search in a, in a way, uh, and then we have all these, like, uh, business products that come out of that.
Interesting. I wanna ask a little bit more about Metaphor side, and then we can go full-
Yeah, sure
... full Exa. When I first met you, which was, uh, really funny 'cause, like, literally, I stayed in your house-
Yeah
... in a very historic, uh, Hayes, Hayes Valley place. You said you were building sort of like link prediction foundation model, and I think there's still a lot of foundation model work in, within Exa today, but what does that even mean?
I cannot be the only person confused by that-
Oh, yeah.
... because, like, there's a limited vocabulary of tokens. You, you're telling me, like, the tokens are the links or, you know, like, it's not, it's not clear.
Yeah. Uh, what we meant by link prediction is that you are literally predicting the, like, given some text, you are predicting the links that follow.
Yes.
That refers to, like... Uh, it, it's how we describe the training procedure, which is that we find links on the web, uh, we take the text surrounding the link, and then we predict- Which link follows. So you're like, uh, you know, similar to transformers where, uh, you're trying to predict the next token.
Here, you're trying to predict the next link. And so you kinda like hide the link from the transformer. So if someone writes-- You know, imagine some article where someone says, "Hey, check out this really cool aerospace startup," and they, they say spacex.com afterwards.
Uh, we hide the spacex.com and ask the model, like, what link came next.
Yeah.
And by doing that many, many times, you know, billions of times, you could actually build a search engine out of that because then, uh, at, at query time, at search time, uh, you type in a, a s- query that's like really cool aerospace startup, and the model will then try to predict what are the most likely links.
So there's a lot of analogs to transformers, but, like, to actually make this work, it does re-require like a different architecture-
Yeah
... than trans-- But it's transformer-inspired, yeah.
What's the design decision between doing that versus extracting the link and the description and then embedding the description and then using-
Typical rank.
Um, yeah. Why do you need to predict the URL versus, like, just describing-- Because you're kinda doing a similar thing in a way, right? It's kinda like based on this description, what's, like, the closest link for it? So one thing is, like, predicting the link.
The other approach is, like, I extract the link and the description, and then based on the query, I search the closest description to it more-
Yeah, that, that-- By the way, that is, that is-- The link refers here to a document. It's not-- I think one confusing thing is it's not-- you're not actually predicting the-
The URL
... the URL itself.
Yeah. Yeah.
Uh, that would require, like, the, the system to have memorized URLs. You're actually, like, getting the actual document. A, a more accurate name could be document prediction.
I see.
This was the initial, like, base model that Exo was trained on, but we've moved beyond that. Similar to, like, how, you know, uh, to train a really good, like, language model, you might start with this, like, self-supervised objective of predicting the next token and then y- uh, just from random stuff on the web.
But then you, you want to, uh, add a bunch of, like, synthetic data and, like, supervised fine-tuning, um, stuff like that to make it really, like, controllable and robust.
Yeah.
Yeah.
Yeah, we just have flow from Lindy and, uh-
Yeah
... their Lindy started to, like, hallucinate Rickrolling YouTube links-
Got it
... instead of, like, a support guide, so.
Oh, interesting.
Yeah.
Exa Rebrand7:58
Yeah.
So round about January, you announced your Series A-
Yeah
... uh, and renamed to Exo. I didn't like the name at the, at the initial, but it's grown on me. I, I liked Metaphor, but, uh, apparently people couldn't spell Metaphor. What would you say is-- are the com-- major components of Exo today, right?
Mm-hmm.
Like, I, I feel like it used to be very model-heavy. Then at the AI engineer conference, Shreyas gave a really good talk on the vector database-
Mm-hmm
... that you guys have. What are the other major moving parts of Exo?
Okay. So Exo overall is a search engine.
Yeah.
And we're trying to make it like a perfect search engine. And to do that, you have to build lots of... And we're doing it from scratch, right?
Yeah.
So to do that, you have to build lots of different-
The crawler
... subsystems. Yeah. You have to crawl a bunch of the web. First of all, you have to find the URLs to crawl. Uh, it, it's connected to the crawler, but yeah, you f-find URLs, you crawl those URLs, then you have to process them with some, you know, it could be an embedding model, it could be something more complex.
But you need to take, you know-- Or like, you know, in the past, it was like a, a keyword-inverted index. Like, you would process all these documents you gather into some processed index, and then you have to serve that at high throughput, at low latency.
And so, and that's like the vector database. And so it's like the crawling system, the AI processing system, and then the serving system. Those are all, like, you know, teams of, like, hundreds, maybe thousands of people at Google.
Um, but for us, it's like one or two people typically. But yeah.
Can you explain the meaning of, uh, Exo? Just the-
Ooh
... story, ten to the sixteenth, uh, and, and all that.
I think ten to the eighteenth.
Eighteen. Eighteen.
Yeah. Yeah, yeah, sure. So Exo means ten to the eighteenth, which, which is in stark contrast to Google, which is ten to the hundredth. Uh, we actually have these, like, awesome shirts that are like ten to the eighteenth is greater than ten to the hundredth.
Yeah, it's great.
And it's great because it's provocative. It's like every engineer in Silicon Valley is like, "What? No, it's not true." Um, like, yeah. And, uh, and then you ask them, "Okay, what does it actually mean?" And, like, the creative ones will, will recognize it.
But yeah, I mean, ten to the eighteenth is better than ten to the hundredth when it comes to search because with search, you want, like, the actual list of, of things that match what you're asking for. You don't want, like, the whole web.
You wanna basically with search filter the wh-- like, everything that humanity has ever created to exactly what you want. And so the idea is, like, smaller is better there. You want, like, the best ten to the eighteenth and not the ten to the hundredth.
Um, like, one way to say this is, like, you know how Google often says at the top, uh, like, you know, th-thirty million results found?
Mm-hmm.
And it's, like, crazy 'cause you're looking for, like-
The first two pages
... startups in San Francisco that work-
Yeah
... on hardware or something, and, like, there are not thirty million results like that. What you want is, like, three hundred and twenty-five results found, and those are all the results. That's what you really want with search, and that's, that's our vision.
It's like it just gives you perfectly what you ask for.
We're recording this ahead of your launch.
Product & Compute10:25
Mm-hmm.
Uh, we haven't released-- We haven't figured out the, the, the name of the launch yet. But what is the product that you're launching, I guess, now that we're coinciding this podcast with?
Yeah. So we basically developed the next version of Exo, which is the ability to get a near-perfect list of results of whatever you want. And what that means is you could make a complex query now to Exo, for example, startups working on hardware in SF, and then just get a huge list of all the things that match.
Mm-hmm.
And, you know, our goal is if there are three hundred and twenty-five startups that match that, we find you all of them. And this is just a ve-- like, this is just, like, a new experience that's never existed before.
It's really... Like, I don't know how you would go about that right now with current tools. And you could apply this same type of, like, technology to anything. Like, let's say you want, uh, you wanna find all the blog posts that talk about Swix's podcast, um, that have come out in the past year.
That is thirty million results.
Yeah, right.
But that, I mean, that would-- I'm sure that would be extremely useful to you guys, and, like, I don't really know how you would get that full comprehensive list.
I just, like, how do you... Well, there's so many questions with regards to how do you know it's complete-
Mm-hmm
... right? 'Cause you're saying there's only thirty million or three-three hundred and twenty-five, whatever. And then how do you do the semantic understanding that it might take, right?
Mm-hmm.
So working on hardware, like, I might not use the words hardware. I might use the words robotics. I might use the words wearables. I might use, like, whatever.
Yes.
So yeah, just tell us more.
Yeah, yeah, sure. So one aspect of this, it's a little subjective.
Yeah.
Um, so, like, certainly- P- providing, yeah, at, at some point we'll provide parameters to the user to, like, you know, some sort of threshold-
Mm
... to, like, uh, gauge like, okay, like, this is a cutoff. Like, this is actually not what I mean. Because it-- sometimes it's subjective and there needs to be a feedback loop. Like, oh, like it might give you, like, a few examples, and you say, like-
Like, "Is this what you want?"
Yeah, exactly.
Yeah.
And so, like, you're, you're kind of like creating a classifier on the fly. But, like, that's ultimately how you solve that problem. So the subject-- there's a subjectivity problem, and then there's the comprehensiveness problem. Those are two different problems.
So to solve the comprehensiveness problem, what you basically have to do is you have to put more compute into the query, into the search until you get the full comprehensiveness.
Yeah.
And I think there's an interesting point here, which is that not all queries are made equal. Some queries, just like this blog post one, might require scanning, like scavenging, like throughout the whole web-
Yeah
... uh, in a way that just, just simply requires more compute. You know, at some point, there's some amount of compute where you will just be comprehensive. You could imagine, for example, running GPT-4 over the entire web and saying like, "Is this a blog post about Swiss pod- Swix's podcast?"
Like, "Is this blog post about Swix's podcast?" And then that would work, right? It would take, you know, a year, maybe cost like a million dollars, but or many more, but, um, it would work. Uh, the point is that, like, given sufficient compute, you can solve the query.
And so it's really a question of, like, how comprehensive do you want it given your compute budget? I think it's very similar to o1, by the way, and one way of thinking about what we built is like o1 for search, uh, because o1 is all about, like, you know, some qu- some questions require more compute than others, and we'll put as much compute into the question as we need to solve it.
So similarly with our search, we will put as much compute into the query in order to get comprehensiveness.
Yeah. Does that mean you have, like, some kind of compute budget that I can specify?
Yes. Yes.
Okay. And, like, what are the upper and lower bounds?
Yeah. This is something we're still figuring out. I think, like, like everyone, there's a new paradigm of, like, variable compute products.
Yeah.
How do you specify the amount of compute? Like, what happens when you run out? Do you just like, ah, do you-- can you, like, keep going with it? Like, do you just put in more credits to get more?
Mm-hmm.
Um, for some, like this can get complex at like the really large compute queries and, like, one thing we do is we give you a preview of what you're gonna get, and then you could then spin up like a much larger job, uh, to get like way more results.
But yes, there is some compute limit, um, at, at least right now.
Yeah.
People think of searches as like, oh, it takes five hundred milliseconds because we've been conditioned, uh, to have search that takes five hundred milliseconds by like search engines like Google, right? No matter how complex your query to Google, it will take like, you know, roughly four hundred milliseconds.
But what if searches can take like a minute or ten minutes or a whole day? What can you then do? And you can do very powerful things.
Hmm.
Um, you know, you can imagine, you know, writing a search, going to get a cup of coffee, coming back, and you have a perfect list. Like, that's okay for a lot of use cases.
B2B Knowledge14:43
Yeah.
I mean, the use case closest to me is venture capital, right? So, uh- Eight, eight... No, I mean, eight years ago, I, I built one of the first like data-driven sourcing platforms, so we would look at GitHub, Twitter, Product Hunt, all these things, look at interesting things, evaluate them.
If you think about some jobs that people have, it's like literally just make a list.
Mm-hmm.
If you're like a analyst-
Yeah
... at a venture firm, your job is to make a list of interesting companies, and then you reach out to them. How do you think about being infrastructure versus like a product? You could say, "Hey, this is like a product to find companies.
This is a product to find things," versus like offering more as a blank canvas that people can build on top of.
Oh, right. Right. Uh, we are a search infrastructure company, so we want people to build, uh, on top of us, uh, build amazing products on top of us. But with this one, we try to build something that makes it really easy for users to just log in, put a few, you know, put some credits in, and just get like amazing results right away and not have to wait to build some API integration.
So we're kinda doing both. Uh, we, we want, we want people to integrate this into all their applications. At the same time, we wanna just make it real easy to use.
Mm-hmm.
Very similar again to OpenAI. Like they'll have-- they have an API, but they also have like a ChatGPT interface, so that you could... It's really easy to use, but you could also build it in your applications.
Yeah. I'm still trying to wrap my head around a lot of the implications. So many businesses run on, like, information arbitrage, you know? Like I know this thing that you don't, especially in investment and financial services. So yeah.
Now all of a sudden you have these tools where like, oh, actually everybody can get the same information at the same time, the same quality level as an API call. You know? It just kinda changes a lot of things.
Yeah. I think, I think what we're grappling with here, what, what, what you're just thinking about is like, what is the world like if knowledge is kind of solved? If like any knowledge request you want is just like right there on your computer.
It's kinda different from when intelligence is solved. There's like a, like a-- I've written before about like a difference between intelligence-
Super intelligence-
Yeah
... super knowledge.
Yeah, like I think that the, the distinction between intelligence and knowledge is actually a pretty good one. They're definitely connected and related in all sorts of ways, but there is a distinction. You could have a world, and we are gonna have this world, where you have like GPT-5 level systems and beyond that could like answer any complex request.
Um, unless it requires some like-- if you say like, uh, you know, give me a list of all the PhDs in New York City who, I don't know, have thought about search before, and even though this, this super intelligence is gonna be like, "I can't find it on Google," right?
Which is kinda crazy. Like we're literally gonna have like super intelligences that are using Google. And so if Google can't find the information, there's nothing they can do. They can't find it. So they-- But if you also have a super knowledge system where it's like, you know, I'm, I'm calling this term super knowledge, where you can just get whatever knowledge you want, then you can pair it with a super intelligence system, and then the super intelligence can, will never be blocked by lack of knowledge.
Yeah. You told me this, uh, uh, when we had lunch. I forget how it came up, but we were talking about AGI and whatnot, and you were like, "Even AGI is gonna need search."
Yeah.
So.
Yeah, right.
Yeah. Um, so we're actually referencing a blog post that you wrote, Super Intelligence and Super Knowledge, uh, and so I would refer people to that. And this is actually a discussion we've had on the podcast a couple of times.
Um, there's so much of model weights that are just memorizing facts. Some of tho- some of those might be outdated, some of them are incomplete or not, y-yeah. So like you just need search and tool use. So I, I do wonder, like is there a maximum language model size that will be the intelligence layer, and then the rest is just search?
Right.
Like a, maybe we should just always use search, and then the sort of workhor- workhorse model is just like an, like, like, like 1B or a 3B parameter model that just drives everything.
Yes, I believe this is a much more optimal system to have a smaller LLM that's really just like an intelligence module, and it makes a call to a search tool.
Mm-hmm.
That's way more efficient 'cause if-- okay, I mean, the, the opposite of that would be, like, the LLM is so big that it can memorize the whole web. That would be, like, way bit... You know, I, it's not practical at all, and it's not possible to train that, at least right now.
And Karpathy has actually written about this, how, like he c- he could-
Yeah
... see models moving more and more towards, like, intelligence modules-
Yeah
... using various tools.
Yeah. So for listeners, that's the-- that was him on the No Priors podcast. And for us, we talked about this at the-- on the Shinyu and Harrison Chase podcast. I-I'm doing search in my head.
That's pretty cool. Right. Dude, do you know-
I told you, there's thirty million results
... do you know the self Exa.ai-
I forgot about our Neural Link-Neural Link integration
... you can get self-hosted, self-hosted Exa.
Yeah. Um, yeah. No, I do see that that is a much more, much more efficient world.
Yeah.
I mean, you could also have GB4-level systems calling search, but it's just because of the cost of inference, it's just better to have a very efficient search tool and a very efficient LLM, and they're built for different things.
Search Architecture19:00
Yeah. I'm just kinda curious, like, it, it is still something so audacious that I don't wanna elide, which is y-you're building a search engine.
Mm-hmm.
Where do you start? How do you...
Yeah.
Like, are there any reference papers or implementations that would really influence your thinking? Anything like that because I, I don't even know where to start- ... apart from just crawl a bunch of shit, but there's gotta be more insight than that.
I mean, yeah. The- ... there's more insight, but I'm always surprised by, like, if you have a group of people who are really focused on solving a problem-
Uh-huh
... um, with the tools today, like there's so m-- and, and in software, like, there are all sorts of creative solutions that just haven't been thought of before.
Uh-huh.
Um, particularly in the information retrieval field-
Yeah
... I think a lot of the techniques are just very old, frankly. Like, I know how Google and Bing work, and they're just not using new methods. There are all sorts of reasons for that. Like, one, like, Google has to be comprehensive over the web, so they're-- and they have to return in four hundred milliseconds.
And those two things combined means they are kind of limit-- and it can't cost too much. They're kind of limited in, uh, what kinds of algorithms they could even deploy at scale. So they end up using, like, a limited keyword-based algorithm.
Also, like, Google was built in a time where, like in, you know, nineteen ninety-eight, where we didn't have LLMs, we didn't have embeddings, and so they never thought to build those things. And so now they have this, like, gigantic system that is built on old technology.
And so a lot of the information retrieval field we found just, like, thinks in terms of that framework.
Yeah.
Whereas we came in as, like, newcomers just thinking, like, "Okay, there-- here's GB3. It's magical. Obviously, we're gonna build search that is using that technology," and we never even thought about using keywords, really ever. Like, uh, like we-we're neural all the way.
We're building an end-to-end neural search engine. And just that whole framing just makes us ask different questions, like, pursue different lines of work, and there's just a lot of low-hanging fruit because no one else is thinking about it.
We're just on the frontier of neural search. We just are, um, for, for, at web scale, um, because there's just not a lot of people thinking that way about it.
Yep. Maybe let's spell this out since-
Yeah, sure
... uh, we're already on this topic. Elephants in the room are Perplexity and-
Mm-hmm
... Search GPT. That's the-- I think that it's all-- it's no longer called Search GPT. I think they call it ChatGPT Search. How would you contrast your approaches to them, uh, based on what we know of how they work?
And yeah, just any-anything in that, in that area.
Yeah. So these systems, there are a few of them now. Uh, they basically rely on, like, traditional search engines like Google or Bing, and then they combine them with, like, LLMs at the end to, you know, output some paragraph ex-- uh, answering your question.
So they... like Search GPT, Perplexity-
I think they have their own-
Yeah
... crawlers, no? Or-
So there's this important distinction between, like, having your own search system and, like, having your own cache of the web.
Okay.
Like, for example, ex-- so you could create-- you could crawl a bunch of the web. Imagine you crawl a hundred billion URLs, and then you create a key value store of, like, mapping from URL to the document. That is technically called an index, but it's not a search algorithm.
So then to actually, like, when you make a query to Search GPT, for example, what is it actually doing? Let's say it's, it's, it could-- it's using the Bing API, uh, getting a list of results, and then it could go ca-- it has this cache of, like, all the contents of those results and then could, like, bring in the cache, like the index cache.
But it's not actually, like-- It's not like they've built a search engine from scratch over, you know, hundreds of billions of pages. Like, is that distinction clear? It's like, yeah, you could have, like, a mapping from URL to documents, but then rely on traditional search engines to actually get the list of results.
Because it's a very hard problem to take... It's not hard to use DynamoDB and, and, and map URLs to documents. It's a very hard problem to take a hundred billion or more documents and, given a query, like, instantly get the list of results that match.
That's a much harder problem that very few entities on, in the, on the planet have done. Like, there's Google, there's Bing, uh, you know, there's Yandex, but, you know, there are not that many companies that are, that are crazy enough to actually build their search engine from scratch when you could just use traditional search APIs.
So Google had PageRank.
Neural PageRank22:44
Mm-hmm.
That's, like, the big thing. Is there a LLM equivalent or, like, any stuff that you're working on that you wanna highlight?
The link prediction objective can be seen as, like, a Neural PageRank because what you're doing is you're predicting the links people share. And so if everyone is sharing some Paul Graham essay about fundraising, then, like, our model is more likely to predict it.
So, like, inherent in our training objective is this, uh, a sense of, like, high canonicity and-
Mm
... like, high quality. But it's more powerful than PageRank. It's strictly more powerful because people might refer to that Paul Graham fundraising essay in, like, a thousand different ways. And so our model learns all the different ways that someone refers to that Paul Graham essay while also learning how important that Paul Graham essay is.
Um, so it's like, it's like PageRank on steroids kinda thing.
Yeah, I think to me, that's the most interesting thing about search today, like with Google and whatnot. It's like it's mostly, like, domain authority. So, like, if you get backlink, like, if you search any AI term, you get this, like, SEO slop websites-
Oh my God. Yeah, yeah
... with like a bu- a bunch of things in them. So this is interesting. But then how do you think about more timeless, maybe, content? So if you think about, you know, maybe the founder mode essay, right? It gets shared by like-
Yeah
... a lot of people.
Yeah.
But then you might have a lot of other essays that are also good, but they just don't really get a lot of traction, even though maybe the people that share them are high quality. How do you kinda solve that thing when, when you don't have the people authority, so to speak, of who's sharing, whether or not they're worth kinda like bumping up?
Yeah, I mean, you do have a lot of control over the training data, so you could, like- Make sure that the training data contains like high quality sources so that ... Okay, like if you, if your training data, I mean, it's very similar to like lang- language model training.
Like, if you train on like a bunch of crap, your prediction-
Right
... will be crap. Our model will match the training distribution it's trained on. And so we could like ... There are lots of ways to tweak the training data to refer to high-quality content that we want. Yeah, I would say also this like, this slop that is returned by, by traditional search engines like Google and Bing, you have ...
The slop is then, uh, transferred into the, these LLMs in like a search GPT or-
Right
... you know, or other systems like that. Like, if slop comes in, slop will go out. And so yeah, that's another answer to how we're different is like we're not like traditional search engines. We wanna give like the highest quality results and like have full control over whatever you want.
If you don't want slop, you get that. And then if you put an LLM on top of that, which our customers do, then you just get higher quality results-
Yeah
... or higher quality output.
And I use Exa Search very often, and it's very good. Especially-
You said Brightwave uses it, too?
Yeah. Yeah, yeah, yeah. Like, the slop is everywhere.
Yeah.
Especially when it comes to AI, when it comes to investment, when it comes to all of these things where like it's valuable to be at the top.
And this problem's only gonna get worse because-
Yeah, no, totally
... like, yeah.
What else is in the toolkit? So you have search API. You have Exa Search, kinda like the web version. Now you have the list builder. I think you also have web scraping.
Mm-hmm.
Maybe just touch on that. Uh, like I, I guess maybe people, they wanna search and then they wanna scrape, right? So is that kind of the use case that people have?
Yeah. Yeah, a lot of our customers, they don't just want ... Because they're building AI applications on top of Exa, they don't just want a list of URLs. They actually want like the full content, uh, like cleaned, parsed-
Markdown
... maybe Markdown, maybe chunked. Whatever they want-
Yep
... we'll, we'll give it to them. Uh, and so that's been like huge for customers, just like-
Yeah
... getting the URLs and instantly getting the content for each URL is like ... And you can do this for 10 or 100 or 1,000 URLs, whatever you want. Um, that's very powerful.
Yeah, I think this is the first thing I asked you for when I-
Yeah. Right, right, right, right
... tried using Exa.
Yeah. Funny story is like, uh, when I, I built, uh, the, the first version of Exa is like we just happened to store the content-
Yes
... like the first 1,024 tokens because I just kinda like kept it 'cause I thought it ... You know, I don't know why. Uh, really for debugging purposes. And so then when people started asking for content, it was actually pretty easy to serve it.
Yeah.
But then, and then we did that, like Exa took off so, uh, because content was so useful, so that was kinda cool.
It is. Uh, I would say there, there are other players, like Gina I think is in this space. Uh, Firecrawl is in this space. There's a bunch of scraper companies, and obviously scraper is just one, one part of your stack, but, like you might as well offer it since you already do it.
Yeah, it makes sense to ha- It's just easy to have an all-in-one solution, and like we are, you know, building the best scraper in the world. So, uh, scraping is a hard problem, and it's easy to get like, you know, a good scraper.
It's very hard to get a great scraper, and it's super hard to get a perfect scraper. So like, and, and scraping really matters to people.
Do you have a perfect scraper? Or-
Not yet, but we're building.
Okay. Um, the web is increasingly closing to the bots and the scrapers, Twitter, Reddit, Quora, Stack Overflow. I don't know what else. How are you dealing with that? How are you navigating those things? Like, you know, OpenAI is like just paying them money.
Yeah, no, I mean, I think it definitely makes it harder for search engines. One response is just that there's so much value in the long tail of sites that are open.
Okay.
Um, and just like even just searching over those well gets you most of the value. But I mean, there, there is definitely a lot of content that is increasingly not unavailable, and so you could get through that through data partnerships.
The bigger we get as a company, the more, the easier it is to just like, uh, make partnerships. But I, I mean, I do see the world as like the, the future where the data, the, the data producers, the content creators will make partnerships with the entities that find that data.
Any other fun use case that maybe people are not thinking about?
Use Cases27:58
Yeah. Oh, I mean, uh, there are so many.
Your, your customers ... Yeah, just like-
Yeah, yeah. What are people doing-
... tell fun stories
... on Exa? Well, I think dating is a really interesting, uh, application of search that is completely underserved because there's a lot of profiles on the web and a lot of people who wanna find love, and that-
Yeah, I'll, I'll use it. Give me like, you know, some age-
Maybe this file just keeps to myself.
... age boundaries, you know.
There's a-
Education level.
Yeah, right.
Location.
Yeah. I mean, you wanna ... What, what do you wanna do with dating? You wanna find like a partner who matches this education level, who like, you know, maybe has written about these types of topics before. Like, if you could get a list of all the people like that, like I think you will unblock a lot of people.
I mean, there ... I mean, I think this is a very Silicon Valley view of dating for sure. And I'm, I'm well aware of that, but it's just an interesting application of like, you know, I would love to meet like an intellectual partner- ...
um, who like shares a lot of my ideas, and I can ... If you could do that through better search, then yeah.
But what is with Jeff? Jeff has already set- sent me off with, with a few people. So like, Jeff, I think, is my personal Exa for dating. Um.
Right, right.
Uh.
My mom's actually a matchmaker and has got a lot of people married.
Wait.
Yeah.
No kidding.
Yeah, yeah. Search is built into the, the blood family line.
It's in your genes.
Yeah, yeah, yeah.
Other than dating, like I, I know you're having quite some success in colleges.
Mm-hmm.
I would just love to map out some more use cases so that our listeners can just use those examples to think about use cases for Exa, right? Because it's such a general technology that it's hard to really pin down, like what should I use it for and, and what kind of products can I build with it?
Yeah, sure. So I mean, there are so many applications of Exa, and we have, you know, many, many companies using us for a very diverse range of use cases, but I'll just highlight some interesting ones. Like one customer, a big customer, is using us to, um, basically build like a, a writing assistant for students who wanna write, uh, research papers.
And basically, like Exa will search for, uh, like a list of research papers related to what the student is writing, and then this product has like an LLM that like summarizes the papers to ... Basically, it's like a next word prediction, but in, you know, prompted by like, you know, 20 research papers that Exa's returned.
Mm. It's like literally just doing their homework for them.
Yeah, yeah. I guess. But the key point is like it's, it's, uh, you know, it's, it's ... You know, research is, is a really hard thing to do, and you need like high-quality content as input.
Also, we've had Elicit on the podcast.
Yeah. Yeah.
I think it's pretty similar.
Okay.
Uh, they, they do focus- Pretty much on just, just research papers-
Mm-hmm
... and, and that research use case. Basically, I think dating, uh, research. Like, I just wanted to like-
Yeah, yeah. Sure, sure
... spell out more things, like, just the big verticals.
Yeah, yeah. No, I mean, there, there are so many use cases, so like-
Finance, we talked about.
Yeah. I mean, one big vertical is just finding a list of companies. Uh, this is useful for VCs, like you said, who wanna find, like, a list of competitors to a specific company they're investigating, or just a list of companies in some field.
Like, uh, there was one VC that told me that him and his team, like, were using Exa for, like, eight hours straight. Like, like for many days on end, just like, like, uh, doing like lots of different queries of different types.
Like, oh, like all the companies in AI for law, or, or all the companies for AI for, uh, construction, and just, like, getting lists of things, because you just can't find this information with, with traditional search engines. And then, you know, all- finding companies is also useful for, for selling.
If you wanna find, you know, like, if we wanna find a list of, uh, writing assistants to sell to, then we can just- we just use Exa ourselves to find... That's actually how we found a lot of our customers.
Ooh, you can find your own customers-
Yeah, yeah, yeah
... using Exa.
Yeah.
Oh, my God.
Uh, in the spirit of, uh, using Exa to bolster Exa, like, recruiting is really helpful, it is a really great use case of Exa-
Mm-hmm
... um, because we can just get, like, a list of, you know, people who've thought about search, and just get, like, a long list and then, you know, reach out to those people.
When you say thought about, are you, are you thinking LinkedIn, Twitter, or are you thinking just blogs?
Or they've wr- written... I mean, it's pretty general, so in that case, like, ideally, Exa would return, like, the, the, really blogs written by people who have just-
So if I don't blog-
Yeah
... I don't show up to Exa, right? Like, I have to blog.
Uh. Well, I mean, you could show up-
It's like an incentive for people to blog.
Well, if you've written about, uh, search in, on Twitter, and we- we do- we do index, uh-
Yeah, yeah, yeah, yeah
... a bunch of tweets, and then we, we should be able to surface that.
Yeah.
Um-
I mean, this is something I, I tell people. Like, you have to make yourself discoverable-
Yeah
... to the web. Uh, uh, you know, it's called learning in public, but, like, it's even more imperative now.
Yes.
'Cause otherwise you don't exist at all.
Yeah. No, this is a huge, uh, thing, which is, like, search engines completely influence-- They have downstream effects. They influence the internet itself. They influence what people choose to create.
Oh.
And so Google, because they're a keyword-based search engine, people, like, kind of like-
Keyword stuff.
Yeah. They're ins- they're, they're incentivized to create things that just match a lot of keywords, which is not very high quality. Uh, whereas Exa is a search algorithm that, uh, optimizes for, like, high quality and actually, like, matching what you mean, and so people are incentivized to create content that is high quality.
Mm.
That, like, they create content that they know will be found by the right person. So, like, you know, if I am a search researcher and I want to be found by Exa, I should blog about search and all the things I'm building because, because now we have a, a search engine like Exa that's powerful enough to find them.
And so the search engine will influence, like, the downstream internet in all sorts of amazing ways.
Yeah, that's interesting.
Yeah. Uh, whatever the search engine optimizes for is what the internet looks like.
Yeah.
Yeah.
Are you familiar with the Mc- term McLuhanism?
No. What's that?
Uh, it's this concept that, uh, like, first we shape tools, and then the tools shape us.
Okay. Yeah.
Uh, so there's, like, this reflexive connection between the things we search for and the things that get searched.
Yes.
So, like, once you change the tool that searches, the, the, the things that get searched also change.
Yes. I mean, there was a clear example of that with-
We've had 30 years of Google's-
Yeah, exactly
... shaping-
Google has basically trained us to think of search- ... like Google has-- Google is search, like, in people's heads, right?
Yeah, yeah.
It's-- One, uh, hard part about Exa is, like, uh, ripping people away from that notion of search and expanding their sense of what search could be. Because, like, when people think search, they think, like, a few keywords, or at least they used to.
They think of a few keywords, and that's it. They don't think to make these, like, really complex paragraph-long requests for information and get a perfect list. ChatGPT was an interesting, like, thing that expanded people's understanding of search, because you start using ChatGPT for a few hours, and you go back to Google, and you, like, paste in your code, and Google just doesn't work, and you're like, "Oh, wait.
It-- Google doesn't d- work that way." So, like, ChatGPT expanded our understanding of what search can be, and I think Exa is, uh, is part of that. We wanna expand people's notion, like, hey, you could actually get whatever you want.
Yeah.
I search on Exa right now, "people writing about learning in public," and I was like, "Is it gonna come up with Swix?"
Auto-Prompting34:07
Am I, am I on there?
You're not, because-
Bro
... it's so-
Exa, Exa failed
... no, it's, it's so about be- because it thinks about learning, like, in public, like public schools, and, like-
Mm
... focuses more on that.
Mm.
You know, it's like, how-- When there are, like, these highly overlapping things-
Mm-hmm
... like, this is, like, a good result-
Yeah
... based on the query, you know? But, like, how do I get to Swix? Right? So i- if you're, like, in these subcultures, I don't think this would work in Google well either, you know? But I, I don't know if you had any learnings from-
No, I'm the first result on Google.
"People writing about learning in public." You're not first result anymore, I guess.
I mean, just type "learn in public" in Google.
Well, yeah, yeah, yeah, but this is also, like, this... In Google, it doesn't work either.
Gotcha, dude.
That's what I'm saying. It's like, how-- When you have, like, a movement-
There's confusion about the, like, what you mean, like, your intention is a little, uh-
Yeah. It's like-
Yeah
... I'm using, I'm using a term that, like, I didn't invent-
Yep
... but I'm kinda taking over.
Yeah.
But, like, there's just so much about that term already that it's hard to overcome, if that makes sense. Because public schools-
Oh
... is like, uh, well, it's, it's hard to overcome public schools, you know? So-
Yeah. So there's the right solution to this, which is to specify more clearly what you mean. And I'm not expecting you to do that. But, so, the, the right interface to search is actually an LLM. Like, you should be talking to an LLM about what you want, and the LLM translates its knowledge of you-
Yeah
... or knowledge of what people usually mean into a query that Exa then uses. And-
Which you have called auto prompts, right? Or-
Yeah, but it's, like, a very light version of that.
Very, very light version.
And really, it's just-- Basically, the right answer is it's the wrong interface. And, like, very soon, the interface to search, and really to everything, will be LLMs, and the LLM just has a full knowledge of you, right?
Yeah.
So we're kind of building for that world. We're skating to where the puck is gonna be. And so since we're moving to a world where, like, LLMs are in the interface to everything, you should build a search engine that can handle complex LLM queries, queries that come from LLMs.
Because L- you're probably too lazy, and I'm too lazy too, to write, like, a whole paragraph explaining, "Okay, this is what I mean by this word," but an LLM is not lazy. And so, like, the LLM will spit out, like, a paragraph or more explaining exactly what it wants.
You need a search engine that can handle that. Traditional search engines like Google or Bing, they're actually designed for humans typing keywords. If you give a paragraph to Google or Bing, they just completely fail. And so Exa can handle paragraphs, and we wanna be able to handle it more and more until it's, like, perfect.
What about opinions? Do you have list-- When you think about the list product, do you think about- Just finding entries. Do you think about ranking entries? I'll, I'll give you a dumb example.
Yeah.
So on Lindy, I've been building the spot that every week gives me, like, the top fantasy football waiver-
Yeah
... pickups.
Okay.
But every website is, like, different opinions on, like, you should pick up these five players, these five players. When you're making lists, do you wanna be kinda, like, also ranking and, like, telling people what's, what's best? Or, like, are you mostly focused on just surfacing information?
There's a really good distinction between, uh, filtering to, like, things that match your query, and then ranking based on, like, what is, like, your preferences. And ranking is-- Like, a filtering is, is objective. It's like, does, does this document match what you asked for?
Uh, whereas ranking is more subjective. It's like, what is the best? Well, it depends what you mean by best, right? So first, first table stakes is let's get the filtering into a perfect place where you actually, like, every document matches what you asked for.
No search engine can do that today. And then ranking, you know, there are all sorts of interesting ways to do that, where you, like, you've maybe for-- you know, have the user, like, specify more clearly what they mean by best.
You could do-- And if the user doesn't specify, you do your best, you do your best based on, uh, what people typically mean by best. But ideally, like, the user can specify, "Oh, when I mean best, I actually mean ranked by the, you know, the number of people who visited that site," let's say, is, is one example ranking.
Or, "Oh, what I mean by best..." Let's say you're listing companies. "What I mean by best is, like, the ones that have, uh, you know, have the most employees," or something like that. Like, there are all sorts of ways to rank a list of results that are not captured by something as subjective as best.
Yeah. Yeah, I mean, it's like, who are the best NBA players in the history?
Yeah, right.
It's like everybody has their own-
Right, right. But I mean, the, the, the search engine should definitely, like, even if you don't specify it, it should do as good of a job as possible.
Yeah, yeah, no, no, totally.
Yeah, yeah.
Agentic Search38:14
Yeah.
It's a new topic to people because we're not used to a search engine that can handle, like, a very-
Mm-hmm
... complex ranking system. Like, you think to type in best basketball players and not something more specific because you know that's the only thing Google could handle.
Mm-hmm.
But if Google could handle, like, oh, basketball players ranked by, like, number of shots scored on average per game, then you would do that, but you know they can't do that, so.
Mm-hmm.
Yeah, that's fascinating. So you haven't word-- used the word agents, but you're kind of building a search agent. Do you believe that that is agentic in feature? Do you think that term is distracting?
I think it's a good term. I do think everything will eventually become agentic-
Mm-hmm
... and so then the term will lose power.
Yeah.
But yes, like what we're building is agentic. It-- in the sense that it takes actions, it, it decides when to go deeper into something.
Has a loop.
It has a loop, right. It, it feels different from traditional search, which is like an algorithm, not an agent.
Yeah.
Ours is a combination of an algorithm and an agent.
I think my reflection from seeing this in the coding space, where there's basically... So the classic framework for thinking about this stuff is the self-driving levels of autonomy, right? Level one to five.
Mm-hmm.
Typically, the level five ones all failed 'cause they're full autonomy, and we're not, we're not there yet, and people like control. People like to be in the loop. So the, the, the level ones was Copilot first, and now it's like Cursor and whatever.
So I feel like if it's too agentic, it's too magical, like, like a, like a one-shot, I stick a, I stick a paragraph into the, the text box, and then it spit, spits it back to me, it might feel like I'm too disconnected to-- from the process and I don't trust it, as opposed to something where I'm more intimately involved with the research product.
I see. So like, uh... Wait, so the earlier versions are-
So, so if trying to stick to the example of the basketball thing, like best basketball player, but instead of best, you, you actually get to customize it with, like, whatever the, the metric is that you, you guys care about.
I'm still not a basketball player .
Yeah, me neither .
But, uh, but, but you know, like, like, like be- people like to be in... My, my thesis is that agents, level five agents failed because people like to-
Mm
... kind of have drive assist rather than full self-driving.
I mean, a lot of this has to do with how good agents are. Like, at some point, if agents for coding are better than humans at all tasks, then-
Sure
... then humans-
We're not there yet
... block. Yeah, yeah, we're not there yet. So, like, in a world-
But, like, the way-
... where we're not there yet-
What you're pitching us-
Yeah
... is, like, you're, you're kinda saying you're going all the way there. Like, I kinda-- I, I think o1 is also very full, full self-driving. You don't get to see the plan.
Mm.
You don't get to affect the plan yet. You just fire off a query, and then it goes away for a couple minutes and comes back, right? Which is effectively what you're saying you're gonna do too.
And you think there's a-
There's a in-between.
I saw, okay, so in building this product, we're exploring new interfaces, because what does it mean to s- like, kick off a search that goes-
Yeah
... and takes ten minutes? Like, is that a good interface? Because what if the search is actually wrong-
Yeah
... like, or it's not exa-exactly specified to what you mean?
Which is why you get previews.
Yeah, you get previews. So it is iterative.
Mm.
But ultimately, once you've specified exactly what you mean, then you kinda do just wanna kick off a batch job, right?
Yeah.
So perhaps what you're getting at is, like, uh, there's this barrier with agents where you have to, like, explain the full context of what you mean, and a lot of failure modes happen when you ha- when you don't.
Yeah.
There's failure modes from the agent just not being smart enough, and then there's failure modes from the agent not understanding exactly what you mean. And there's a lot of context that is shared between humans that is, like, lost between, like, humans and, and this, like, new creature .
Yeah. Yeah, because people don't know what's going on. I mean, to me, the best example of, like, system prompts is, like, why are you writing you're a helpful assistant? Like, of course, you should be a helpful assistant.
Right.
But, but people don't yet know, like, can I assume that you know that? You know, it's like, why did the... And now people write, "Oh, you're a very smart software engineer." Like-
200 IQ-
You never make-
... expert at Python
... Yeah, like you never make mistakes. Like, were you gonna try and make mistakes before? So I think people don't yet have an understanding. Like with, with driving, people know what good driving is.
Mm.
It's like, don't crash. Stay within kind of like a certain speed range. It's like follow the directions. It's like I don't really have to explain all of those things, I hope. But with AI and, like, models and, like, search, people are like, "Okay, what do you actually know?
What are, like, your assumptions about how search, how you're gonna do search?"
Mm.
And, like, can I trust it? You know, can I influence it? So I think that's kind of the, the middle ground. Like, before you go ahead and, like, do all the search, it's like, can I see how you're doing it?
And that maybe helps you-
Show your work
... kinda like, yeah.
Yeah.
Steer you.
Yeah, yeah. No, I mean- Yeah, so you're saying even if you've crafted a great system prompt, you want to be part of the process itself, uh, because the system prompt doesn't-- it doesn't capture everything, right?
Yeah.
So yeah. A system prompt is like you get to choose the person you work with. It's like, "Oh, like I want, I want a software engineer who thinks this way about code." But then even once you've chosen that person, you can't just give them a high-level command, and they go do it perfectly.
You have to be part of that process. So yeah, I agree.
Just a side note. For my sy- my favorite system prompt programming anecdote now is the Apple Intelligence system prompt that someone, someone's, uh, prompt injected it and seen it. And, like, the Apple Intelligence has the words like, "Please don't, don't hallucinate."
And it's like, of course, we don't want you to hallucinate, right? Like, so-
Yeah
... it's exactly that, that what you're talking about. Like, we should train this behavior into the model, but somehow we still feel the need to, to inject into the prompt. And I still don't even think that we're very scientific about it.
Like it-- I think it's almost like cargo culting. Like-
Yeah
... we have this, like, magical, like, turn around three times, throw salt over your shoulder before you do something, and, like, it worked the last time, so let's just do it the same time now. And like, we do-- there's no science to this.
I do think a lot of these problems might be ironed out in future versions, right?
Yeah.
So... And like, they might, they might hide the details from you. So it's like they actually all of them have a system prompt that's like, "You are a helpful assistant." You don't actually have to include it- ... uh, even though it might actually be the way they've implemented it in the back end.
Yeah. It should be done in RLEF. Okay, uh, one question I, I always just kind of curious about just so I, I-- this episode is I'm gonna try to frame this in terms of just the general AI search wars.
Mm-hmm.
You know, you're, you're one player in that. Um, there's Perplexity, ChatGPT Search, and Google, but there's also, like, the B2B side. Uh, we had Drew Houston from Dropbox on, and, and he's competing with Glean w-who've, uh, we've also had DD from, from Glean on.
Is there appetite for Exa for my company's documents?
o1 & AGI44:19
There is appetite, but I think we have to be-
Focused
... disciplined. Focused-
Yeah
... disciplined. I mean, we're already taking on, like, perfect web search-
Yeah
... which is a lot. Um, but I mean, ultimately, we wanna build a perfect search engine, which definitely for a lot of queries involves your, your personal information, your company's information. And so, yeah, I mean, the grandest vision of Exa is perfect search really over everything.
Every domain.
You know, we're gonna have an Exa satellite, uh, because, because satellites can gather information that, uh, is not available publicly. Uh-
Gotcha.
Yeah.
Can we talk about AGI?
Of course.
We never, we never talk about AGI, but you had-
Right
... uh, this whole tweet about o1 being the biggest-
Mm-hmm
... kinda like AI step function towards it. Why does it feel so important to you? I know there's kinda like always criticism and saying, "Hey, it's not the smartest. Something is better." It's like blah, blah, blah. What did you see?
So you say this is what Ilya see- ... what Sam see, what, what they will see.
I've just, I've just, you know, been connecting the dots. I mean, this was the key thing that a bunch of labs were working on, which is like, can you create a reward signal? Can you teach yourself based on a reward signal?
Whether you're-- if you're trying to learn coding or math, if you could have one model, say, uh, be a grading system that says, like, "You have successfully solved this programming assessment", and then one model, like, be the generative system that's like, "Here are a bunch of programming assessments."
You could train on that. So basically, whenever you could create a r-reward signal for some task, you could just generate a bunch of tasks for yourself, see that like, oh, on two of these thousand you did well, and then you just train on that data.
It's basically like creating your own data for yourself. And like, you know, all the labs working on that, OpenAI built the most impressive product doing that, and it's just v-very-- it's very easy now to see how that could, like, scale to just solving, like, like solving programming or solving mathematics.
Uh, which sounds crazy, but everything about our world right now is crazy. Um, and so I think if you remove that whole like, "Oh, that's impossible," and you just think really clearly about, like, what's now possible with like what they, what they've done with o1, it, it's easy to see how that scales.
How do you think about older GPD models then? Should people still work on them? You know, if like, uh, obviously they just had the new Haiku. Like, is it even worth spending time, like, making these models better versus just...
You know, Sam talked about o2 at that day. Uh, so obviously, they're, they're spending a lot of time in it. But then you have maybe the GPU poor, which are still working on making Llama good. Uh, and then you have the follower labs that do not have an o1-like model out yet.
Yeah, this kinda gets into like, uh, what will the ecosystem of, of models be like in the future, and is there room-- is, is everything just gonna be o1-like models? I think, well, I mean, there's definitely a question of like inference speed, and if certain things...
Like o1 takes a long time because it has to think. Well, I mean, o1 is, is two things. It's like, one, it's, it's use-- it's bootstrapping itself. It's teaching itself, and so the base model is smarter. But then it also has this, like, inference time compute where it could, like, spend, like, many minutes or many hours thinking.
And so even the base model, which is also fast, it doesn't have to take minutes. It could take, uh, is, is better, smarter. I believe all models will be trained with this paradigm. Like, you'll want to train on the best data, but there will be many different size models from different, very many different, like, companies, I believe.
Yeah. Because like I don't-- Yeah. I mean, it's hard, hard to predict, but I don't think OpenAI is gonna dominate like every possible LLM for every possible use case. I think for a lot of things, like you just want the fastest model, and that might not involve o1 methods at all.
I would say if you were to take the Exa being o1 for search literally, you really need to prioritize search trajectories. Like almost maybe paying a bunch of grad students to go research things, and then you kind of track what they search and like what, what, what the sequence of searching is.
Because it seems like that is the, the gold mine here, like the chain of thought or the s-thinking trajectory.
Yeah. When it comes to search, I've always been skeptical of human-labeled data.
Okay.
Uh,
Yeah, please.
We, we tried something at our company, uh, at Exa recently where, uh, me and a, a bunch, a bunch of engineers on the team like, uh, labeled a bunch of queries, and it was really hard. Like, you know, you have all these niche queries, and you're looking at a bunch of results, and you're trying to identify which matches the query.
Like it's talking about, you know- The intricacies of like some biological experiment or something. I have no idea, like I, I don't know what matches. And what, what labelers like me tend to do is just match by keyword.
I'm like, "Oh, like this document matches a bunch of keywords, so it must be good." But then you're actually completely missing the meaning of the document. Whereas an LLM like GPT-4 is really good at labeling, and so I, I, I actually think like you could just-- we could- we could get by, and which we are right now doing, uh, using like LLMs as the labelers.
Specifically for search, I think it's interesting-- it's different between like search and like GPT-5 are different because GPT-5 might benefit from training on a lot of PhD notes because like GPT-5 might have to do like very, very complex like, uh, problem-solving, uh, in, in, af- when it's given an input.
But with search, it's actually a very different problem. You're, you're asking simple questions about billions of things. Uh, so like whereas like GPT-5 is asking a really hard... it's like solving a really hard question, but it's one. It's like one question, a PhD level question.
With search, you're asking like simple questions about billions of things like, "Is this a startup? Did this person write a blog post about search?" You know, those are actually simple questions. You don't need like PhD level training data.
Does that make sense?
Yeah. What else we got here? Uh, nap pods.
Oh, yeah.
Culture & Business49:37
What's the-
Yeah, so like just generally, I think, uh, Exa has a very interesting company building vibe. Like you, you have a meme lord CTO, um-
I guess.
Yeah.
I don't know. Like, a- and you, you have-- you, you just generally, um, are counter-consensus in a bunch of things.
Mm.
What is the culture at Exa like?
Yeah. I, me and Jeff are-- I mean, we've been best friends since like, like we met, like met like first day of college and been best friends ever since, and we have a really good vibe, I think, that's like intense but also really fun and like, like funny honestly.
We have a ton of-- like we just laugh a lot, a ton at Exa, and I think that's just like you see that in every part of our culture. We don't really care about how the world sees anything.
Like me and Jeff are just like that. Like we're just thinking really just like, like, "What should we do here? Like, what do we need?" And so in the nap pod case it was like people get tired a lot when they're coding or doing anything really, and like why can't we just sleep here?
Or, or like nap and, uh, okay, if we need a nap, then we should get a nap pod. It's crazy to me that there aren't nap pods in lots of companies because like I get tired all the time.
I, I take a nap like every other day probably for like twenty minutes. I'm actually never actually napping. I'm just thinking about a problem, but closing my eyes really like-
Yeah
... um, first of all, it makes me come up with more creative solutions and then also actually, uh, gives me some rest. So which is awesome.
But Google was the original company that had the nap pods at work, right?
Oh, okay. Well, then-- well, at, at one point Google was thinking from first principles at everything too. Um, and that was reflected in their nap pod.
But also you, you like, you like didn't just get a nap pod for your office. You like found something from China-
Oh, yeah
... and you were like, "Who wants to get in on this? Let's get a container full of them."
Yeah. Well, we tried, we tried to be frugal. So like we were- ... we were looking at like different nap pods and then, uh, at some point we were like, "Wait, China probably has solved this problem." And so then we ordered it from China, and it was actually so heavy.
Like when it came off the truck it was like five hundred pounds, and I-- like the truck was like having trouble like putting it on the ground. And so like me and the, the delivery guy were like trying to hold it in and we couldn't.
We were struggling, so someone came out from on- from the street and like hard-- started helping us.
You could've hurt yourself, man.
I, I, I know. It was really dangerous. But we did it, and then it was awesome. And-
It's funny, I was reading the TechCrunch article about it.
Oh, yeah.
There was a TechCrunch article on the nap pods?
Yeah. And then-
Oh, God
... Jeff explained-- well, they quote Jeff. And this paragraph says, "So the nap pods maintain employees' ability to stop work and sleep rather than the idea that," in quotes, "employees are slaves," close quote. I don't know.
Jeff has a way with words.
What, what-- I, I'm like, I'm sure that's not what he meant, you know, but I'm curious like just like how people-- there's always like this... I think for a little bit it went away about like startups and kinda like hustle culture-
Hustle culture. Yeah
... and like all of that, and I think now with AI people are like have all these feelings towards AI that are kinda like-
I think Exa are pro hustle culture, right?
Yeah, but I, I mean, I mean, ideally the hustle is like people are just having fun, which is like-
Yeah, yeah
... people, people are just having fun.
Yeah, but I would say from the outside it's like people don't like it, you know? I'm saying people not in-
Mm
... in AI and kind of like in tech are kinda like, "Oh, these guys are at it again. These are like the same people that gave us underpaid drivers," like whatever. It's like so-
I see
... it, it was just funny to see somehow they wanted to make it sound like Jeff was saying employees are slaves, but like that doesn't-
Oh, yeah. I don't know.
That doesn't make sense.
But yeah, I mean, okay, I can't imagine a more exciting experience than like building something from scratch that's like a huge deal with a bunch of your friends. Our team is going to look back in ten years and think this was like the most beautiful experience that you could have in life, and like that's how I think about it.
And, um, yeah, that's just so... I, it's not, it's not a hustle or not. It's like, is this like, like does this satisfy like your core desire to like build things in the world?
Mm-hmm.
And it does.
Yep. Anything else we didn't cover? Any parting thoughts? Are you hiring? Are you-- obviously you're looking for more people to use it, but...
Yeah, yeah. We're definitely hiring. We're, we're growing quite fast, and we have a really smart team of engineers and researchers, and we now have a-- we just purchased a five million dollar, uh, H200 cluster, so we have a lot more, uh, compute to play with.
Do you run all your own inference?
We do a mix of our cluster and like AWS.
Okay.
Inference, we, we use these, our-- so we have our current cluster, which is like A100s, and now we've updated to the new one. Uh, we use it for training and research.
What's the training versus inference budget? Like is it like a-
Ooh
... is it fifty-fifty? Is it-
Uh, yeah. We-- there will be more inf- inference for search for sure. Um-
The, the other thing I mentioned, so by the way, I'm, I'm like sidetracking, but I'm just kinda throwing this in there 'cause I, I always think about the economics of AI search. Like for those, I think, I think if you look up, there's the upper limit is going to be whatever you can monetize off of ads, right?
So for Google, let's say it's like a one cent per thousand views, something like that.
Mm.
I don't know the, the exact number. The exact number is floating around out there. That means that's your revenue, right? Then your cost has to be lower than that.
Mm-hmm.
And so at some point, like for an LLM inference call to be made for every page view, you, you need to get it lower than the money that you take in for-
Yes
... for that. And like one of the things that I was very surprised, surprised for Perplexity and Character as well was that they couldn't get it so low that it would be, uh, reasonable. I think for you guys it is a mix of front-loading it by indexing, so you only run that compute like once a month, once a, once a quarter, whatever you, you do re-indexing, and then it's just a little bit more when you, when you do inference, when the search actually gets done, right?
Like so I think when people work out like the economics of such a business, they, they have to kinda think about where do you put the, the costs.
Yes, yes. I mean, uh, definitely you have to-- you cannot run LLMs over the whole index, you know, billions of things at query time, so you have to pre-process things-
Yeah
... us- usually with LLMs. But then you, you can do a re-rank over like, you know, ten, 30, 100 depending on 1,000, depending on how... You know, you could, you could play with different sizes of L- of transformers to get the cost to work out.
I mean, one really interesting thing is like we're building a search engine at a time where LLM costs are going down like crazy. When some very useful tool goes down in cost by 200X in like the space of, I don't know, a couple years, there are going to be new opportunities in search.
Right?
Mm-hmm.
So like to, to not integrate this and build a-- to not like rethink search from scratch, the search algorithm itself, given the fact that things are going down 200X, is crazy.
Thank you so much for coming on, man. It was fun.
Yeah, thank you. This was so fun. Really fun.






