Intro0:00
Okay, um, I think we're gonna kick this off. Um, thanks to everyone who made it, uh, early morning. Um, uh, uh, this is like really weird experiments that we wanted to try because one, we saw this space, uh, and but two, also I've been to a, a number of these things now, and, um, I always felt like there was not enough like industry content for, for people, and we wanted an opportunity while everyone is in town in like one central spot to get everyone together, um, to talk about the best stuff of the year, review the year.
It's, uh, very nice that Europe is always the end of the year. Um, and so I'm very honored that, uh, Sarah and Pranav have, have agreed to help us kick this off. Um, Sarah, I've known for, I was actually counting seventeen years.
Sounds good.
Um, and but she's, she's gone-- she's, uh, been enormously successful as an AI investor. Um, even, uh, even when you-- during your Greylock days, I was tracking your, your investing, and it was, it was, uh, it's come a long way since then.
Um, and Pranav, uh, I, I've known, I've known, uh, shorter, but he's also starting to write, uh, really incredible posts and opinions about, uh, what he's seeing as a, as an investor. So I wanted to kick this off with an industry session.
Um, we have, uh, a great day of sort of like best of year recaps, uh, for, uh, lined up. I think Vic is here as well, um, and, uh, and, and the Roboflow guys. So, uh, I would just let you keep, uh, kick it, kick it off.
Thank you.
Conviction Fund1:30
Hi, everyone. Uh, my name is Sarah Guo, and thanks to, uh, Sean and friends here for having me and Pranav. So, um, uh, I, I'd start by just giving thirty seconds of intro. I promise this isn't an ad.
Uh, we started a venture fund called Conviction about two years ago. Here is a set of the investments we've made. Uh, they range from, uh, companies at the infrastructure level, um, in terms of feeding the revolution to, uh, foundation model companies, alternative architectures, domain-specific training efforts, and of course, applications.
Um, and the premise of the fund, Sean mentioned I worked at Greylock for about a decade before that and came from the product engineering side, was that, uh, we, we thought that there was a really interesting technical revolution happening, uh, that it would probably be the biggest change in how people use technology in our lifetimes, and that represented huge economic opportunity.
And, and maybe that there'd be an advantage versus the incumbent venture firms in that when the floor is lava, the dynamics of the markets change, the types of products and founders that you back change, uh, it's a lot for existing firms to ingest, and a lot of their mental models may not apply in the same way.
Uh, and so there was an opportunity for first principles thinking, and if we were right, we would do really well and get to work with amazing people. And so we are two years into that journey, and we can share some of the opinions and predictions we have with all of you.
Um, sorry, I'm just making sure that isn't actually blocking the whole presentation. Uh, and Pranav's gonna start us off.
Um, so quick agenda for today. We'll cover some of the model landscapes and themes that we've seen in twenty twenty-four, uh, what we think is happening in AI startups, and then some of our latent priors, uh, on what we think is working in investing.
So the-- Um, I, I thought it'd be useful to start from like what was happening at NeurIPS last year in December twenty twenty-three. So in October twenty twenty-three, OpenAI had just launched the ability to upload images to ChatGPT, which means up until that moment, it's hard to believe, but like roughly a year ago, you could only input text and get text out of ChatGPT.
Um, the Mistral folks had just launched the Mixtral model right before the beginning of NeurIPS. Google had just announced Gemini. I very genuinely forgot about the existence of Bard before making these slides. And Europe had just announced that they were doing their first round of AI regulation, but not to be their last.
And when we were thinking about like what's changed in twenty twenty-four, there's at least five themes that we could come up with that feel like they were descriptive of, of what twenty twenty-four has meant for AI and for startups.
Model Battle4:01
And so we'd start with, um, first, it's a much closer race on the foundation model side than it was in twenty twenty-three. So this is LM Arena. They're, uh, ask users to rate the evaluations from, uh, from ge-- of generations from specific prompts.
So you get two responses from two language models, answer which one of them is better. The way to interpret this is like roughly a hundred ELO difference means that you're preferred two-thirds of the time. And a year ago, every OpenAI model was like more than a hundred points better than anything else.
And the view from the ground was roughly like OpenAI is the IBM. There is no point in competing. Everyone should just give up, go work at OpenAI, or attempt to use OpenAI models. And I think the story today is not that.
Um, I think it would have been unbelievable a year ago if you told people that A, the best model today on this, at least on this eval, is not OpenAI, and B, that it was Google, would have been pretty unimaginable to the majority of researchers.
But actually, there are a, a variety of, of op-- of proprietary language model options and some set of open source options that are increasingly competitive. And this seems true not just on the eval side, but also in actual spend.
So this is ramp data. There's a bunch of colors, but it's actually just OpenAI and Anthropic spend. And the OpenAI spend at the beginning, at the end of last year in November of twenty-three was close to ninety percent of total volume.
And today, less than a year later, it's closer to sixty percent of total volume, um, which I think is indicative both that language models are pretty easy APIs to switch out and people are trialing a variety of different options to figure out what works best for them.
Related second trend that we've noticed is that open source is increasingly competitive. So this is from the scale leaderboards, which is, uh, a set of independent evals that are not contaminated. And on, on a number of topics that actually the, the foundation models clearly care a great deal about, open source models are pretty good on math instruction following and adversarial robustness.
The Llama model is amongst the top three of, of evaluated models. Uh, I included the agenting tool use here just to point out that this isn't true across the board. There are clearly some areas where foundation model companies have had more data or more expertise in training against these use cases.
But models are surprisingly an increasing-- open source models are surprisingly increasingly effective. Um, this feels true across evals. This is the MMLU eval. Um, I wanna call out two things here. One is that It's pretty remarkable that the ninth-best model and two points behind and, uh, the, the best state-of-the-art model is, is actually a seventy billion parameter model.
I think, um, this would've been surprising to a bunch of people who were-- the belief was largely that most intelligence is just an emergent property, and there's a limit to how much intelligence you can push into smaller form factors.
Uh, in fact, a year ago, the, the best small model or under ten billion parameter model would've been Mistral 7B, which on this eval, if memory serves, is somewhere around a sixty. Um, and today, that's the Llama 8B model, which is more than ten points better.
The, the gap between what is state-of-the-art and what you can fit into a fairly small, uh, form factor is ac-actually shrinking. Um, and again, related, the-- we think the price of intelligence has come down substantially. This is, this is a graph of flagship OpenAI model costs, where the cost of the API has come down roughly eighty, eighty-five percent in, call it the last year, year and a half, um, which is pretty remarkable.
This isn't just OpenAI 2. This is also, like, the full set of models. This is from Artificial Analysis, which tracks cost per token across a variety of different APIs and public inference options. And, like, we were doing some math on this.
Token Deflation7:07
If you wanted to recreate, like, what a-- the kind of data that a text editor had or that, like, something like Notion or Coda that's somewhere in the volume of a couple thousand dollars to create that volume of tokens, that's pretty remarkable and impressive.
New Modalities7:30
It's clearly not the same distribution of data. But just as, like, a sense of scope, the-- there's an enormous volume of data that you can create. Uh, and then fourth, we think new modalities are beginning to work. Um, let's start quickly with biology.
We're lucky to work with the folks at Tri Discovery, um, who just released Tri-1, which is open source model that outperforms AlphaFold3. It's impressive that this is, like, roughly a year of work with a pretty specific dataset and then pretty specific technical beliefs.
But, um, models in domains like biology are beginning to work. We think that's true on the voice side as well. Um, point out that there were voice models before. Things like ElevenLabs have existed for a while. But we think low latency voice is more than just a feature.
It's actually a net new experience interaction. Um, using voice mode feels very different than the historical transcription first models. Same thing with many of the Cartesian models. Um, and then a new nascent use case is execution. So Cloud launched Computer Use.
OpenAI launched Code Execution inside of Canvas yesterday, and then I think Devin just announced that they-- you can all try it for five hundred dollars a month, um, which is pretty remarkable. It's a, a set of capabilities that have historically never been available to vast majority of population, and I think we're still in early innings.
Cognition, the company, was founded under a year ago. First product was roughly nine months ago, which is pretty impressive.
And if you recall, like, a, a year ago, the point of view on SWE-bench was, like, it was impossible to surpass what-
Yeah
... team percent or so. Um, and I think the, the whole industry now considers that, uh, if not trivial, um, accessible.
Yeah.
Um, last new modality w-we wanted to call out, although there are many more, is video. Um, I took the liberty-- I got early access to Sora and managed to sign up before they cut off accesses. So, um, here is my favorite joke in the form of a video.
Hopefully, someone here can guess it.
Yeah. You're telling me a shrimp fried this rice? It's, uh, pretty bad joke, but I really like it. Um, and I think this one, the next video here is, uh, one of our portfolio companies, HeyGen, that, um, translated and does the dubbing for, uh, or lip sync and dubbing for, um, live speeches.
So this is Javier Milei, who speaks in Spanish, but here you will hear him in English if this, if this plays. Um, and you can see that you can capture the original tonality of, of his speech and performance.
I think audio here doesn't work, but we'll, we'll- Does it? -post something publicly. Sure. Um...
Let's give it a shot.
Yeah. Let's give it a shot. We are supposed to defend- Excellent ... the West from the Western world. Yeah. And you can hear that this captures, like, his original tone, uh, and, like, the emotion in his speech, which is definitely new and pretty impressive from, from new models.
Um, so the last, uh, the-- yeah, that makes sense. Um, the last point that we wanted to call out is, uh, the much purported end of scaling. I think there's a great debate happening here later today on the question of this.
Scaling Limits10:07
But we think at minimum, uh, it's hard to deny that there are at least some limits to the, the clear benefits to increasing scale. Um, but there also seems like there are new scaling paradigms. So the question of test time compute scaling is a pretty interesting one.
It seems like OpenAI has cracked a version of this that works, and we think, A, foundation model labs will come up with better ways of doing this, and B, so far, it largely works for very verifiable domains, things that look like math and physics and maybe secondarily software engineering, where we can get an objective value function.
Um, and I think an open question for the next year is going to be how do we generate those value functions for spaces that are not as well constrained or well defined. Um, and so the question that this leaves us in is, like, well, what does that mean for startups?
Startup Metrics11:00
And I-I think a prevailing view has been that we live in an AI bubble. There's an enormous amount of funding that goes towards AI companies and startups that is largely unjustified based on outcomes and what's actually working on the ground.
Uh, and startups are largely raising money on hype. And so we pulled some PitchBook data, and the twenty twenty-four number is, like, probably incomplete since not all rounds are being reported, and largely suggests, like, actually there is a, a substantial recovery in funding, and maybe twenty twenty-five looks something like twenty twenty-one.
But if you break out the numbers here a bit more, um, the red is actually just a small number of foundation model labs, like what you would think of as the largest labs raising money, which is upwards of thirty to forty billion dollars this year.
And so the reality of the funding environment actually seems, like, much more sane and rational. It doesn't look like we're headed to a version of twenty twenty-one. In fact, the, the foundation model labs account for an outsized amount of, of money being raised, but the, the set of money going to companies that are working seems much more rational.
And we wanted to give you-- we can't share numbers for every company, but this is one of our portfolio companies growing really, really quickly. Um, we think zero to twenty in just PLG style spending is pretty impressive. If any of you are doing better than that, you should come find us.
We'd love to chat.
Winning Patterns12:12
And so what we wanted to try and center a discussion on, this is certainly not all of the companies that are making ten million more or revenue and growing, but we took a selection of them and wanted to give you a couple ideas of patterns that we've noticed that seem to be working across the board, um- The first one that we've noticed is, like, first wave service automation.
So we think there's a large amount of work that doesn't get done at companies today, either because it is too expensive to hire someone to do it, it's too expensive to provide them context and enable them to be successful, uh, at, uh, at whatever the specific role is, or it's too hard to manage, um, those set of people.
So for scribing, it's too expensive to hire the specific set of people. For Sierra and Decagon, for customer support sell companies, it's really useful to do, like, next level automation, and then there's obviously growth in that. And for Harvey and Even Up, the story is, um, you can do first wave professional services and then grow beyond that.
Second trend that we've noticed is, uh, better search, new friends. So we think that there's a... It, it's pretty impressive, like, how effective text modalities have been. So Character and Replica have been remarkably successful companies, and there's a whole host of not-safe-for-work chatbots as well, um, that are pretty effective at just text generation.
They're pretty compelling mechanisms. And on the productivity side, Perplexity and Glean have demonstrated this as well. I worked at a search company for a while. I think the changing paradigms of the how people capture and learn information is pretty interesting.
We think it's likely text isn't the last medium. There are infographics or sets of information that seem more useful, or sets of engagement that are more engaging. Um, but this feels like a pretty interesting place to start.
Oh. Y- yeah. Okay. Mike. So o- one thing that I- I've worked on investing in in a long time is democratization of different skills, be they creative or technical. This has been an amazing few years for that. Across, uh, different modalities, audio, um, video, uh, general image media, text, and now code and, and, and really fully functioning applications.
Um, o- one thing that's really interesting about the growth driver for all of these companies is the, the end users in large part are not people that we thought of as... We the venture industry, you know, the royal we, thought of as important markets before.
Um, and, and so a premise we have as a fund is that n- there's actually much more instinct for creativity, um, visual creativity, audio creativity, technical creativity than, um, uh, like, there's latent demand for it. And AI applications can really serve that.
Uh, I think in particular, Midjourney was a company that is in the vanguard here, and nobody understood for a long time because the perhaps outside view is, like, how many people wanna generate images that are not, um, easily...
You know, they're raster, they're not easily editable. They can't be used in these professional contexts in a complete way. And the answer is, like, an awful lot, right? For a whole range of use cases, and I think we'll continue to find that, especially as the capabilities improve and, uh, we think the, the range of, um, uh, quality and, uh, controllability that you can get in these different domains is still...
It's very deep and we're still very early. Um, a- and then I, I think as, if, if we're in the first or second inning of this AI wave, um, one obvious place to go invest, uh, and to go build companies is the enabling layers, right?
Um, shorthand for this is obviously compute and data. Uh, I think the, the needs for, uh, data are largely changed now as well. You need more expert data. Um, you need different forms of data. We'll talk about that later in terms of who has, like, let's say, reasoning traces in different domains that are interesting to, uh, companies doing their own training.
But this is, this is an area that has seen explosive growth, and we continue to invest here. Um, okay. So maybe time for some opinions. Uh, uh, there was a prevailing narrative that, um, you know, some part from companies, some part from investors.
Value Debates16:05
It's a fun debate, uh, as to where is the value in the ecosystem and can there be opportunities for startups? Um, if you guys remember the phrase GPT wrapper, it was, like, the dominant phrase in the tech ecosystem for a while of...
And, and what it, what it represented was this idea that there was no value at the application layer. You had to do pre-training, and then, like, nobody's gonna catch OpenAI in pre-training. A- and you know, this isn't, this isn't, like, a, a knock on OpenAI at all.
These, these labs have done amazing work enabling the ecosystem, and we continue to partner with them and, and others. But, um, uh, but it, it's simply untrue as a narrative, right? Uh, the odds are clearly in favor of a very, very rich ecosystem of innovation.
You have a bunch of choices of models that are good at different things. Um, you have price competition. You have open source. Uh, I, I think an underappreciated impact of test time scaling is you're going to better match user value with your spend on compute.
And so if you are a new company that can figure out how to make these models useful to somebody, the customer can pay for the compute instead of you taking as a, as a startup the CapEx for pre-training or, um, or RL upfront.
Uh, and um, uh, a- a- as Pranav mentioned, you know, small models, especially if you know the domain, can be unreasonably effective. Uh, and the product layer has, if we look at the sort of cluster of companies that we described, shown that it is creating and capturing value and that it's actually a pretty hard thing to build great products that leverage AI.
Um, uh, so, so broadly, like, we have a point of view that I think is actually shared by many of the labs that the world is full of problems in the last mile to go take even AGI into all of those use cases is quite long.
Okay. Another prevailing belief is that, um... Or, or you know, another great debate that Shawn could host is, like, does the value go to startups or incumbents? Uh, we must admit some bias here, even though we have, uh, you know, friends and portfolio, former portfolio companies that would be considered incumbents now.
But, um- Uh, oh, sorry. Swap, swap, uh, swap views. Sorry, uh, you know, there are, there are markets in venture that have been considered traditionally, like, too hard, right? Like, just bad markets for the, the venture capital spec, which is capital efficient, rapid growth.
That's a venture backable company, um, where the end output is a, you know, a tens of billions of dollars of enterprise value company. Um, and, and these included areas like legal, healthcare, defense, pharma, education, um, uh, you know, any traditional venture firm would say, like, "Bad market.
Nobody makes money there. It's really hard to sell. There's no budget," et cetera. And, and one of the things that's interesting is if you look at the cluster of companies that has actually been effective over the past year, some of them are in these markets that were traditionally non-obvious, right?
And so perhaps one of our more optimistic views is that AI's really useful, and if you make a capability that is novel, that is several magnitudes, um, orders of magnitude cheaper, then actually you can change the buying pattern and the structure of these markets.
And maybe the legal industry didn't buy anything 'cause there wasn't anything worth buying for a really long time. That's one example. Um, we, we also think that, like, what was the last great consumer company? Um, maybe it was Discord or Roblox in terms of things that started that have just, like, really, um, enormous user bases and engagement, uh, until, you know, we had these consumer chatbots of different kinds and, and, like, the next...
perhaps the next generation of search. As Pranav mentioned, we think that the, um, uh, opportunity for social and media generation in games is, uh, large and new in a, in a totally different way. Um, and, and finally, uh, in terms of the markets that we look at, uh, I, I think there's broad recognition now that you can sell against outcomes and services rather than software spend with AI, because you're doing work versus just giving people the ability to do a workflow.
But, um, if you take that one step further, we think there's elastic demand for many services, right? Uh, our, our classic example is, um, there's on order of 20 to 25 million professional software developers in the world. Uh, you know, m- I imagine much of this audience is technical.
Uh, demand for software is not being met, right? If we take the cost of software and high-quality software down two orders of magnitude, we're just gonna end up with more software in the world. We're not gonna end up with fewer people doing development.
Um, at least that's what we would argue. Um, and then finally on the incumbent versus, uh, startup question, uh, the prevailing narrative is incumbents have the distribution, the product surfaces, and the data. Don't bother competing with them. They're gonna create and capture the value and share some of it back with their customers.
I think this is only partially true. Um, they... Incumbents have the distribution. Uh, they have always had the distribution. Like, the point of the startup is you have to go fight with a better product or a more clever product, um, and maybe a different business model to go get new distribution.
But the specifics around the product surface and the data I think are actually worth understanding. There's a really strong innovator's dilemma. If you look at the SaaS companies that are dominant, they sell by seat, and if I'm doing the work for you, I don't necessarily wanna sell you seats.
I might actually decrease the number of seats. Um, the tens, uh, the, the decades of, um, years and millions of man and woman hours of code that have been, um, written to, uh, enable a particular workflow in CRM, for example, may not matter if I don't want people to do that workflow of filling out the database every Friday anymore.
Uh, and, and so I, I do think that this sunk cost or the incumbent advantage gets highly challenged by new UX and code generation as well. Um, and then one disappointing learning that we found in our own portfolio is no one has the data we want in many cases, right?
So imagine you are trying to automate, um, a specific type of knowledge work, uh, and what you want is the reasoning trace, um, all of the inputs, and the output decision. Um, like, that sounds like a very useful set of data, and the incumbent companies in any given domain, they never save that data, right?
Like, they have a database with the outputs some of the time. A- and so I, I would say, uh, one of the things that is worth thinking through as a startup is, um, when an incumbent says they have the data, like, what is the data you actually need to make your product higher quality?
Okay, so in, in summary, um, you know, our, our shorthand for the set of changes that are happening is Software 3.0. We think it is a full stack rethinking, and it enables, um, and a, a new generation of companies to have a huge advantage.
Software 3.022:54
The speed of change, um, favors startups. If the floor is lava, it's really hard to turn a really big ship. Uh, I think that some of the CEOs of large companies now are incredibly capable, but they're still trying to make 100,000 people move very quickly in a new paradigm.
Um, the market opportunities are different, right? These markets that we think are interesting and very large, like represent a trillion dollars of value, are not just the replacement software markets of the last two decades. Um, it's not clear what the business model for many of these companies should be.
Uh, Sierra just started talking about charging for outcomes. Um, outcomes-based pricing has been this holy grail idea in software, and it's been very hard, but now we do more work. Um, uh, uh, there are other business model challenges, um, and so, you know, our companies, they spend a lot more on compute than they have in the past.
They spend a lot with the foundation model providers. They think about gross margin. Uh, they think about where to get the data. Uh, it's a time where you need to be really creative about product, um, uh, versus just replace the workflows o- of the past, uh, and it might require ripping out those workflows entirely.
It's a different development cycle. I bet most of the people in this room have written evals, um, and like compared to, you know, the academic benchmark to a real world eval and said like, you know, "That's not it," and how do I make a user, um, understand, uh, the- Um, non-deterministic nature of these outputs or gracefully fail.
I think that's in a, like, a different way to think about product than in the past. Um, and we, we need to think about infrastructure again, right? Um, there was this middle period where the cloud providers, the hyperscalers took this problem away from software developers, and it was all just gonna be like, I don't know, front end people at some point.
And it's like we are not there anymore. We're back in the hardware era where people are, um, acquiring and managing and optimizing compute, and I think that will really matter in terms of capability in companies. Um, so, uh, I, I guess we'll, we'll end with a call to action here and, and encourage all of you to seize the opportunity.
Um, it is the greatest technical and economic opportunity that we've ever seen. Like, we made a decade plus career-type bet on it. And, um, uh, we do a lot of work with the foundation model companies. Uh, we think they are doing amazing work, and they're great partners and even co-investors in some of our efforts.
But, uh, I, I think all of the focus on their interesting missions around AGI and safety, um, do not mean that there are not opportunities in other parts of the economy. The world is very large, and we think much of the value will be distributed in the world through an unbundling and eventually a rebundling, uh, as often happens in technology cycles.
Um, so we think this is a market that is structurally supportive of startups. We're really excited to try to work with the more ambitious ones. And the theme of 2024, um, to us has been like, well, thank goodness this is a, this is an ecosystem that is much friendlier to startups than 2023.
It is what we hoped. Um, and, and so, uh, you know, please, uh, ask us questions and take advantage of the opportunity.
Outro26:06
Thanks for joining our first talk from Latent Space Live at NeurIPS 2024 in Vancouver. As always, this is your AI co-host Charlie. I want to give a huge thank you to Sarah Guo and Pranav Reddy for sharing their invaluable insights on the state of AI in 2024.
Be sure to check the link in the description for their presentation slides, social media links, and additional resources. Watch out and take care. You're giving me wind-






