# Outlasting Noam Shazeer, Crowdsourcing Chai AI w/ 1.4m DAU — with William Beauchamp, Chai Research

Latent Space · 2025-01-26

<https://addtry.com/da562bdf-4a5b-4fdf-b83b-86cfc7993338>

William Beauchamp, founder-CEO of Chai AI, explains how he pivoted from algorithmic trading to build a character chatbot platform before Character.ai, growing to 1.4M DAU and $22M+ revenue by crowdsourcing model improvements through Chaiverse. Starting with GPT-J in 2021, he found product-market fit with a therapist bot and shifted to user-generated content, letting users define prompts, images, and names. Competing against well-funded rivals like Character AI and Talkie, Chai ships over 100 LLMs weekly via its developer platform, spending $10M on compute in 2024 and tripling it. Beauchamp argues AI follows an S-curve, not scaling laws, and focuses on inference optimization using rejection sampling and reward models to serve better responses. He prioritizes 'insanely great' products over technology-driven features, noting that audio and image features failed to move metrics, while a data flywheel and aggressive user acquisition drove rapid growth.

## Questions this episode answers

### How did Chai AI discover its product-market fit?

William Beauchamp initially built bots for news and recipes, but no one used them. The breakthrough came when his sister created a therapist bot; the next day, Chai had 20 active users averaging 20 minutes of engagement. That revealed users wanted judgment-free, immediate conversation with human-like AI for fun and companionship, not utilitarian tasks.

[14:10](https://addtry.com/da562bdf-4a5b-4fdf-b83b-86cfc7993338?t=850000)

### What were the key events that accelerated Chai's user growth in 2024?

Four main drivers: migrating the backend past 500K DAU after outages; Character AI's acquisition by Google, which likely reduced their ad spend and freed market share; perfecting a data flywheel that enabled evaluating 20–50 models daily, shortening feedback from 30 days to 3 hours; and hiring a ByteDance growth expert who ramped user acquisition spend to $40K/day, boosting growth from 2x to 3x.

[27:22](https://addtry.com/da562bdf-4a5b-4fdf-b83b-86cfc7993338?t=1642000)

### How does Chai's Chaiverse platform evaluate and blend AI models?

Developers submit LLMs to Chaiverse, where each model gets 5,000 completions served to users for A/B comparisons. This produces an ELO ranking; only the top 20% of models are useful. Chai blends multiple top models—for instance, serving a smart model 50% of the time and a funny model 50%—to combine strengths without expensive routing.

[46:39](https://addtry.com/da562bdf-4a5b-4fdf-b83b-86cfc7993338?t=2799000)

### Why doesn't Chai use streaming for its chat completions?

Streaming prevents rejection sampling. Chai generates 16 full completions for each message, then uses a reward model trained on 50 million user messages to select the one most likely to engage the user, delivering the entire response at once. This improves quality and allows serving larger models without latency penalties.

[1:08:41](https://addtry.com/da562bdf-4a5b-4fdf-b83b-86cfc7993338?t=4121000)

## Key moments

- **[0:00] Poker to AI**
  - [1:19] William Beauchamp accumulated $100,000 through poker to start algorithmic trading after Cambridge.
  - [3:52] At 30, William Beauchamp left his £5M/year quant trading firm to pursue AI.
  - [6:09] 'Crypto is the most fun form of gambling invented ever,' says William Beauchamp.
  - [7:38] William Beauchamp aimed to build a platform for AI, analogous to YouTube for video, to decentralize LLMs.
  - [9:09] William Beauchamp disagrees with the 'brute force' AI scaling view, citing S-curves that plateau near human performance.
- **[11:36] Finding PMF**
  - [13:56] Chai's product-market fit: William's sister's therapist bot drew 20 users spending 20 minutes on average.
  - [18:45] Chai's breakthrough: letting users create bot persona prompts with GPT-J in 2021, before it was obvious.
  - [20:22] Chai was the #1 AI app in the App Store with 100K DAU before Character AI launched.
  - [21:33] Character AI's $100M+ funding let them serve larger models, making Chai's 6B model feel dumber.
- **[22:01] Underdog Growth**
  - [24:01] Chai's average user session lasts 90 minutes with 150 messages, far exceeding typical ChatGPT usage.
  - [26:00] Chai AI reached 1.4M DAU and $22M revenue, tripling users and doubling revenue in 2024.
  - [27:26] Chai's back-end on Firebase hit scaling limits at 500K DAU, causing a 3-month growth plateau.
  - [31:05] William Beauchamp suspects Character AI's Google acquisition and reduced ad spend boosted Chai's growth.
  - [32:27] Chai's team of 5 AI researchers ships over 100 LLMs per week by tightening feedback loops.
  - [33:30] Chai reduced model evaluation from 30-day retention tests to 3-hour human feedback with 5,000 user samples.
- **[34:53] Inflection Points**
  - [34:54] Chai users flipped from preferring Character AI to Chai after October 2024, per Reddit sentiment.
  - [35:56] Hiring a former ByteDance head of growth, amazed by Chai's 1M DAU without ads, ignited hypergrowth.
  - [37:09] Chai now spends $40K/day on user acquisition, half of competitors, achieving 3x annual growth.
  - [38:10] Chai launched voice mode before Character AI but saw no retention or monetization gain, so removed it.
- **[42:10] Multimodality & UGC**
  - [43:01] The #1 problem in AI companions: all content is created by middle-aged men in Silicon Valley, not users.
  - [44:35] 'Until a team earns $100M/year creating content on Chai, we're not finished,' says William Beauchamp.
  - [46:42] Chaiverse needs only 5,000 completions to accurately rank an LLM's entertainment value via blind comparisons.
- **[46:49] Chaiverse ELO**
- **[51:58] AGI Debate**
  - [52:02] William Beauchamp predicts AGI is further away, citing S-curves and LLMs' weakness at true reasoning.
  - [53:15] William Beauchamp proposes 'super knowledge' rather than 'super intelligence' to describe AI's true strength.
- **[57:14] Evals**
  - [57:20] ELO-based human feedback serves as Chai's north star, a self-scaling alternative to static benchmarks.
- **[1:02:01] Content & Culture**
  - [1:05:45] William Beauchamp on Chai's culture: 'One person said they couldn't get work done because of PTO on Friday. What is PTO?'
  - [1:07:06] Chai AI spent $10 million on compute in 2024 and expects to triple that in 2025.
- **[1:07:20] Inference Tricks**
  - [1:08:14] Chai switched from vLLM to MK1, an inference engine with custom CUDA kernels, for much faster inference.
  - [1:08:41] Chai never streams completions because streaming prevents rejection sampling, which improves response quality.
  - [1:11:43] 'We've never had a technology before that can just make stuff,' says William Beauchamp on AI's generative uniqueness.
- **[1:11:48] Closing**

## Speakers

- **Alessio** (host)
- **Swyx** (host)
- **William Beauchamp** (guest)

## Topics

Agent Platforms, Startups, Inference

## Mentioned

Chai (company), Character AI (company), DeepSeek (company), Google (company), Hugging Face (company), LMSys (company), MK1 (company), OpenAI (company), BERT (product), Chaiverse (product), Cobalt AI (product), Firebase (product), GCP (product), GPT (product), GPT-J (product), Llama (product), Silly Tavern (product), o1 (product), vLLM (product)

## Transcript

### Poker to AI

**Alessio** [0:05]
Hey, everyone. Welcome to the Latent Space Podcast. This is Alessio, partner and CTO at Decibel, and today we're in the Chai AI office with my usual co-host, Swyx.

**Swyx** [0:14]
Hey, thanks for having us. Uh, and, um, we are-- It's rare that we get to get out the office, so, uh, thanks for inviting us to your home. Uh, we're in the office of Chai with William Beauchamp.

**William Beauchamp** [0:23]
Yeah, that's right.

**Swyx** [0:24]
You're a founder of Chai AI, but previously you... I mean, I think you're concurrently also running your, your fund.

**William Beauchamp** [0:30]
Yep. So, um, I was simultaneously running an algorithmic trading company.

**Swyx** [0:35]
Yeah.

**William Beauchamp** [0:36]
But I fortunately was able to kinda exit from that, um, I think just in Q3 last year.

**Swyx** [0:42]
Yeah.

**William Beauchamp** [0:42]
So, yeah.

**Swyx** [0:43]
Congrats.

**William Beauchamp** [0:44]
Yeah, thanks.

**Swyx** [0:44]
So Chai is-- has always been on my, my radar because, well, first of all, you do a lot of advertising, I guess, in, in the Bay Area, so it's working.

**William Beauchamp** [0:52]
Yep.

**Swyx** [0:52]
Uh . And second of all, the reason I reached out through our mutual friend Joyce was because I, I'm just generally interested in the consumer AI space, chat platforms in general. I think there's a lot of inference, uh, insights that we can get on, uh, from that, as well as human psychology insights.

**William Beauchamp** [1:08]
Mm.

**Swyx** [1:08]
Kind of like a weird blend of the two. And we also share a, share a bit of a history as former finance people tr-tra-crossing over.

**William Beauchamp** [1:15]
Yeah.

**Swyx** [1:15]
I guess we can just kinda start it off with, like, the origin story or, or the-

**William Beauchamp** [1:19]
Sure

**Swyx** [1:19]
... uh, of Chai. Like, why decide working on a consumer AI platform rather than, uh, you know, B2B SaaS?

**William Beauchamp** [1:25]
So, um, just quickly touching on the background in finance. Originally, I'm from the UK, uh, born in London, and I was fortunate enough to go study economics at Cambridge. And I graduated in 2012, and at that time, everyone in the UK and everyone on my course, HFT, quant trading-

**Swyx** [1:46]
Yeah

**William Beauchamp** [1:46]
... was really the big thing. It was, like, the big wave that was happening, so there was a lot of opportunity in that space. And throughout college, I'd sort of played poker, so I'd, uh, you know, I was-- I dabbled as a professional poker player, and I was able to accumulate this, this sort of, um, you know, say, $100,000 through playing poker.

And at the time, as my friends would go work at companies like Jane Street or Citadel, I kind of did the maths, and I just thought, "Well, maybe if I traded my own capital, I'd probably come out ahead.

I'd make more money than just going to, to work at Jane Street."

**Swyx** [2:20]
That's a 100K base as capital?

**William Beauchamp** [2:22]
Yes, yes. Yep.

**Swyx** [2:23]
That's not a lot .

**William Beauchamp** [2:24]
Well, it depends what strategies you're doing and it, you know... There is an advantage to being small, right? 'Cause there are-- If you have a 10-

**Swyx** [2:31]
Strategies that don't work at size.

**William Beauchamp** [2:32]
Exactly, exactly. So if, if you have a fund of $10 million, if you find a little anomaly in the market that you might be able to make 100K a year from, that's a 1% return on your 10 million fund.

If your fund is 100K, that's 100% return, right? So being small, in some sense, was an advantage. So it started off, and the-- taught myself Python, and machine learning was, like, the big thing as well. Machine learning had really-- It was the first big time machine learning was being used for image recognition.

Neural networks come out. You get dropout and, you know. So this, this was the big thing that's going on at the time. So I probably spent my first three years out of Cambridge just building neural networks, um, building random forests to try and predict asset prices, right?

And then trade that using my own money. And that went well. And, you know, if you, if you start something and it goes well, you try and hire more people, and the, the, you know, the first people that came to mind was the talented people I went to college with.

And so I hired some friends, um, and that went well, and hired some more. And eventually, you kinda run out of friends to hire, and so that was when I formed the company. And, you know, from that point on, you know, we had our ups and we had our downs and that, that was a whole long story and journey in itself.

But after doing that for about eight or nine years, on my 30th birthday, which was four years ago now, I kinda took a step back to s- to just, um, evaluate my life, right? This is what one does when one turns 30.

**Swyx** [4:06]
Yeah.

**William Beauchamp** [4:06]
Um-

**Swyx** [4:06]
Who knows?

**Alessio** [4:07]
I just turned 30. Yeah. I, I hear you.

**William Beauchamp** [4:09]
And, you know, I looked at my 20s and I loved it. It was a really special time. I was really lucky and fortunate to have worked with this amazing team, been successful, had a lot of hard times, and through the hard times, learnt wisdom and then a lot of success and, you know, was able to enjoy it.

And so the company was making about £5 million a year, and it was just me and a team of, say, 15, like, Oxford and Cambridge-educated mathematicians and physicists. It was, like, the real dream that you would have if you wanted to start a quant trading firm.

It was like-

**Swyx** [4:41]
Your own-- all your own money.

**William Beauchamp** [4:41]
Yeah, exactly. It was, it was all the team's own, own money. We had no customers complaining to us about issues. There was no investors, you know, saying, you know, they don't like the risks that we're taking. We could really, uh, run the thing exactly as we wanted it.

**Swyx** [4:53]
It was like Susquehanna or-

**William Beauchamp** [4:55]
Exactly

**Swyx** [4:55]
... like Rentech or...

**William Beauchamp** [4:56]
Yep, exactly, yeah. And that-- And, and they're the companies that we would kind of look towards as we were building that thing out. But on my 30th birthday, I look and I say, "Okay, great. This thing is making as much money as kind of anyone would really need."

And I thought, "Well, what's gonna happen if we keep, keep going in this direction?" And it was clear that there-- we would never have a kind of a big, big impact on the world. We could enrich ourselves. We could make really good money.

Everyone on the team would be paid very, very well. Presumably, I could make enough money to buy a yacht or something. But this stuff wasn't that important to me. And so I felt a sort of obligation that if you have this much talent, and if you have a talented team, especially as a founder, you wanna be putting all that talent towards a good use.

I looked at the time of, like, getting into crypto, and I had a really strong view on crypto, which was that as far as a gambling device, it's, like, the most fun form of gambling invented in, like, ever.

Super fun. I thought as a way to evade monetary regulations and banking restrictions, I think it's also absolutely amazing. So it has two, like, killer u- uh, use cases.

**Swyx** [6:07]
Not so much banking the unbanked.

**William Beauchamp** [6:08]
But everything else- But everything else to do with, like, the blockchain and, and, you know, web, was it Web 3.0 or web two? You know, that I-- that didn't-- it didn't really make much sense. And so instead of going into crypto, which I thought even if I was successful, I would end up in a lot of trouble, I thought maybe it'd be better to, to build something that governments wouldn't have a problem with.

I knew that LLMs were, like, a thing. I think OpenAI had said... They hadn't released GPT-3 yet, but they'd said, "GPT-3 is so powerful we can't release it to the world," or something. Or was it GPT-2? And then I started interacting with-- I think Google had open sourced some language models.

They weren't necessarily LLMs, but they, but they were, um-

**Alessio** [6:49]
BERT.

**William Beauchamp** [6:49]
Yeah, exactly. So I was able to play around with BERT. Nowadays, so many people have interacted with ChatGPT. They get it, but it's like the first time you, you, you can just talk to a computer, and it talks back.

It's kind of a special moment, and, you know, everyone who's done that goes like, "Wow, this is how it should be." Right? It should be like rather than having to type on Google and, and search, you should just be able to ask Google a question.

When I saw that, I read the literature, I kind of came across the scaling laws, and I think even four years ago, all the pieces of the puzzle were there, right? Google had done this amazing research and published, you know, a lot of it.

OpenAI was still open, and so they'd published a lot of their research. And so you really could be fully informed on, on the state of AI and where it was going. And so at that point, I was confident enough it was worth a shot.

I think LLMs are gonna be the next big thing, and so I-- that's the thing I wanna be building in, in that space. And I thought, "What's the most impactful product I can possibly build?" And I thought it should be a platform.

So I myself love platforms. I think they're fantastic because they open up an ecosystem where anyone can contribute to it, right? So if you think of a platform like a YouTube, instead of it being like a Hollywood situation where you have to-- if you wanna make a TV show, you have to convince Disney to give you the money to produce it.

Instead, anyone in the world can post any content they want to YouTube, and if people wanna view it, the algorithm is gonna promote it. Nowadays, you can look at creators like MrBeast or Joe Rogan. They would have never have had that opportunity unless it was for this platform.

Other ones, like Twitter's a great one, right? But I would consider Wikipedia to be a platform, where instead of the Britannica Encyclopedia, which is this-- it's like a, a monolithic... You get all the, the researchers together, you get all the data together, and you combine it in this, in this one monolithic source.

Instead, you have this distributed thing. You can say anyone can host their content on Wikipedia, anyone can contribute to it, and anyone can-- maybe their contribution is they delete stuff. When I was hearing, like, the kind of the Sam Altman and kind of the, the Muskian perspective of AI, it was a very kind of monolithic thing.

It was all about AI is, basically, a single thing, which is intelligence, and the more data, the more intelligent; the more compute, the more intelligent; and the more and better AI researchers, the more intelligent, right? They would speak about it as a kind of a race to, like, who can get the most data, the most compute, and the most researchers, and that would end up with the most intelligent AI.

But I didn't believe in any of that. I thought that's, like, the total, like... I thought that perspective is the perspective of someone who's never actually done machine learning. Because with machine learning, first of all, you see the, the performance of the models follows an S curve.

So it's not like it just goes off to infinity, right?

**Alessio** [9:48]
Mm-hmm.

**William Beauchamp** [9:48]
And the, the S curve, it kind of plateaus around human-level performance. And you can look at all the, all the machine learning that was going on in the twenty tens, everything kind of plateaued around the human-level performance. And we can think about the self-driving car promises, you know, how Elon Musk kept saying the self-driving car is gonna happen next year, it's gonna happen next next year.

Or you can look at the image recognition, the speech recognition. You can look at all of these things. There was almost nothing that went superhuman except for something like AlphaGo, and we can speak about why AlphaGo was able to go, like, su- uh, superhuman.

So I thought the most likely thing was gonna be this. I thought, it's not gonna be a monolithic thing that's like an Encyclopedia Britannica. I thought it must be a distributed thing. And I actually liked to look at the world of finance for what I think a mature machine learning ecosystem would look like.

So finance is a machine learning ecosystem because all of these quant trading firms are running machine learning algorithms, but they're running it on a centralized platform like a marketplace. And it's not the case that there's one giant quant trading company with all the data and all the, the quant researchers and all the algorithms and compute.

But instead, they all specialize. So one will specialize on high-frequency trading, another will specialize on mid frequency, another one will specialize on equities, another one will specialize... And I thought that's the way the world works. That's how it is.

And so there must exist a platform where a small team can produce an AI for a unique purpose, and they can iterate and build the best thing for that, right? And so that was the vision for Chai. So we wanted to build a platform for LLMs.

### Finding PMF

**Alessio** [11:36]
That's kind of the maybe inside versus contrarian view-

**William Beauchamp** [11:40]
Mm-hmm

**Alessio** [11:40]
... that led you to start the company.

**William Beauchamp** [11:42]
Yep.

**Alessio** [11:42]
Then what was maybe the initial idea maze? Because, you know, if you think about-- If somebody told you that was the Hugging Face founding story-

**William Beauchamp** [11:48]
Mm-hmm

**Alessio** [11:48]
... people might believe it. It's kinda like a similar-

**William Beauchamp** [11:51]
Yep

**Alessio** [11:51]
... ethos behind it.

**William Beauchamp** [11:52]
Yep.

**Alessio** [11:52]
How did you land on the product featured today? And, like, maybe what were some of the ideas that you discarded that initially you thought about?

**William Beauchamp** [11:59]
So the first thing we built, it was fundamentally an, an API. So nowadays, people would describe it as, like, agents, right? But anyone could write a Python script, they could submit it to the Chai back end, and we would then host this code and execute it.

So that's like the developer side of the platform. On their Python script, the interface was essentially text in and text out. An example would be the very first bot that I created, I think it was a, like, a Reddit news bot, and so it would-- first, it would pull the popular news, then it would prompt whatever...

Like, I just used some external API for, like, Bur or GPT-2 or, like, it was a very, very small thing, and then the user could talk to it. So you could say to the bot, "Hi, bot, what's the news today?"

And it would say, "This is the top stories," and you could chat with it. Now, four years later, that's like Perplexity or something. That's like the-- right?

**Alessio** [12:56]
Mm.

**William Beauchamp** [12:56]
But back then, the models were, first of all, like, really, really dumb. You know, they had an IQ of, like, a four-year-old. And users-- There, there really wasn't any demand or any PMF for interacting with the, with the news.

So then I was like, "Okay, um, clearly no PMF for that, so let's make another one." And I made a bot which was like you could talk to it about a recipe. So you could say, "I'm making eggs," like, "I've got eggs in my fridge.

What should I cook?" And it'll say, "You should make an omelet," right? There was no PMF for that. No one used it. And so I just kept creating bots, and so every single night after work, I'd be like: Okay, I-- like, we have AI, we have this platform.

I can create any text in, text out sort of agent and put it on the platform. And so we just create stuff night after night, and then all the coders I knew, I would say to them, "Look, there's this platform.

You can create any, like, chat AI. You should put it on." And, you know, everyone's like, "Will, chatbots are super lame. We want absolutely nothing to do with your chatbot app." No one who knew Python wanted to, to build on it.

I'm, like, trying to build all these bots, and no consumers wanna talk to any of them. And then my sister, who at the time was, like, just finishing college or something, I said to her, I was like, "If you wanna learn Python, you should just submit a bot for my platform."

And she, she built a therapist bot, and then the next day, I checked the performance of the app, and I'm like, "Oh my God, we've got twenty active users," and they spent, they spent, like, an average of twenty minutes on the app.

And I was like, "Oh my God, what, what bot were they speaking to for an average of twenty minutes?" And I looked, and it was the therapist bot. And I went, "Oh, this is where the PMF is." There was no demand for, for recipe help.

There was no demand for news. There was no demand for dad jokes or pub quiz or fun facts or... What they wanted was they wanted the therapist bot. At the time, I kind of reflected on that, and I thought, "Well, if I want to consume news, the most fun thing, the most fun way to consume news is, like, Twitter."

It's not like-- The value of there being a back and forth wasn't that high.

**Alessio** [14:53]
Mm.

**William Beauchamp** [14:54]
Right? And I thought, "If I need help with a recipe, I actually just go, like, The New York Times has a good recipe section." Right? It's not actually that hard. And so I just thought the thing that AI is 10X better at is a sort of a, a conversation, right, that's not intrinsically informative, but is more about an opportunity.

You can say whatever you want. You're not gonna get judged. If it's three AM, you don't have to wait for your friend to text back. It's like-- it's immediate. They're gonna reply immediately. You can say whatever you want.

It's judgment free, and it's much more like a playground. It's much more like a fun experience. And you could see that if the AI gave a person a compliment, they would love it. It's much easier to get the AI to give you a compliment than a human.

From that day on, I said, "Okay, I get it. Humans wanna speak to, like, humans or human-like entities, and they wanna have fun." And that was when I started to look less at platforms like Google, and I started to look more at platforms like Ins- uh, Instagram.

And I was trying to think about: Why do people use In- uh, Instagram? And I could see that I think Chai was, was filling the same desire or the same drive. If you go on Instagram, typically, you wanna look at the faces of other humans, or you wanna hear about other people's lives.

So if it's like The Rock is making himself pancakes on a cheat day, you kind of feel a little bit like you're The Rock's friend or you're, like, having pancakes with him or something, right? But if you do it too much, you feel like you're sad and, like, a lonely person.

But with AI, you can talk to it and tell it stories, and it tell you stories, and you can play with it for as long as you want, and you don't feel like you're, like, a sad and lonely person.

You feel like you, you actually have a friend.

**Alessio** [16:30]
And why, why is that? Do you have any insight on that-

**William Beauchamp** [16:33]
I think it's just from-

**Alessio** [16:33]
Like, from using it

**William Beauchamp** [16:34]
... human psychology. I think it's just the, the idea that with old school social media, you're just consuming passively, right? So you'll just swipe. If I'm watching TikTok, just like swipe and swipe and swipe. And even though I'm getting the dopamine of, like, watching an engaging video, there's this other thing that's building in my head which, like, I'm feeling lazier and lazier and lazier.

And after a certain period of time, I'm like, "Man, I just wasted forty minutes. I achieved nothing." But with AI, because you're interacting, you feel like you're-- It's not like work, but you feel like you're participating and contributing to the thing.

You don't feel like you're just consuming, so you don't have a sense of remorse, basically. And you know, I think on the whole, people-- the way people talk about Chai and, and interact with the AI, they speak about in a incredibly positive sense.

Like, we get people who say they have eating disorders saying that the AI helps them with their eating disorders. People who say they're depressed, it helps them through, like, the rough patches. So I think there's something intrinsically healthy about interacting that TikTok and Instagram and YouTube doesn't quite tick.

From that point on, it was about building more and more kind of like human-centric AI for people to interact with. And I was like, "Okay, let's make a, a Kanye West bot," right?

**Alessio** [17:51]
Mm.

**William Beauchamp** [17:51]
And then no one wanted to talk to the Kanye West bot. And I was like, "Oh, who's like a cool persona for teenagers to wanna interact with?" And I was like-- I was trying to find the influencers and stuff like that, but no one cared.

Like, they didn't wanna interact with the in-

**Alessio** [18:05]
No one.

**William Beauchamp** [18:06]
Yeah. And it-- Instead, it was really just the special moment was when we said the realization that developers and software engineers aren't interested in building this sort of AI, but the consumers are. Right? And rather than me trying to guess every day, like, what's the right bot to submit to the platform, why don't we just create the tools for the users to build it themselves?

And so nowadays, this is, like, the most obvious thing in the world. But when Chai first did it, it was not an obvious thing at all, right? And so we took an API for, let's just say it was-- I think it was GPT-J, which was this six billion parameter open source transformer-style LLM.

We took GPT-J, we let users create the prompt, we let users select the image, and we let users choose the name, and then that was the bot. And through that, they could shape the experience, right? So if they said, "This bot's gonna be really mean, and it's gonna be called, like, Bully in the Playground," right?

That was like a whole category that I, I never would've guessed, right? People love to fight. They love to have a disagreement, right? And then they would create-- There'd be all these romantic archetypes that I didn't know existed.

And so as the users could create the content that they wanted, that was when Chai was able to, to get this huge variety of content. And rather than appealing to, you know, one percent of the population that I'd figured out what they wanted, you could a-appeal to a much, much broader thing.

And so from that moment on, it was very, very crystal clear. It's like Chai-- Just as Instagram is this social media platform that lets people create images and upload images, videos, and upload that, uh, f- Chai was really about how can we let the users create this experience in AI and then share it and interact and search.

So it's really, you know, I say it's like a platform for social AI.

**Alessio** [20:01]
Where did the Chai name come from?

**William Beauchamp** [20:03]
Yep.

**Alessio** [20:03]
Because you started the same time as Character AI-

**William Beauchamp** [20:05]
Chat AI

**Alessio** [20:05]
... so-

**William Beauchamp** [20:06]
Chat AI.

**Alessio** [20:07]
Oh, okay.

**William Beauchamp** [20:08]
Yeah.

**Alessio** [20:08]
I was like, "Is it Character AI shortened?" Uh, you started at the same time-

**Swyx** [20:11]
Maybe he likes the tea

**Alessio** [20:12]
... so I was curious.

**William Beauchamp** [20:12]
No, no, no.

**Swyx** [20:13]
He likes Chai tea.

**William Beauchamp** [20:13]
We started-

**Alessio** [20:13]
Well, that's the, the UK origin was, like, the second, the Chai.

**William Beauchamp** [20:16]
We started way before Character AI, and there's an interesting story that, um-

**Alessio** [20:22]
Yeah, yeah, yeah

**William Beauchamp** [20:22]
... Chai's, um, Chai's numbers were very, very strong, right? So I think in even '20-- I think late 2022. Was it late 2022? Or maybe early 2023. Chai was, like, the number one AI app in the App Store, so we, we would have something like 100,000 daily active users.

And then one day we kind of saw there was this website, and we were like, "Oh, this website looks just like Chai," and it was the Character AI, uh, website. And I think that nowadays it's-- I think it's much more common knowledge that when they left Google with the funding, I think they knew what was the most trending, the number one app, and I think they sort of, um-

**Alessio** [21:03]
Oh, okay

**William Beauchamp** [21:04]
... built their-

**Swyx** [21:04]
You found the PMF for them.

**William Beauchamp** [21:05]
We found the PMF for them. Exactly. Yeah. So I worked a year very, very hard, and then they-- And then that was when I learnt a lesson, which is that if you're VC-backed and if, you know... So Chai, we'd kinda ran-- We'd got to this point, I was the only person who'd invested.

I'd invested maybe £2 million in the business, and w- you know, from that, we were able to build this thing, get to, say, 100,000 daily active users. And then when Character AI, AI came along, the first version, we sort of laughed.

We were like, "Oh, man, this thing sucks." Like, "They don't know what they're building. They're building the wrong thing anyway." But then I saw, oh, they've raised $100 million. Oh, they've raised another $100 million. And then our users started saying, "Oh, guys, your AI sucks."

'Cause we were serving a six billion parameter model, right? How big was the model that Character AI could afford to serve, right? So we would be spending-- Let's say we would spend a dollar per, per user, right? Um, over the, the, you know, the entire lifetime-

### Underdog Growth

**Alessio** [22:03]
Dollar per, per session? Per chat? Per-

**William Beauchamp** [22:04]
No, no, no, no

**Alessio** [22:04]
... per month?

**William Beauchamp** [22:04]
Um, let's say we'd get-- Over the course of a year, we'd have a million users, and we'd spend a million dollars on the AI throughout a year.

**Swyx** [22:11]
See, yeah, yeah.

**William Beauchamp** [22:11]
Right.

**Swyx** [22:12]
Like aggregated.

**William Beauchamp** [22:12]
Exactly. Exactly. Right. They could spend 100 times that. So people would say, "Why is your AI much dumber than Character AI's?" And then I was like, "Oh, okay, I get it. This is like the Silicon Valley style, um, hyperscale business."

And so yeah, we moved to Silicon Valley and, uh, got some funding and iterated and built the flywheels and, um, yeah, I, I'm very proud that we were able to compete with that, right? So-- And I think the reason we were able to do it was just customer obsession.

And it's similar, I guess, to how DeepSeek have been able to produce such a compelling model when compared to someone like an OpenAI, right? So DeepSeek, you know, their latest, um, V2-

**Swyx** [22:54]
V3, yeah

**William Beauchamp** [22:55]
... they-- Yeah. They claim to have spent five million training it. It may be a bit more, but yeah. Um-

**Swyx** [23:01]
Like, why are you making such a big deal out of this?

**William Beauchamp** [23:03]
Yeah.

**Swyx** [23:03]
There's an agenda there.

**William Beauchamp** [23:04]
Yeah.

**Swyx** [23:05]
You specific-- You brought up DeepSeek, so we have to ask. You had a call with them.

**William Beauchamp** [23:08]
We did. We did. We did. Um, let me think what to say about that. I think for one, they have an amazing story, right? So their background is, again, in finance.

**Swyx** [23:17]
They're, they're the Chinese version of you.

**William Beauchamp** [23:18]
Exactly. Well, there's a lot of similarities. Yes. Yes. I have a great affinity for companies which are like, um, founder-led, customer-obsessed, and just try and build something great. And I think what DeepSeek have achieved that's quite special is they've got this amazing inference engine.

They've been able to reduce the size of the KV cache significantly, and then by being able to do that, they're able to significantly reduce their inference costs. And I think with, kind of with AI, people get really focused on, like, the kind of the foundation model or, like, the model itself, and they sort of don't pay much attention to the inference.

To give you an example, with Chai, let's say a typical user session is 90 minutes, which is like f-- you know, is, is very, very long. For comparison, let's say the average session length on TikTok is 70 minutes, so people are spending a lot of time.

And in that time, they're able to send, say, 150 messages. That's a lot of completions, right? It's quite different from an OpenAI scenario where people might come in, they'll have a particular question in mind, and they'll ask, like, one question and a few follow-up questions, right?

So because they're consuming, say, 30 times as many requests for a chat or a conversational experience, you've got to figure out how to, how to get the right balance between the cost of that and the quality. And so, you know, I think with AI, it's always been the case that if you want a better experience, you can throw compute at the problem, right?

So if you want a better model, you can just make it bigger. If you want it to remember better, give it a longer context. And now, uh, what OpenAI is doing to, um, great fanfare is with rejection sampling, you can generate many candidates, right?

And then with some sort of reward model or some sort of scoring system, you can serve the most promising of these many candidates, and so that's kind of scaling up on the inference time compute side of things. And so for us, it doesn't make sense to think of AI as just the absolute performance.

So if you look at, like, the MMLU score or the, you know, any of these, um, benchmarks that people like to look at, if you just get that score, it doesn't really tell, tell you anything 'cause it's really like progress is made by improving the performance per dollar.

And so I think that's an area where DeepSeek have been able to perform very, very well, um, surprisingly so. And so I'm very interested in what Llama 4 is gonna look like and if they're able to sort of match what DeepSeek have been able to achieve with this performance per dollar gain.

**Alessio** [26:00]
Before we go into the inference, some of the deeper stuff, can you give people overview of, like, some of the numbers?

**William Beauchamp** [26:06]
Mm-hmm.

**Alessio** [26:06]
So I think last I checked-

**William Beauchamp** [26:07]
Yeah

**Alessio** [26:07]
... you had, like, 1.4 million daily active now.

**William Beauchamp** [26:10]
Yeah.

**Alessio** [26:10]
It's like over 22 million of revenue, so it's a-- quite a business.

**William Beauchamp** [26:14]
Yeah. Um, I think we grew by a factor of, um... You know, users grew by a factor of three last year. Revenue o- over doubled. You know, it's very exciting. We're competing with some really big, really well-funded companies.

Character.ai got this, I think it was almost a $3 billion valuation, and they have 5 million DAU, is the number that I last heard. Talkie, which is a Chinese-built app owned by a company called Minimax, they're incredibly well-funded, and these, these companies didn't grow by a factor of three last year, right?

And so when you've got this company and this team that's able to keep building something that gets users excited and they wanna tell their friend about it, and then they wanna come and they wanna stick on the platform, I think that's very special.

And so, um, last year was a, a great year for the team. And yeah, I think the numbers reflect the hard work that we put in, and then fundamentally, the quality of the AI is the quality of the experience that you have.

**Swyx** [27:16]
You actually publish your DAU growth chart, which is-

**William Beauchamp** [27:19]
Yep

**Swyx** [27:19]
... unusual.

**William Beauchamp** [27:20]
Yep.

**Swyx** [27:20]
Uh, and I see some inflections.

**William Beauchamp** [27:22]
Yep.

**Swyx** [27:22]
Like, it's not just a straight line.

**William Beauchamp** [27:23]
Yes. Yep.

**Swyx** [27:23]
There's some things that actually inflect.

**William Beauchamp** [27:25]
Yes.

**Swyx** [27:25]
What were the big ones?

**William Beauchamp** [27:26]
Cool. That's a great, great, great question. Let me think of, uh, a good answer.

**Swyx** [27:30]
I'm basically looking to annotate this chart which doesn't have annotations on it.

**William Beauchamp** [27:33]
Cool. The first thing I would say is this, is I think the most important thing to know about success is that success is born out of failures, right? 'Cause it's only really through failures that we learn. You know, if you think something's a good idea and you do it and it works, great, but you didn't actually learn anything 'cause everything went exactly as you imagined.

But, uh, if you have an idea and you think it's gonna be good, you try it and it fails, there's a gap between the reality and the expectation, and that's an opportunity to learn. The flat periods, that's us learning.

**Swyx** [28:05]
Mm.

**William Beauchamp** [28:05]
And then the, the, the... And then the up periods is that's us reaping the rewards of that. So I think the big-- of the growth chart of just 2024, I think the first thing that really kind of put a dent in our growth was our back end.

So we just reached the scale. So we'd-- from day one, we'd built on top of Google's GCP, which is Google's, uh, cloud, cloud platform, and they were fantastic. We used them when we had one daily active user, and they worked f- pretty good all the way up till we had about 500,000.

It was never the cheapest, but it-- from an engineering perspective, man, that thing scaled insanely good.

**Swyx** [28:44]
Like, not Vertex?

**William Beauchamp** [28:46]
Not Vertex.

**Swyx** [28:47]
Like GKE, that kind of stuff?

**William Beauchamp** [28:48]
We use Firebase.

**Swyx** [28:50]
Oh, okay.

**William Beauchamp** [28:50]
So we use Firebase. I'm pretty sure we're the biggest user ever on Firebase.

**Swyx** [28:55]
That's expensive.

**William Beauchamp** [28:56]
Yeah. We had, we had calls with engineers and they're like, "We wouldn't recommend using this product beyond this point, and you're three X over that." Um, so we, we pushed Google to, like, their absolute limits. You know, it was, uh, it was fantastic for us 'cause we could focus on the AI.

We could focus on just adding as much value as possible. But then what happened was, after 500,000, just the thing-- the way we were using it, and it would just-- it wouldn't scale any further. And so we had a really, really painful at least three-month period as we kind of migrated between different services, figuring out, like, what requests do we wanna keep on Firebase and what ones do we wanna move onto something else?

And then, you know, making mistakes and learning things the hard way. And then after about three months, we got that right, so that we would then be able to scale to the 1.5 million DAU without any further issues from the GCP.

But what hap- what happens is, is if you have an outage, new users who go on your app experience a dysfunctional app, and then they're going to exit. And so your next day, the key metrics that the app stores track are gonna be something like, um, retention rates, money spent, and the star, like the rating that they give you.

**Swyx** [30:14]
In the app store.

**William Beauchamp** [30:14]
So in the app store, yeah. So-

**Swyx** [30:16]
Tyranny

**William Beauchamp** [30:16]
... so if you're ranked top 50 in entertainment, you're going to acquire a, a certain rate of users organically. If you go in and have a bad experience, it's gonna tank Where you're positioned in the algorithm, and then it can take a long time to kind of earn your way back up, at least if you wanted to do it organically.

If you throw money at it, you can jump to the top, and I can talk about that. But broadly speaking, if we look at 2024, the first kink in the graph was outages due to hitting 500K DAU. The backend didn't want to scale past that, so then we, we just had to do the engineering and build through it.

Okay, so we built through that, and then we get a little bit of growth, and so, okay, that's feeling a little bit good. I think the next thing, I think it's, um... I'm not gonna lie, I have a feeling that when Character AI got-

**Swyx** [31:05]
I was thinking

**William Beauchamp** [31:05]
... I think so. I think-- So the Character AI team fundamentally got acquired by Google. And I don't know what they changed in their business. I don't know if they dialed down their ad spend, their user acquisition.

**Swyx** [31:17]
Product didn't change, right? Product's just what it is.

**William Beauchamp** [31:18]
I don't think so, yeah. I thi- I think the product is what it is.

**Swyx** [31:20]
It's like maintenance mode.

**William Beauchamp** [31:21]
Yes. But I think the, the issue that people... You know, some people may think this is an obvious fact, but running a business can be very competitive, right? Because other businesses can see what you're doing, and they can imitate you, and then there's this question of if you've got one company that's spending $100,000 a day on advertising, and you've got another company that's spending zero, if you consider market share, and if you're considering new users which are entering the market, the guy that's spending 100,000 a day is gonna be getting 90% of those new users.

And so I have a suspicion that when the founders of Character AI left, they dialed down their spending on user acquisition, and I think that kind of gave oxygen to, like, the, the other apps. And so Chai was able to then start growing again in a really healthy fashion.

Um, I think that's kind of like the second thing. I think a third thing is we've really built a great data flywheel. Like, the AI team sort of perfected their flywheel, I would say, in Q-- end of Q2, and I could speak about that at length.

But fundamentally, the way I would describe it is when you're building anything in life, you need to be able to evaluate it, and through evaluations, you can iterate. We can look at benchmarks, and we can say the issues with benchmarks and why they may not generalize as well as one would hope and the challenges of working with them, but something that works incredibly well is getting h- uh, feedback from humans.

And so we built this thing where anyone can submit a model to our developer backend, and it gets put in front of 5,000 users, and the users can rate it, and we can then have a really accurate ranking of, like, which models are users finding more engaging or more entertaining.

And it gets... You know, it's at this point now where every day we're able to... I mean, we evaluate between 20 and 50 models, LLMs every single day, right? So even though we've got, only got a team of, say, five AI researchers, they're able to iterate a huge quantity of LLMs, right?

So our team ships, let's just say minimum 100 LLMs a week is what we're able to iterate through. Now, before that moment in time, we might iterate through three a week. We m- you know, there was a time when even doing, like, five a month was a challenge, right?

**Swyx** [33:47]
Mm-hmm.

**William Beauchamp** [33:47]
By being able to change the feedback loops to the point where it's not let's launch these three models, let's do an A/B test, let's reassign, let's do different cohorts, let's wait 30 days to see what the day 30 retention is, which is the, the, kind of the...

If you're doing an app, that's like A/B testing 101, would be do a 30-day retention test, assign different treatments to different cohorts, and come back in 30 days. So that's insanely slow. That's just, it's too slow. And so we were able to get that 30-day feedback loop all the way down to something like three hours.

And when we did that, we could really, really, really perfect techniques like DPO, fine-tuning, prompt engineering, blending, rejection sampling, training a reward model, right? Really successively, like boom, boom, boom, boom, boom. And so I think in Q3 and Q4 we got-- The amount of AI improvements we got was, like, astounding.

It was sort-- It was getting to the point I thought, like, how much more, how much more edge is there to be had here? But the team could-- just could, could keep going and going and going. That was, like, number three for the inflection point.

### Inflection Points

**Swyx** [34:54]
There's a fourth?

**William Beauchamp** [34:54]
The important thing about the third one is if you go on our Reddit or you, you talk to users of, of AI, there's, like, a clear date. It's, like, somewhere in October or something. The users, they flipped. They-- Before October, the users would say, "Character AI is better than you," for the most part, to then from October onwards, they would say, "Y- wow, you guys are better than Character AI."

And that, that, that was, like, a really clear, positive signal that we'd, we'd sort of done it. And I think people, you can't, um, cheat consumers. You can't trick them. You can't bullshit them. They know, right? They-- If you're gonna spend 90 minutes, right, on a platform...

And with apps, there's-- the barriers to switching is pretty low. Like, you can try Character AI for a day. If you get bored, you can try Chai. If you get bored of Chai, you can go back to Character.

So the users, the loyalty is not strong, right? What keeps them on the app is the experience. If you deliver a better experience, they're gonna stay, and they can tell. So that was the third. The fourth one was we were fortunate enough to get this hire.

He was-- I hired one really talented engineer, and then they said, "Oh, at my last company, we had a head of growth. He was really, really good, and he was the head of growth for ByteDance for two years.

Would you like to speak to him?" And I was like, "Um, yes. Yes. Do you know what? Yes, I think I would." Um, and so I spoke to him, and he just blew me away with what he knew about user acquisition.

You know, it was like a 3D chess sort of thing. You know, as much as, as I know about AI-

**Swyx** [36:25]
Like ByteDance as in TikTok US?

**William Beauchamp** [36:27]
Yes. Yes.

**Swyx** [36:27]
Not-

**William Beauchamp** [36:28]
Yes.

**Swyx** [36:28]
ByteDance has other stuff.

**William Beauchamp** [36:29]
Yep. He was interviewing us as we were interviewing him, right? And so-

**Swyx** [36:33]
He has his pick of options.

**William Beauchamp** [36:34]
Yeah, exactly. And so he was kind of looking at our metrics, and he was like- I saw him get really excited when he said, "Guys, you've got a million daily active users, and you've done no advertising." And I said, "Correct."

And he, he was like, "That's unheard of." He's like, "I've never heard of anyone doing that." And then he started looking at our metrics, and he was like, "If you've got all of this organically, if you start spending money, this is gonna be very exciting."

I was like, "Let's give it a go." So then he came in, we've just started ramping up the user acquisition, so that looks like spending, you know, let's say we're spending-- we started off spending $10,000 a day. It looked very promising.

Then 20,000. Right now we're spending $40,000 a day on user acquisition. That's still only half of what, like, Character AI or Talky may be spending. But from that y- it's, it's sort of -- we were growing at a rate of maybe, say, two X a year.

And that, that got us growing at a rate of three X a year. So I'm b- I'm growing, I'm evolving more and more to, like, a Silicon Valley style hypergrowth. Like, you know, you build something decent, and then you can s-slap on a huge of-

**Swyx** [37:35]
You, you did the important thing. You built the product first.

**William Beauchamp** [37:37]
Of course. But, but, but then you can slap on, like, like, the rocket or the jet engine or something, which is just this cash in. You pour in as much cash, you buy a lot of ads, and your growth is faster.

**Swyx** [37:49]
Not to, you know... I'm just kinda curious what, what's working right now versus what surprisingly doesn't work.

**William Beauchamp** [37:54]
Oh, there's a long, long list of surprising stuff that doesn't work. Yeah.

**Swyx** [37:57]
Yeah.

**William Beauchamp** [37:57]
The surprising thing, like, the most surprising thing what doesn't work is almost everything doesn't work. That's what's surprising. And I'll give you an example. So about a year and a half ago, we were super excited by audio. I was like, "Audio's gonna be the next killer feature.

We have to get it in the app, and I wanna be the first." So everything Chai does, I want us to be the first. We may not be the, the company that's strongest at execution, but we can always be the most innovative.

**Swyx** [38:24]
Interesting.

**William Beauchamp** [38:24]
Right? So we can-

**Swyx** [38:25]
And you're pretty strong at execution.

**William Beauchamp** [38:27]
We're much stronger in- we're much stronger... A lot of the, a lot of the reason we're here is because we were first.

**Swyx** [38:32]
Huh.

**William Beauchamp** [38:32]
If we launched today, it'd be so hard to get the traction-

**Swyx** [38:35]
Okay

**William Beauchamp** [38:35]
... because it's like, to get the flywheel, to get the users, to get the-- to build a product people are excited about, if you're first, people are naturally excited about it.

**Swyx** [38:44]
Yeah, novelty.

**William Beauchamp** [38:44]
But if you're fifth or tenth, man, you've, you've got to be insanely good at execution.

**Swyx** [38:48]
So you were first with voice?

**William Beauchamp** [38:49]
We were first. We were first. Like-

**Swyx** [38:51]
I only know Character. When Character launched voice, that, that made a noise.

**William Beauchamp** [38:54]
They launched it, I think they launched it at least nine months after us.

**Swyx** [38:57]
Okay.

**William Beauchamp** [38:57]
Okay. But we laun- ... I mean, the team worked so, so hard for it. Like, at the time we did it, latency is a huge problem. Cost is a huge problem. Getting the right quality of the voice is a huge problem, right?

Then there's this user interface and getting the right user experience, because you don't just want it to start blurting out, right? You want to kind of activate it.

**Swyx** [39:20]
Mm.

**William Beauchamp** [39:21]
But then you don't want to have to keep pressing a button every single time. There's a lot that goes into getting a really smooth audio experience. So we went ahead, we invested the three months, we built it all, and then when we did the AB test, there was, like, no change in any of the numbers.

And I was like, "This can't be right. There must be a bug." And we spent, like, a week just checking everything, checking it again, checking it again, and it was like the users just did not care. And it was something like only 10 or 15% of users even clicked the button to, like, they wanted to engage the audio, and they would only use it for 10 or 15% of the time.

So if it, if you do the maths, if it's just, like, something that one in seven people use it for one seventh of their time, you've changed, like, 2% of the experience. So even if that, that 2% of the time is, like, m- insanely good, it doesn't translate much when you look at the retention-

**Swyx** [40:12]
Mm

**William Beauchamp** [40:12]
... when you look at the engagement, and when you look at the monetization rates. So audio did not have a big impact.

**Swyx** [40:18]
Notable. I, I'm pretty big on audio, but, uh, yeah. It's hard.

**William Beauchamp** [40:21]
Yeah, I like it too, but it's, you know... So a lot of the stuff which I do, I'm a big you can have a theory, you know-

**Swyx** [40:26]
Empiricist, yeah.

**William Beauchamp** [40:27]
Y- exactly, exactly. So I think if you want to make audio work, it has to be a unique, compelling, exciting experience that they can't have anywhere else.

**Swyx** [40:37]
Yeah. It could be your models were just, weren't good enough.

**William Beauchamp** [40:40]
No, no, no. They were great.

**Swyx** [40:41]
Oh, yeah?

**William Beauchamp** [40:42]
They were very good.

**Swyx** [40:43]
Okay.

**William Beauchamp** [40:43]
But it was like, it was kinda like just the, you know, if you listen to, like, an Audible or a Kindle or sort of like, you just hear this voice, and it's like, you don't go like, "Wow, this is, this is special," right?

It's like a convenience thing.

**Swyx** [40:55]
Yeah.

**William Beauchamp** [40:56]
But, um, the idea is that if you can-- if Chai's the only platform... Like, let's say you have a MrBeast, and YouTube is the only platform you can watch a MrBeast video, and it's the most engaging, fun video that you want to watch.

You'll go to a YouTube. And so it's like for audio, you can't just put the audio on there, and people go, "Oh, yeah, it's, like, 2% better." Or, like, 5% of users think it's 20% better, right? It has to be something that the majority of people for the majority of the experience go like, "Wow, this is a big deal."

That's the features you need to be shipping. If it's not gonna appeal to the majority of people for the majority of the experience, and it's not a big deal, it's not gonna move the needle.

**Swyx** [41:34]
Cool. So you killed it.

**William Beauchamp** [41:35]
Yeah, okay.

**Swyx** [41:36]
I don't see it anymore.

**William Beauchamp** [41:37]
Exactly, yeah. So I love this. The longer-- This is-- It's kinda cheesy, I guess, but the longer I've been working at Chai, and I think the team agrees with this, all the platitudes, at least I thought they were platitudes, that you would get from, like, the Steve Jobs, which is like, "Build something insanely great," right?

Or, um, "Be maniacally focused," or, you know, "The most important thing is saying no to, not to work on." All of these sort of lessons, they just are, like, painfully true. They're painfully true. So now I'm just like, everything I say, I'm either quoting Steve Jobs or Zuckerberg.

I'm like, "Guys, move fast and break things."

### Multimodality & UGC

**Swyx** [42:11]
You've, you've joined the Palo Alto Kool-Aid now.

**William Beauchamp** [42:12]
Yeah, it's just so- Everything they said is so, so true.

**Swyx** [42:15]
It's in the water. You need a turtleneck.

**William Beauchamp** [42:17]
Yeah. Yeah.

**Swyx** [42:18]
Uh, yeah. Uh-

**William Beauchamp** [42:18]
Everything is so true.

**Swyx** [42:19]
This last question from, on my side, and I want to pass this to Alessio-

**William Beauchamp** [42:22]
Mm-hmm

**Swyx** [42:22]
... is on just m- just multimodality in general.

**William Beauchamp** [42:25]
Yeah.

**Swyx** [42:25]
This actually comes from Justine Moore from, uh-

**William Beauchamp** [42:27]
Yeah

**Swyx** [42:27]
... A16Z, who's a friend of ours. And a lot of people are trying to do voice image video for AI companions.

**William Beauchamp** [42:32]
Yes.

**Swyx** [42:32]
And you just said voice didn't work.

**William Beauchamp** [42:33]
Yep.

**Swyx** [42:34]
What would make you revisit?

**William Beauchamp** [42:37]
Steve Jobs is very-- Listen, he was very, very clear on this. There's a habit of engineers who, once they've got some cool technology, they wanna find a way to package up the cool technology and sell it to consumers.

**Alessio** [42:50]
Yeah.

**William Beauchamp** [42:50]
Right? It-- That does not work. So you're free to try and build a startup with you've got your cool tech, and you wanna find someone to sell it to. That's not what we do at Chai. At Chai, we start with the consumer.

What does the consumer want? What is their problem, and how do we solve it? So right now, the number one problems for the users, it's not the audio. That's not the number one problem. It's not the image generation either.

That's not their problem either. The number one problem for users in AI is this: All the AI is being generated by middle-aged men in Silicon Valley, right? That's all the content. You're interacting with this AI, you're speaking to it for ninety minutes on average.

It's f- been trained by middle-aged men. The guys out there, they're out there saying like, "Oh, what should the AI say in this situation?" Right? "What's funny?" Right? "What's cool? What's boring? What's, what's entertaining?" That's not the way it should be.

The way it should be is that the users should be creating the AI, right? And so the way I speak about it is this: Chai, we have this AI engine in which sits atop a thin layer of UGC.

So the thin layer of UGC is absolutely essential, right?

**Alessio** [44:00]
It's just prompts.

**William Beauchamp** [44:01]
But it's, but it's just prompts. It's just prompts. It's just an image. It's just a name.

**Alessio** [44:05]
Yeah.

**William Beauchamp** [44:05]
It's like we've done one percent of what we could do. So we need to keep thickening up that layer of UGC. It must be the case that the users can train the AI, and if reinforcement learning's powerful and important, they have to be able to do that.

And so it's gotta be the case that there exists... You know, I say to the team, just as MrBeast is able to spend a hundred million a year or whatever it is on his production company, and he's got a team building the content, the MrBeast content, which then he shares on the YouTube platform.

Until there's a team that's earning a hundred million a year or spending a hundred million on the content that they're producing for the Chai platform, we're not finished, right? So that's the problem. That's what we're excited to build.

And getting too caught up in the tech, I think, is a fool's errand. It does not work.

**Alessio** [44:53]
As an aside, I saw the Beast Games thing on Amazon Prime.

**William Beauchamp** [44:56]
Yeah.

**Alessio** [44:56]
It's not doing well, and I'm curious. It's kinda like, even the-

**William Beauchamp** [44:59]
I know the audience rating is high. The run to middle sucks, but audience rating is high.

**Alessio** [45:03]
But it's not, like, in the top ten. I saw it dropped off-

**William Beauchamp** [45:05]
Oh, okay

**Alessio** [45:05]
... of, like, the-

**William Beauchamp** [45:06]
Yeah, that one I don't know.

**Alessio** [45:06]
I- I'm curious, like, you know, it's kinda, like, similar content, but different platform. And then going back to, like-

**William Beauchamp** [45:11]
Yep

**Alessio** [45:11]
... some of what you were saying is, like, you know, people come to Chai expecting-

**William Beauchamp** [45:14]
Yep

**Alessio** [45:15]
... some type of content.

**William Beauchamp** [45:15]
Yeah, I th- I think it's, um, something that's interesting to discuss is, like, is moats and what, what, what is the moat? And so, you know, if you look at a, a com- a platform like a YouTube, the moat, I think, is in-- first is, is really is in the ecosystem, and the ecosystem c- is comprised of, you have the content creators, you have the users, the consumers, and then you have the algorithms.

**Alessio** [45:40]
Mm-hmm.

**William Beauchamp** [45:40]
And so this, this creates a sort of a flywheel where the al- the algorithms are able to be trained on the users and the users' data. Um, the, the recommended systems can then feed information to the content creators.

So MrBeast, he knows which thumbnail does the best. He knows the first ten seconds of the video has to be this particular way, and so his content is super optimized for the YouTube platform, so that's why it doesn't do well on Amazon.

If he wants to do well on Amazon, how many videos has he created on the YouTube platform, right? Thousands, tens of thousands, I would guess.

**Alessio** [46:12]
Yeah.

**William Beauchamp** [46:12]
He needs to get those iterations in on the Amazon. So at Chai, I think it's all about how can we get the most compelling, rich, user-generated content, stick that on top of the AI engine, the recommender systems, in such that we get this beautiful data flywheel, more users-

**Alessio** [46:30]
Mm-hmm

**William Beauchamp** [46:30]
... better recommendations, more creators, more content, more users.

**Alessio** [46:35]
You mentioned the algorithm. You have this idea of the Chaiverse-

**William Beauchamp** [46:38]
Yep

**Alessio** [46:38]
... on Chai.

**William Beauchamp** [46:39]
Yep.

**Alessio** [46:39]
And you have your own kinda like-

**William Beauchamp** [46:40]
Yes

**Alessio** [46:40]
... LMSys-like ELO-

**William Beauchamp** [46:41]
Yeah

**Alessio** [46:41]
... system.

**William Beauchamp** [46:42]
Yeah.

**Alessio** [46:42]
Um, yeah. What are things that your model's optimized for, like, your user is optimized for? And maybe talk about how you build it, how people submit models.

### Chaiverse ELO

**William Beauchamp** [46:50]
So Chaiverse is what I would describe as a developer platform. More often when, when we're speaking about Chai, we're thinking about the Chai app, and the Chai app is, is really this product for consumers. And so consumers can come on the Chai app, they can interact with our AI, and they can interact with other UGC, and it's really just these, these kind of bots, um, and it's a thin layer of UGC.

Okay. Our mission is not to just have a very thin layer of UGC. Our mission is to have as much UGC as possible, so we must have-- I don't want people at Chai training the AI. I want people, not middle-aged men, building AI.

I want everyone building the AI. As many people building the AI as possible. Okay, so what we built was we built Chaiverse, and Chaiverse is kind of... It's kinda like a prototype, is the way to think about it.

And it started with this, this observation that, well, how many models get submitted into Hugging Face a day? It's hundreds. It is hundreds, right? So there's hundreds of LLMs submitted each day. Now, consider that what does it take to build an LLM?

It take, it takes a lot of work, actually. It's like someone devoted several hours of compute, several hours of their time, prepared a dataset, launched it, ran it, evaluated it, submitted it, right? So there's a lot of work that's going into that.

So what we did was we said, "Well, why can't we host their models for them and serve them to users?" And then what would that look like? The first issue is, well, how do you know if a model's good or not?

Right, like, we don't wanna serve users the crappy models, right? So what we would do is we would-- I love the, the, the LMSys style. I think it's really cool. It's really simple. It's a very intuitive thing, which is you simply present the users with two completions.

You can say, "Look, this is from model A, this is from model B, which is better?" And so if someone submits a model to Chaiverse, what we do is we, um, we spin up a GPU, we download the model- We're gonna now host that model on this GPU, and we're gonna, um, we're gonna start routing traffic to it, and we're gonna send-- We think it takes about 5,000 completions to get an accurate signal.

**Swyx** [48:56]
That's roughly what LLMs says.

**William Beauchamp** [48:57]
Right. And from that, we're able to get an accurate ranking of which models are people finding entertaining and which models are not entertaining. If you look, the bottom 80% are kind of-- they all suck. You can just disregard them.

They totally suck. Yeah. Then when you get the top 20%, you know you've got a decent model, but you can break it down into more nuance. Like, there might be one that's really descriptive, right? There might be one that's got a lot of personality to it.

There might be one that's really lo- uh, logical. Then the question is, well, what do you do with these top models? From that, you can do more sophisticated things. You can try and do, like, a routing thing where you say, for a user-- given user request, we're gonna try and predict which of these end models that users enjoy the most.

That turns out to be pretty expensive and not a huge source of, of, like, edge or improvement. Something that we love to do at Chai is blending, which is... You know, it's-- The, the simplest way, way to think about it is you're gonna end up and you're gonna pretty, pretty quickly see you've got one model that's really smart, one mo- one model that's really funny.

How do you get the user an experience that is both smart and funny? Well, just 50% of the requests, you can serve them the smart model. 50% of the requests, you serve them the funny model.

**Swyx** [50:15]
Just a random 50/50?

**William Beauchamp** [50:16]
Just a random, yeah. And then-

**Swyx** [50:18]
That's blending? Okay.

**William Beauchamp** [50:19]
That's blending. You can do more sophisticated things on top of that-

**Swyx** [50:22]
Uh-huh

**William Beauchamp** [50:22]
... as in all things in life. But the 80/20 solution, if you just do that- ... you get, you get a pretty powerful effect out of the gate.

**Swyx** [50:28]
Random number generator.

**William Beauchamp** [50:30]
I think it's like the robustness of randomness. Random is a very powerful optimization technique-

**Swyx** [50:36]
Yeah

**William Beauchamp** [50:36]
... and it's a very robust thing. So you can explore a lot of the space very efficiently. There's one thing that's really, really important to, to share, and this is the most exciting thing for me, is after you do the, the ranking, you get an ELO score, and you can track a user's first join date.

The first date they submit a model to Chaiverse, they almost always get a terrible ELO, right? So let's say the first submission, they get an ELO of 1,100 or t- or 1,000 or something, and you can see that they iterate, and they iterate, and they iterate, and it, it'll be like no improvement, no improvement, no improvement, and then boom.

**Swyx** [51:11]
Do you give them any data, or do you have to come up with this themselves?

**William Beauchamp** [51:13]
We do. We do. We do. We do.

**Swyx** [51:15]
Okay. Yeah, yeah.

**William Beauchamp** [51:15]
We try and strike a balance between giving them, um, data that's very useful. You've got to be compliant with GDPR-

**Swyx** [51:21]
Yeah

**William Beauchamp** [51:21]
... which is like you, you have to work very hard to preserve the privacy of, um, users of your app.

**Swyx** [51:27]
Okay.

**William Beauchamp** [51:27]
So we try to give them as much signal as possible to be helpful. The minimum is we j- we're just gonna give you a score, right? That's the minimum. But that, that alone is-- people can optimize a score pretty well because they're able to come up with theories, submit it.

Does it work? No. A new theory. Does it work? No. And then boom, as soon as they figure something out, they keep it, and then they iterate, and then boom, they figure something out, and they keep it.

**Swyx** [51:49]
Last year, you had this post on your blog, crowdsourcing the lead to the-

**William Beauchamp** [51:53]
Yes

**Swyx** [51:53]
... 10 trillion parameter AGI.

**William Beauchamp** [51:54]
Yeah. Yeah.

**Swyx** [51:54]
And you call mixture of experts recommenders.

**William Beauchamp** [51:57]
Yeah.

**Swyx** [51:58]
So since-

### AGI Debate

**William Beauchamp** [51:58]
Yeah

**Swyx** [51:58]
... any updated thoughts 12 months later?

**William Beauchamp** [52:02]
Yeah. I think the odds, the timeline for AGI has certainly been pushed out. Right now, this is in-- I'm a controversial person. I don't know. Like-

**Swyx** [52:08]
Let's do it

**William Beauchamp** [52:09]
... I just think-

**Swyx** [52:10]
You don't believe in scaling laws. You, you, you think AGI is further away.

**William Beauchamp** [52:13]
I think it's an S-curve. I think everything's an S-curve. And I think that the models have proven to just be far worse at reasoning than people sort of thought. And I think whenever I hear people talk about LLMs as, as reasoning engines, I sort of cringe a bit.

I don't think that's what they are. I think of them more as like a simulator. I think of them as like a s- like a s- right? So they get trained to predict the next most likely token. It's like a physics simulation engine, so you get these, like, games where you can, like, construct a bridge, and you drop a car down, and then it predicts what, what, what should happen, and that's really what LLMs are doing.

Um, it's not so much that they're reasoning, it's more that they're just doing the most likely thing. So fundamentally, the ability for people to add in intelligence, I think is very limited. What most people would consider in- intelligence, I think the AI is, is not a crowdsourcing problem, right?

Now, with Wikipedia, Wikipedia crowdsources knowledge.

**Swyx** [53:14]
Mm.

**William Beauchamp** [53:15]
It doesn't crowdsource intelligence.

**Swyx** [53:16]
Mm.

**William Beauchamp** [53:17]
So it's a subtle distinction. AI is fantastic at knowledge. I think it's weak at intelligence. And a lot-- It's easy to conflate the two, 'cause if you ask it a question and it gives you... You know, if you said, "Who was the, the seventh president of the United States?"

And it gives you the correct answer, I'd say, "Well, I don't know the answer to that." And you can conflate that with intelligence, but really that's a question of knowledge. And knowledge is really this thing about saying, "How can I store all of this information, and then how can I retrieve something that's relevant?"

Okay. They're fantastic at that. They're fantastic at storing knowledge and retrieving the relevant knowledge. They're superior to humans in that regard. And so I think we need to come up for a new word for, like, how does one describe...

AI should contain more knowledge than any individual human. It should be more accessible than any individual human. That's a very powerful thing. That's super powerful. But what words do we use to describe that?

**Swyx** [54:11]
Uh, we had a previous guest on Hexa AI that does-

**William Beauchamp** [54:15]
Yeah

**Swyx** [54:15]
... uh, search, uh, and they-- he tried to coin super knowledge as-

**William Beauchamp** [54:18]
Super knowledge, yeah

**Swyx** [54:19]
... as, as the opposite-

**William Beauchamp** [54:20]
I think that's a better, yeah

**Swyx** [54:20]
... of super intelligence.

**William Beauchamp** [54:20]
Exactly. I think, I think super knowledge is a, is a more accurate word for it.

**Swyx** [54:25]
Yeah, you can store more things than any human can.

**William Beauchamp** [54:27]
Exactly.

**Swyx** [54:27]
But you may not be more intelligent.

**William Beauchamp** [54:28]
And you can, and you can retrieve it better than any human can as well.

**Swyx** [54:31]
Yeah.

**William Beauchamp** [54:31]
And I think it's those two things combined that's special.

**Swyx** [54:34]
Yeah.

**William Beauchamp** [54:34]
I think that thing will exist. That thing can be built, and I think you can start with something that's entertaining and fun, and I think, um- I often think it's like, look, it's gonna be a 20-year journey, and we're in, like, year four.

Or it's like the web, and this is like 1998 or something, you know? You've got a long, long way to go before the Amazon.coms are like these huge multi-trillion dollar businesses that every single person uses every day. And so AI today is very simplistic, and it's fundamentally the way we're using it, the flywheels and this, this, this ability for how can everyone contribute to it to really magnify the value that it brings.

Right now, like, I think it's a bit sad almost, like right now you have big labs, I'm gonna pick on OpenAI, and they kind of go to like these human labelers and they say, "We're gonna pay you to just label this, like, subset of questions that we want to get a really high-quality dataset.

Then we're gonna get, like, our own computers that are really powerful," and that's kinda like the thing. For me, it's so much like Encyclopedia Britannica. It's, like, insane. All the people that were interested in blockchain, it's like, well, this is, this is what needs to be decentralized.

You need to decentralize that thing. Yeah. Because if you distribute it, people can generate way more data in a distributed fashion. Way more. Right? You need the incentives. Yeah, of course, yeah. But I mean, the c- the ...

That's kind of the, the exciting thing about Wikipedia, was it's this understanding, like the incentives, you don't need money to incentivize people- You don't need dog coins? No. No, no. Sometimes, sometimes people get the satisfaction from just seeing the correct thing.

Number go up. Yeah, yeah. So- But, I mean, you do pay money for Chaiverse. We've, we've paid out over $100,000 to model creators, but do you know what we saw? It's not motivating. We saw that it didn't really make a difference.

Like, if they were submitting models at a certain rate, if you paid them a bunch of money, they didn't change the rate. What the money let them do was if they wanted to fine-tune a Llama 70B on eight H100s overnight, if you give them money, then they can do it.

Yeah, or you give them compute. Yeah. So, so I think the most exciting person we ever saw from interacting with Chai, Chaiverse was we gave some kid who was, like, um, like 17 years old, I think we gave him $1,000, and he spent all the money on buying a physical GPU, and he took a picture of it and said, "This is what I bought, and I'm gonna be training more models with it."

So that's why I, that's why I love platforms. Should you hire him or That's the temptation. Yeah. That's the temptation. But you wanna keep the team small. No, no. As a platform- Yeah ... we can't just hire every good content creator.

We've got to build the systems. And the best content creator today isn't gonna be the best content creator next year. What about evals? So you've talked about reasoning and knowledge. Most of the benchmarks that people use- Yep ...

### Evals

**William Beauchamp** [57:20]
wanna mimic reasoning. I wanna register a disagree on the reasoning, but we have to- Yeah, yeah, yeah ... we have to keep going. No, no, but yeah, I'm curious, like how, how do you think about the evals that matter to you?

So yeah, like is ... ELO cannot be your only eval. You, you must have internal evals. Um- You mentioned evals ... I think ELO is a fantastic north star- Yeah ... and the reason for ... Or like it's the main one we wanna see go up because it's this human feedback.

The humans know what they want. It's beautiful because when you come up with an eval, you're further removing yourself away from the true problem, right? So whatever it is you're trying to optimize or, or figure out, you kinda have to, have to slice it, and then you've got this ...

It's like a snapshot. Like, as soon as you saturate one eval, you need to figure out a new eval. But with, by saying to humans just which is better, A or B, it's super robust, it's super generalizable, it just keeps, keeps scaling.

So we've, in the past, used evals to get through a, to get through a blocker. I mean, a great example is, you know, is like having like a safety filter or something where you wanna make sure your models ...

Because listen, users find ... You'll be shocked at the correlation between not family-friendly content, whether that's just like swearing, like y- people find it funny when the AI swears. Mm-hmm. So if you have two completions, A or B, like if you give me any LLM, I can make it 20% funnier just by training it to throw in swear words.

So the issue with that is it's like how ... Are we measuring like quality improvements? Are we measuring superficial improvements, right? And this actually links back to the LMSys. They did a- Style control ... style control. Exactly. We actually had them on the podcast.

Yeah, yeah. And so that's the way ... I, I would rather just lean on human feedback and just continue to make that more and more robust and more and more useful. And, you know, you can say some people are like GPU poor and GPU rich.

We're like, we're feedback rich. Like, when you've got one and a half million people a day, we get as much feedback from humans as we want. So we're not in a position where we needed to have the evals very much.

And when we do, we saturate them pretty quick. So a safety one, you know, within a month we don't need to use it anymore 'cause it's sort of- Yeah ... it's, you know, the issue has been addressed. I, I think one problem I have, um, and this is a broader products question maybe, is- Yep ...

the, the ELOs apply to the whole user population. That's right. Clearly, the user behavior, there's segments- Yep ... that have, like, I'm, I'm a role play person- Yep ... I'm a thera- therapy person, I'm a not safe for work person.

Exactly. You don't split them? This is why I say, like, I think we're in year four of like a 20-year thing, where it's like at the end of the day, if we all go on like Spotify or like imagine if Spotify only had the top five musicians, I think it would retain over 85% of its existing users.

Yeah. Right? And I think if YouTube, if YouTube only kept the top five content creators, it would be enough for the vast majority of people. The thing I'm just trying to share here is there's ... One surprising thing about humans is their preferences are pretty correlated.

What you find funny and entertaining, I find funny and entertaining, and he finds funny and entertaining. There might be degrees of variation in it. I might find it super funny, you might find it only slightly funny, but optimizing to a global works very, very well.

And for segmentation to be really powerful, segmentation will work amazing. If you found a comment super boring and I found it super fun, if we could segment that, then that would unlock really powerful stuff. But unfortunately, that's not the shape of human behavior Right?

It's like I might rank it 10 out of 10 funny, you might rank it seven out of ten funny. And it's like it doesn't give you as much space to play as you would hope. It's an element of the diversity of content that AI can produce right now, which is it's the-- it's not as diverse as if you consider a platform like YouTube, you can watch a MrBeast video that's totally different to a makeup tutorial.

So there's enough diversity there where if you go on my YouTube feed, it is totally different to my sister's one. My sister's one, it's all like women and, you know, mine-- if you go on mine, it's all like bald, middle-aged men either talking about MMA or...

Right? But I think with, with AI, it's still, it's still a bit too early for that degree of segmentation. So I think it all comes, the recommended systems, the personalization. But this is why I like the don't start with the technology, start with the problem.

The problem is UGC. We must give users the tools to build more variety and more engaging content.

**Swyx** [1:01:43]
Yeah. I feel like there's-- I was surprised at how thin it was when I tried out Chai.

**William Beauchamp** [1:01:47]
Yeah. Yeah.

**Swyx** [1:01:48]
Um, it's very thin.

**William Beauchamp** [1:01:49]
Agree.

**Swyx** [1:01:49]
Haven't you been tempted, like there's, there's this ecosystem of Cobalt, Silly Tavern-

**William Beauchamp** [1:01:53]
Yep

**Swyx** [1:01:53]
... those guys.

**William Beauchamp** [1:01:54]
Yep.

**Swyx** [1:01:54]
They have model cards.

**William Beauchamp** [1:01:55]
Yep.

**Swyx** [1:01:55]
It, it seems like an industry standard almost.

**William Beauchamp** [1:01:57]
Yep. Yep. Agreed.

**Swyx** [1:01:58]
You can-- can I just import those?

**William Beauchamp** [1:02:00]
Um, let me think what to say.

### Content & Culture

**Swyx** [1:02:02]
Oh, you're already working on it.

**William Beauchamp** [1:02:03]
No, it's like they-- I remember when Chai man-- Chai, Silly Tavern and like Co- uh, Cobalt, Cobalt AI is, is basically as old as Chai. So when Chai was-- when we just existed, they just existed and both of us were using GPT, uh, J.

Yeah, yeah, yeah. And I remember very early on I was like, "These guys shouldn't even exist because if we build a good enough platform, they should just be posting their content on our platform."

**Swyx** [1:02:29]
Yeah, but they're open source, so you can-

**William Beauchamp** [1:02:30]
No, exactly. That was what I learned. Uh, eventually I learned like their-- what they're excited about is slightly different from a typical-

**Swyx** [1:02:36]
Yeah

**William Beauchamp** [1:02:36]
... um, consumer. My answer is, um, it's kinda like a complex thing where it's, it's really down to the content creator wants-- typically they, they're building it for themself and typically they want to create an experience for themself.

One content creator might have to write a thousand words describing... They'll take a science fiction scenario. They say, "Okay, you're on a spaceship and you're going off into space and your crew, these are your crew members. You've got one that's really friendly, one that's really mean, and you're the new cadet and you wanna rise to the top."

And they can really go into great de- Right? And then you can give that to like a Llama 70B and Llama 70B will do a pretty good job of adhering to the prompt and the user will have a good experience.

Okay. Now on Chai, very few users will ever go to that level of content creation. If instead the user, we can really make our AI understand the user more so that rather than having to use a thousand characters or a thousand tokens to describe the scenario, we can just say, "Look, you're in a spaceship, you've got three crew mate.

It's gonna be dramatic and there should be some fighting." And then the AI gives you an even better experience, then the content creator is happier. And so fundamentally the way I'd kind of think about it is there's the steerability of the AI and so a lot of the work we do at Chai is, is really about saying we want the AI to react to the user and react to the content creator in the way that they most want.

One kind of like analog would be TikTok. I think the thing that TikTok did insanely good was they made it really easy for like anyone-- If you make a video on TikTok, almost anyone can make a kind of fun video really easy.

You just put some music on the top of it, you throw some of the animations on top and it's not hard to have a pretty fun thing. And I think that's much more like the Chai style, where it's like users don't wanna have to work, you know-- If your content is only good if you have like Shakespeare, it's better if, if just anyone at home can make the-

**Swyx** [1:04:35]
Sure

**William Beauchamp** [1:04:35]
... can make the thing. So that's, that's kinda like my answer to the Silly Tavern style and I think the right answer is how do you get the Silly Tavern people fine-tuning models that create a really special effect?

**Alessio** [1:04:47]
As we wrap, this is kinda like the call for action.

**William Beauchamp** [1:04:50]
Mm-hmm.

**Alessio** [1:04:50]
Uh, part one, you have Chai Grant-

**William Beauchamp** [1:04:52]
Yep

**Alessio** [1:04:53]
... which I think a lot of people don't know about, which is grants for open source projects.

**William Beauchamp** [1:04:56]
Absolutely.

**Alessio** [1:04:56]
Any ideas, any projects that you want to see people work on that should apply or-

**William Beauchamp** [1:05:02]
Yeah. Let me think. I think, um, so we do Chai, Chai Grant and fundamentally, you know, we give cash, no strings attached. It's kind of our way of, of doing two things. One, giving back and supporting the community.

We've benefited from a lot of open source packages. A lot of our developers and engineers are like really pro open source. And then also it's a great way to just meet talented people and, and like expand connections. So with respect to Chai Grant, if anyone's got any sort of, um, GitHub project, any sort of thing they've built that they're proud of, just apply.

Just apply. It's like no strings attached cash and people have a pretty high success rate. So that's the first thing. Other call to actions would be, I think Chai is this, you know, it's a startup. We're a small team.

It's like 15 people. We work very intense. It's a very hardcore sort of environment which we found that a lot of people don't like. They don't like the, you know, they'll ask us this concept of work-life balance. One time a person said-- They said something like, "I can't get this done because I'm taking PTO on Friday."

And I said, "What is PTO?"

**Swyx** [1:06:05]
Okay.

**William Beauchamp** [1:06:06]
Um, it stands for paid time off. And this-

**Swyx** [1:06:09]
I know what it is.

**William Beauchamp** [1:06:10]
And this person was gone. They didn't-- Like, they were no longer in the company four weeks on.

**Alessio** [1:06:15]
I mean, legally I think you have to.

**William Beauchamp** [1:06:18]
Well, it's true. There's no problem. Look, if you've gotta take a day off, right? We all have personal lives, right? But it's about this idea of responsibility.

**Swyx** [1:06:26]
Yeah.

**William Beauchamp** [1:06:26]
If you're not in the office on Friday, you still have your responsibilities. So I don't care if you work hard Thursday to get it wrapped up. I don't care if you're working hard Saturday to get it wrapped up.

**Swyx** [1:06:34]
It's not an excuse to-

**William Beauchamp** [1:06:35]
It's not an excuse. The way this individual spoke about it, it was like an excuse. I think it's an environment, very talented engineers working very hard in an intense space. It's the thing that gets me excited. It's, it's why I think, you know, I really love working at Chai, is because it's a place of talented people working super hard.

So yeah, I think people who have got-- who've worked at startups and they, they lo- that's what they, they want the taste of, I think they should reach out, they should apply and I think 90% of people are gonna say that sounds terrible Don't apply

**Alessio** [1:07:03]
It's not for them.

**William Beauchamp** [1:07:04]
Yeah. It's exactly, exactly.

**Alessio** [1:07:06]
Yeah. I just realized we skipped one important part. So you spent $10 million on compute-

**William Beauchamp** [1:07:11]
Yep

**Alessio** [1:07:11]
... last year.

**William Beauchamp** [1:07:11]
Yep.

**Alessio** [1:07:12]
You say you're gonna probably triple that.

**William Beauchamp** [1:07:13]
Yep.

**Alessio** [1:07:14]
I'm sure you're doing a lot of work on custom kernels, kinda like inference optimization.

**William Beauchamp** [1:07:18]
Yes. Yep.

**Alessio** [1:07:19]
Any cool stuff that you wanna share there?

### Inference Tricks

**William Beauchamp** [1:07:21]
Yeah, lots of cool stuff. So really quickly, I think inference is very, very important. It's super important. It's massively underlooked, and we can look at all the different foundation models and the techniques, the differences in the foundation models on how well they perform from a cost perspective with inference.

Mixture of experts, for example, tend to do really, really good from, like, a cost perspective. We've worked with a very talented team called MK1, and-

**Alessio** [1:07:51]
Yeah, we-- So I saw, I saw them in the Chaiverse logs. What are they?

**William Beauchamp** [1:07:54]
We were using-- We were running vLLM for a while, and vLLM's really fantastic, absolutely amazing the work that they've done and achieved. And at some point, I got introduced to, the founder's name is Paul Merolla, and he was a co-founder at Neuralink, really, really expert in, like, hardware.

He kind of explained to me, he was like, "Look, if you know hardware really well, you can write the CUDA kernels really well." He said, "You should check out our inference engine," and they kind of blew vLLM out the water when we evaluated it.

Much, much, much faster. And I think the special thing that he was able to do with us is we love rejection sampling, so we do much more rejection sampling than maybe typical. And, you know, generate it so we, we never, ever, ever just generate a single completion, right?

This is why we don't do streaming. A lot of people-- Like, ChatGPT use- used to do a lot of streaming. Like, the-

**Alessio** [1:08:47]
Yeah

**William Beauchamp** [1:08:47]
... completion would come out one thing at a time.

**Alessio** [1:08:49]
I did notice that in your UX.

**William Beauchamp** [1:08:50]
Yeah.

**Alessio** [1:08:50]
Normally chat you have to stream.

**William Beauchamp** [1:08:51]
Exactly, but Chai has never done streaming because if you stream, you're unable to do re- rejection sampling.

**Alessio** [1:08:57]
Mm-hmm.

**William Beauchamp** [1:08:57]
The benefit of that is you can serve a larger model. The reason why you can serve a larger model is because let's say instead of generating a completion in four seconds, because the user gets the first token faster, you can generate in 10 seconds.

Well, if you've got 10 seconds to generate a completion, you can serve a much larger model. So typically, the people that are streaming, the benefit that they're getting is they're serving a larger model. With Chai, we give you, uh, w- you know, the second the answer comes, boom, you get the full completion, and the reason for that is because we wanna generate 16 completions, see the entire response, and then we wanna evaluate which one we think is the best.

Um-

**Alessio** [1:09:35]
Do you have a separate LLM evaluator?

**William Beauchamp** [1:09:37]
Of-- Yes, we do. Yeah. So, um, typically they're referred to as a reward model, and that's a, you know, that's like a term from reinforcement learning.

**Alessio** [1:09:44]
Reinforcement learning.

**William Beauchamp** [1:09:44]
And for that you can start off with something very simple, which is, do you think the user is going to respond to it? That's a simple one. So you can, you can train, you can take 50 million messages and, and look at all the sorts of messages users reply to, which ones they don't, and then you can train this, this reward model to evaluate completions.

**Alessio** [1:10:03]
Mm.

**William Beauchamp** [1:10:03]
And so it knows like, okay, if you say this, the user's not gonna respond. So don't bother sending it to the user. If you say this, the user's definitely gonna engage with this, so send them-

**Alessio** [1:10:11]
Mm.

**William Beauchamp** [1:10:11]
Send them that.

**Alessio** [1:10:12]
There's an interesting parallel between MoEs at the top-

**William Beauchamp** [1:10:14]
Yep

**Alessio** [1:10:15]
... spreading out to different experts-

**William Beauchamp** [1:10:16]
Yep

**Alessio** [1:10:16]
... and then at the bottom with rejection sampling-

**William Beauchamp** [1:10:18]
Yep

**Alessio** [1:10:19]
... con- con, uh, choosing from different-

**William Beauchamp** [1:10:21]
Yep

**Alessio** [1:10:21]
... paths.

**William Beauchamp** [1:10:22]
Yep.

**Alessio** [1:10:22]
I don't know, just-

**William Beauchamp** [1:10:22]
I totally agree. That's the stuff that is the future of AI. I think that's the exciting stuff. And there's a parallel between that. Why was AlphaGo able to be superhuman?

**Alessio** [1:10:33]
Okay.

**William Beauchamp** [1:10:33]
Right? It's this ability to generate many different paths-

**Alessio** [1:10:37]
Tree search. Yeah

**William Beauchamp** [1:10:38]
... and tree search. Exactly. So I think if you wanna talk about what would intelligence look like, it looks much more like tree search, combining the generative nature of these LLMs with a really good tree search, and that's what openly I've done with o1 and o3.

**Alessio** [1:10:52]
I don't know that they do tree search. They never said.

**William Beauchamp** [1:10:55]
They do.

**Alessio** [1:10:55]
It's implied.

**William Beauchamp** [1:10:56]
Yes.

**Alessio** [1:10:56]
Okay.

**William Beauchamp** [1:10:57]
Yes. Yes.

**Alessio** [1:10:58]
You-- Are you comfortable with o1 being a reasoning engine?

**William Beauchamp** [1:11:01]
No. No, no, no. I'm, I'm kind of saying it's better at reasoning-

**Alessio** [1:11:04]
Okay

**William Beauchamp** [1:11:04]
... because they leverage the tree search well. And the, the issue of the reasoning is they're saying, is this-- Like, they train, um, they have the models to say, "Is this logically correct, and what's the likelihood of it being logically correct?"

So you can build up the sophisticated mechanisms to get it less bad at reasoning, but you'll see, like, eventually what, what AI is really, really good at, people won't say it's r- It's always gonna be better at retrieving.

It's always gonna be better at storing knowledge, which is so highly correlated with intelligence that we often assume it's the same. What, what AI is truly special at and gets consumers really excited is it's generative. It can just make stuff.

We've never had a technology before that can just make stuff.

**Alessio** [1:11:45]
Simulate, yeah.

**William Beauchamp** [1:11:46]
Yeah. So that's the special, that's the exciting thing.

### Closing

**Alessio** [1:11:48]
Awesome. Well, any parting, parting thoughts?

**William Beauchamp** [1:11:52]
No, it's b- it's been a pleasure. I guess the only thing I'd add is, like, our office is in Palo Alto, so, um-

**Alessio** [1:11:57]
Yeah

**William Beauchamp** [1:11:58]
... you know, people with startup experience looking to join a fast-growing, high-impact startup-

**Alessio** [1:12:03]
Yeah

**William Beauchamp** [1:12:04]
... in Palo Alto

**Alessio** [1:12:04]
... uh, we'll find your culture deck-

**William Beauchamp** [1:12:05]
Yeah

**Alessio** [1:12:05]
... which is, which is good.

**William Beauchamp** [1:12:06]
Oh, great. Fantastic.

**Alessio** [1:12:06]
And then also-

**William Beauchamp** [1:12:07]
So what-

**Alessio** [1:12:07]
Your-- Yeah, 100K

**William Beauchamp** [1:12:08]
... yeah, what's the story where if you made 100K trading, we'll fast-track your application? Like-

**Alessio** [1:12:12]
You mean, I kinda qualify. Yeah.

**William Beauchamp** [1:12:16]
We just looked at the team-

**Alessio** [1:12:17]
Mm-hmm

**William Beauchamp** [1:12:18]
... and it, it got to the point where almost every single person on the team you could point to, and they had done something special before joining the team. Like, they, they had strong markers of, like, there was something special about them.

**Alessio** [1:12:28]
Mm-hmm.

**William Beauchamp** [1:12:28]
That's not to say it's, like, like an exclusive thing. You have to-

**Alessio** [1:12:31]
Yeah, yeah

**William Beauchamp** [1:12:31]
... have achieved something special. But it's just, uh, we got this one engineer, and she, she started going to college, she went to CMU when she was, like, 15 years old or something, and it's like, that's a bit special.

There's another engineer, he created a Git repo, and I think it got, like, 1,500 stars, and it was, like, a repo for, like, there were some drivers that he wrote. It was, like, a super low, low-level thing. I was like, "That's a bit special."

We had this other guy, he joined the team, and he'd, he had made 100K buying and selling sneakers, right?

**Alessio** [1:13:01]
Trading.

**William Beauchamp** [1:13:02]
Yeah.

**Alessio** [1:13:02]
Okay.

**William Beauchamp** [1:13:03]
So, so it's like, it's just this thing, like, like, if you've been to Harvard, cool. That's great. It shows that you're really smart, and you work really hard. Cool. That's good. But if you've actually built something and, and, and done something that's a bit more tangible, that gets us even more excited.

**Alessio** [1:13:16]
Awesome.

**William Beauchamp** [1:13:17]
Yeah.

**Alessio** [1:13:17]
Well, thanks for having us-

**William Beauchamp** [1:13:18]
Yeah, pleasure

**Alessio** [1:13:18]
... at Chai HQ.

**William Beauchamp** [1:13:19]
Yeah. Thanks, guys.

---

This library is powered by PodHood (https://podhood.com), the podcast website platform.
