# A Brief History of the Open Source AI Hacker - with Ben Firshman of Replicate

Latent Space · 2024-02-28

<https://addtry.com/16d27d42-1b09-40bc-87f8-f8f7cf028c13>

Ben Firshman, CEO of Replicate, explains how the inference platform grew from a research reproducibility tool into a 2M-user API business by embracing the generative image community and treating open-source AI as a hacker-friendly ecosystem. They accidentally discovered their API when a user reverse-engineered their web form, leading to their first $1k/month customer. Cog, their container standard for ML models, was born from lessons at Docker and the need to make models tinkerable. Ben argues that fine-tuning's low cost makes open-source models sustainable, and that AI engineers (orders of magnitude more than ML engineers) just need to start playing with models. He also discusses GPU scarcity, preferring sustainable pricing over price wars, and reveals that demand is not outpacing supply thanks to aggregating demand.

## Questions this episode answers

### How did Replicate accidentally discover their business model and get their first paying customer?

Ben Firshman explains that early Replicate had no public API; users ran models through a web form. One user reverse-engineered the internal JavaScript API to generate images in bulk. Instead of blocking them, the team documented the API, reached out, and offered it for about $1,000 a month. The user, who was creating NFT art, agreed, becoming their first paying customer and sparking the API business.

[31:57](https://addtry.com/16d27d42-1b09-40bc-87f8-f8f7cf028c13?t=1917000)

### What failures did the Replicate founders experience before arriving at the idea that worked?

After creating ArxivVanity to convert PDFs, Ben and Andreas entered YC with a CI-like benchmarking platform for ML researchers. They couldn’t find a business model, built COVID tools that went nowhere, and then made an open-source experiment tracker that researchers didn’t adopt. Running out of money, they returned to their original idea: packaging ML models in containers, which became Cog and Replicate.

[17:19](https://addtry.com/16d27d42-1b09-40bc-87f8-f8f7cf028c13?t=1039000)

### What design principles from Docker did Replicate apply when creating Cog, and why choose a higher-level abstraction?

Ben, who worked on Docker Compose, says Cog treats a model as a Docker container with an OpenAPI spec defining inputs and outputs. They avoided low-level Dockerfiles, using a YAML config for Python dependencies instead, because ML researchers are not Linux experts. Cog automatically handles CUDA versions, inference servers, and batching, serving as an open standard for packaging and sharing models.

[44:57](https://addtry.com/16d27d42-1b09-40bc-87f8-f8f7cf028c13?t=2697000)

## Key moments

- **[0:00] Intro**
  - [0:36] Ben Firshman says building physical things like fixing cars transfers skills to building software tools.
  - [1:22] Ben Firshman's CLI design principle: 'Your command line interface should feel like a big iron machine, where you pull a lever and it goes clunk.'
  - [2:49] Q: How did Ben operationalize low latency in Fig/Docker? A: He switched from Python to Go because Python's runtime startup was too slow.
  - [4:58] Ben Firshman: CLIs have evolved from machine interfaces to human interfaces, and can now be conversational like a language model.
- **[6:47] Arxiv Vanity**
  - [6:47] Ben Firshman's Archive Vanity project was born from frustration with reading ML papers as PDFs on his phone during a vacation in Greece.
  - [10:33] Ben's friend Andreas reading an ML paper on his phone: 'This is fucking stupid.' This sparked the idea to convert arXiv papers to HTML.
  - [11:49] Ben Firshman's Archive Vanity caught arXiv's eye, but he got distracted by a new idea that became Replicate.
  - [13:13] Andreas, a Spotify ML engineer, found it hard to replicate research because papers omitted details and code often didn't run.
  - [15:54] The initial Replicate concept was a CI system for ML to generate reproducible baselines, solving the problem of unreliable paper benchmarks.
  - [17:19] During YC, Ben and his co-founder realized they had no business model for their ML tool and failed to sell a non-existent product.
- **[19:47] YC & Pivot**
- **[23:11] Community Launch**
  - [23:11] After a failed COVID tools pivot, the team built Cog and Replicate as an open-source experiment tracking tool, later pivoting to model sharing.
  - [26:06] Ben raised a seed round pre-launch and pre-team to build Cog and Replicate, aiming to make ML research accessible via containers and web demos.
  - [28:39] Ben got captivated by the early generative image community on Discord, like the VQGAN-CLIP bot, which felt like the start of something big.
  - [29:26] Ben on the generative image community: 'This feels like open source software,' with people tinkering and forking models.
- **[31:58] API & Growth**
  - [31:58] Replicate's API business started when a user reverse-engineered their internal API to generate NFT art, leading to the first customer paying $1,000/month.
  - [33:09] Ben on the first customer: 'Here's the documentation. That'll be like 1,000 bucks a month, please.' And they said, 'Sure, that sounds great.'
  - [35:54] Stable Diffusion's release in August 2022 was an inflection point: it was open source, tinkerable, and good enough to attract a flood of users to Replicate.
  - [39:37] Indie hackers and enterprise customers have surprisingly similar needs: both want to use AI without becoming infrastructure experts, says Ben.
  - [42:58] Q: How do companies use Replicate? A: Unsplash uses BLIP to annotate millions of images; others fine-tune or run custom inference pipelines.
- **[44:30] Cog & Standards**
  - [44:30] Cog was designed to be higher-level than Docker, using a YAML config and OpenAPI spec to abstract away CUDA, inference servers, and batching.
  - [52:24] Replicate's Cog vs Ollama's Model File vs Llama File: Ben says they are complementary, with Cog for servers and Ollama for local.
- **[57:49] Compute & Waste**
  - [57:49] Replicate acts as a 'GPU bank,' aggregating demand to commit long-term with cloud providers, then offering short-term access like a maturity transformation.
  - [1:02:38] Most models on Replicate are unoptimized; only popular ones like Stable Diffusion are optimized, leaving a long tail of GPU waste.
  - [1:08:37] Replicate prices models sustainably, avoiding price wars, because the real value is the ability to customize and fine-tune open models on the platform.
- **[1:10:58] Open Source & Future**
  - [1:10:58] Ben on open source AI licensing: it's okay for model companies to restrict commercial use if the license still allows tinkering and fine-tuning.
  - [1:15:19] Ben predicts that software engineers who ignore AI will be like those who ignored the internet in the 90s: it's a must-learn skill.
  - [1:17:33] Q: Who should apply to Replicate's hacker-in-residence role? A: People who love building and showing others how to use AI, like Charlie Holtz.

## Speakers

- **Alessio** (host)
- **Swyx** (host)
- **Ben Firshman** (guest)

## Topics

Inference, Open Source Tools, Image Generation

## Mentioned

Hugging Face (company), Ollama (company), OpenAI (company), Replicate (company), Y Combinator (company), Arxiv (product), ArxivVanity (product), BLIP (product), Big Sleep (product), CLIP (product), Cog (product), Diffusers (product), Docker (product), Docker Compose (product), Llama (product), Llama File (product), Pixray (product), Stable Diffusion (product), Transformers (product), VQGAN-CLIP (product)

## Transcript

### Intro

**Alessio** [0:00]
Hey everyone, welcome to the Latent Space Podcast. This is Alessio, partner and CTO in residence at Decibel Partners, and I'm joined by my co-host Swyx, founder of Small AI.

**Swyx** [0:09]
Hey, and today we have Ben Firshman in the studio. Welcome, Ben.

**Ben Firshman** [0:13]
Hey, good to be here.

**Swyx** [0:15]
Uh, Ben, you're a co-founder and CEO, uh, CEO of Replicate. Uh, before that you were most notably cr- uh, creator of Fig, or founder of Fig, which got... Which became Docker Compose. Um, you also did a couple other things, um, before that, but, uh, that's what a lot of people know you for.

Um, what should people know about you that, you know, outside of your, your sort of LinkedIn- ... profile?

**Ben Firshman** [0:36]
Yeah, good question. I think I'm a builder and tinkerer, like in a very broad sense, and I love using my hands to make things. So like I work on, you know, things maybe a bit closer to tech, like electronics, but I also like build things out of wood and I like, uh, fix cars, and I fix my bike and build bicycles and all this kind of stuff.

And there's so much I think I've learnt from transferable skills from just like working in the real world to building things, building things, uh, in software. Um, and you know, so much about being a builder both in real life and, and in software that's, uh, that, uh, crosses over.

**Swyx** [1:14]
Is there a real world analogy that you use often when you're thinking about like a code ar- architecture or, or problem?

**Ben Firshman** [1:22]
I like to build software tools as if they were a, um, as if they were something real. I like to imagine, uh... So I wrote this thing called the Command Line Interface Guidelines, which was a bit like sort of-

**Swyx** [1:38]
Yes

**Ben Firshman** [1:39]
... the Mac Human Interface Guidelines, but for command line interfaces. I did it with, um, uh, the guy I created Docker Compose with, um, and, and a few other people. And I think something in there, I think I described that your command line interface should feel like a big iron machine, where you pull a lever and it goes clunk.

And like things should respond within like 50 milliseconds, um, as if it was like a real life thing. Um, and like another analogy here is like in the real life, you know when you press a button on an electronic device and it's like a soft switch, and you press it and nothing happens, and there's no physical feedback-

**Swyx** [2:18]
Mm-hmm

**Ben Firshman** [2:18]
... of anything happening, and then like half a second later something happens? Like that's how a lot of software feels, but instead, like software should feel more like something that's real, where you touch, you pull a physical lever and the physical lever moves, you know?

And I've taken that lesson of kind of human interface, um, to, to software a ton. You know, it's all about kind of low latency, it feeling, things feeling really solid and robust, um, both the command lines and, and uh, and user interfaces as well.

**Swyx** [2:44]
And how did you operationalize that for Fig or Docker?

**Ben Firshman** [2:49]
Uh, a lot of it's just low latency. Actually, we didn't do it very well for, for Fig and Compose- ... in the first place. We used Python, which was a big mistake, where Python's really hard to get booting up fast, because you have to load up the whole Python runtime before it can run anything.

**Swyx** [3:02]
Okay.

**Ben Firshman** [3:03]
Um, Go is much better at this, where like Go just instantly starts, um...

**Swyx** [3:08]
So, so what's the... Like you have to be under 500 milliseconds to start up?

**Ben Firshman** [3:12]
Yeah, effectively.

**Swyx** [3:13]
Yeah. Okay.

**Ben Firshman** [3:13]
I mean, I mean, you know, human, perception of human things being immediate is, you know, something like 100 milliseconds.

**Swyx** [3:18]
Okay.

**Ben Firshman** [3:19]
Uh, so any- anything like that is, is, is yeah, good enough.

**Swyx** [3:24]
Yeah. Um, um, also I should mention since we're talking about your side projects, uh... Well, one thing is, um, I am maybe one of a few fellow people who have actually written something about CLI design principles, uh, because I was, uh, in charge of the Netlify CLI back in the day, um, and, uh, had many thoughts.

Uh, but one of my fun thoughts I'll just share in case you have thoughts is, um, I think CLIs are effectively starting points for scripts that are, that are then run, and the moment one of the script's preconditions are fa- are not fulfilled, typically they end.

So the, the CLI developer will just, will just exit the-

**Ben Firshman** [3:57]
Mm

**Swyx** [3:58]
... the program. Um, and the way that I desi- I, I really wanted to, to create the Netlify dev workflow was for it to be kind of a state machine that would resolve the, itself. Um, if it detected a precondition wasn't fulfilled, it would actually delegate to a subprogram that would then fulfill that pre- precondition, asking for more info or waiting until a condition is fulfilled, then it would go back to the original flow and continue, continue that.

**Ben Firshman** [4:20]
Mm.

**Swyx** [4:20]
Uh, don't know if you like that was ever tried or is there a more formal definition of it, because I just came up with it randomly. But it felt like the beginnings of AI, in the sense that when you run a CLI command, you have an intent to do something, and you may not have given the CLI all the things that it needs to do-

**Ben Firshman** [4:36]
Mm-hmm

**Swyx** [4:36]
... to execute that intent.

**Ben Firshman** [4:37]
Mm-hmm.

**Swyx** [4:38]
So that was my two cents.

**Ben Firshman** [4:40]
Yeah, that reminds me of a thing we f- we sort of, um, thought about when writing the CLI guidelines, where CLIs were designed in a wor- world where the CLI was really a programming environment, and it was primarily designed for like machines, like to use all of these commands and scripts.

Whereas over time, the CLI has evolved to humans. It was, you know, it's back in a world where like the primary way of like using an... computers was like writing shell scripts effectively. Um, and, uh, we've, we've transitioned to a world where actually humans are using CLI programs much more than they used to, and, um, and the current, the current sort of best practices about how Unix was designed, like, you know, there's, there's lots of sort of design documents about Unix from, from the '70s and '80s, where they, they say things like, "Command line commands should not output s- anything on success.

It should be completely silent." And which makes sense if you're using it in a shell script.

**Swyx** [5:42]
Yeah.

**Ben Firshman** [5:42]
But if a user is using that, it just looks like it's broken.

**Swyx** [5:45]
Yeah.

**Ben Firshman** [5:45]
Like if you type copy and it just doesn't say anything-

**Swyx** [5:48]
Yeah

**Ben Firshman** [5:48]
... you assume that it didn't work as a new user. Um, and yeah, so I, I think what's really interesting about the CLI is for, is that it's actually a really good, to your point, it's a really good user interface Where it can be like a conversation, where it feels like you're, instead of just like you telling the computer to do this thing, and either silently succeeding or saying, "No, you did-- failed," you know?

It can, like, guide you in the right direction and tell you what your intent might be and, and that kind of thing, in a way that's actually -- it's almost more natural to a CLI than it is in a graphical user interface, 'cause it feels like this back and forth with the computer.

**Swyx** [6:31]
Yeah.

**Ben Firshman** [6:32]
Um, almost funnily like, like a language model. Um, uh, so I think there's some, some interesting intersection of like CLIs and language models actually being very, very sort of, uh, uh, you know, closely related and a good fit for each other.

**Swyx** [6:47]
Yeah. I'll, I'll say, uh, one of the surprises from last year, um, you know, I worked on a coding agent, uh, but I think the most successful coding agent of my cohort was Open Interpreter, which was a CLI implementation.

### Arxiv Vanity

**Swyx** [6:57]
And, uh, I have chronically, even as a CLI person, I have chronically underestimated the CLI as a useful interface.

**Ben Firshman** [7:04]
Yeah. Yeah, totally.

**Swyx** [7:05]
Um, you also developed Archive Vanity-

**Ben Firshman** [7:06]
Yep

**Swyx** [7:06]
... which you recently retired after a glorious seven years of-

**Ben Firshman** [7:10]
Something like that, yeah

**Swyx** [7:11]
... something like that. Um, which is nice, and I guess HTML PDFs.

**Ben Firshman** [7:16]
Yep. That, that was actually the, the start of where Replicate came from.

**Swyx** [7:21]
Okay.

**Ben Firshman** [7:21]
Um-

**Swyx** [7:21]
We can tell that story

**Ben Firshman** [7:22]
... which, uh... So when I quit Docker, I got really interested in, um, science infrastructure, just as like a problem area, um, because it is... Like, science has created so much progress in the world. The fact that we're, you know, can talk to each other on a podcast and we use computers, and the fact that we're alive is probably thanks to medical research, you know?

But science is just like completely archaic and broken, and it's like 19th century processes that just happen to be, you know, copied to the internet rather than take into account that, you know, we can transfer information at the speed of light now.

Um, and the whole way science is funded, and all this kind of thing is all kind of very broken. Um, there's just so much potential for making science work better, and I realized that I wasn't a scientist, and I didn't really have any time to go and get a PhD and become a researcher, but I'm a tool builder, and I could make existing scientists better at their job.

And if I could make like a bunch of scientists a little bit better at their job, maybe, you know, that's the kind of equivalent of being a researcher. Um, so, um, one particular thing I dialed in on is just how science is disseminated, in that, um, uh, it's all of these, um, PDFs quite often behind paywalls, you know, on the internet.

Um, but-

**Swyx** [8:40]
And that's a whole thing.

**Ben Firshman** [8:41]
That's-

**Swyx** [8:41]
Because it's funded by national grants.

**Ben Firshman** [8:43]
Yep.

**Swyx** [8:44]
Government grants that are then put behind paywalls.

**Ben Firshman** [8:46]
Yeah, exactly. That's, that's like a whole... Yeah, I could talk for hours about that. But the particular thing we got, we got dialed in on was, um, or I, I got kind of, um, but interestingly, these PDFs are also, there's a bunch of open science that happens as well.

So maths, physics, computer science, machine learning notably, is all published on the Arxiv, which is, um, actually a surprisingly old institution.

**Swyx** [9:11]
Some random Cornell side project.

**Ben Firshman** [9:12]
Yeah, it was just like somebody in Cornell who started a mailing list in the '80s. And then when the web was invented, they built a web interface around it. Like it's super old. Um, and it-

**Swyx** [9:23]
It, it's like kind of like a use- user group thing, right? That's why there are all these like numbers and stuff.

**Ben Firshman** [9:26]
Yeah, exactly.

**Swyx** [9:27]
Uh-

**Ben Firshman** [9:27]
Like it's, uh, it's a bit like-

**Swyx** [9:29]
The special interest group

**Ben Firshman** [9:29]
... or something. Yeah. Um, and, um, that's where all, basically all of maths, physics, and computer science happens. Um, but it's still PDFs published to this thing.

**Swyx** [9:39]
Yeah.

**Ben Firshman** [9:39]
You know? Which is just so infuriating. Um, so, um, you know, like the, the, the, the web was invented at CERN, a physics institution, to share academic writing.

**Swyx** [9:52]
Mm.

**Ben Firshman** [9:52]
Like there are these, there are figure tags, there are like author tags, there are heading tags, there are cite tags. You know, hyperlinks are effectively citations-

**Swyx** [10:00]
Yes

**Ben Firshman** [10:00]
... because you want to link to another academic paper. But instead you have to like copy and paste these things and try and get around paywalls. Like it's absurd, you know? Um, like the, and now, now we have like social media and things, but still like academic papers as PDFs, you know?

It's just like why? This is not what the web was for. Um, so anyway, I got really frustrated with that, and I went on vacation with my old friend Andreas. So we were, we used to work together in London on a startup, at somebody else's startup, and we were just on vacation in Greece for fun, and he was like trying to read a machine learning paper on his phone.

You know, like we had to like zoom in and like scroll line by line on the PDF. And he was like, "This is fucking stupid." So- ... and I was like, "I know." Like this is something, we discovered our mutual hatred for, for this, you know?

And, uh, we spent our vacation sitting by the pool like making LaTeX to HTML like converters and making the first version of Archive Vanity. Um, anyway, that was, uh, been a whole thing, and, um, the story, we shut it down recently because I- they caught, caught the eye of Arxiv who were like, "Oh, this is great.

We just haven't had the time to work on this." And what's tragic about the Arxiv is it is, um, it is like this department of, it's like this project of Cornell that's like they can barely scrounge together enough money to survive.

I think it might be better funded now than it was when we were, we were collaborating with them. Um, and compared to these like scientific journals, it's just like this is actually where the work happens.

**Swyx** [11:23]
Yeah.

**Ben Firshman** [11:23]
But they just have a fraction of the money that like the, these, these big scientific journals have, which is just so tragic. Um, but anyway, they were like, "Yeah, this is great. We can't afford to like do it, but do you wanna like as a volunteer integrate Archive Vanity into Arxiv?"

Um-

**Swyx** [11:37]
Oh, you did the work.

**Ben Firshman** [11:38]
We didn't do the work.

**Swyx** [11:39]
Oh.

**Ben Firshman** [11:39]
We started doing the work.

**Swyx** [11:40]
Okay.

**Ben Firshman** [11:40]
We did some. I think we worked on this for like a few months to actually get it integrated into Arxiv.

**Swyx** [11:44]
Huh.

**Ben Firshman** [11:44]
Um, and then we got like distracted by Replicate.

**Swyx** [11:49]
Ah.

**Ben Firshman** [11:49]
So, um, a guy called Dayan picked up the work and made it happen. Um, like somebody who works on one of the, the piece, the libraries that powers Archive Vanity, um-

**Swyx** [11:58]
Okay

**Ben Firshman** [11:59]
... but yeah.

**Swyx** [11:59]
And the relationship with Archive Sanity?

**Ben Firshman** [12:02]
Um, none

**Alessio** [12:03]
Did, did you predate them? I, I actually don't know-

**Ben Firshman** [12:05]
We-

**Alessio** [12:05]
... the, the lineage

**Ben Firshman** [12:06]
... we were after-- We both were both users of ArxivSanity-

**Alessio** [12:08]
Okay

**Ben Firshman** [12:08]
... which is like a sort of archive an- aggregates by-

**Alessio** [12:10]
Which is Andreas', uh-

**Ben Firshman** [12:11]
I'm sure Andreas Babathi, yeah

**Alessio** [12:12]
... like Rexus on top of Arxiv.

**Ben Firshman** [12:13]
Yeah, yeah. And we were both users of that.

**Alessio** [12:15]
Yeah.

**Ben Firshman** [12:15]
And I think we were trying to come up with a working name for Arxiv.

**Alessio** [12:18]
Okay. All right.

**Ben Firshman** [12:18]
And Andreas just, like, cracked a joke of like, "Oh, let's call it ArxivVanity, 'cause it's making the papers look nice."

**Alessio** [12:23]
Yeah, yeah.

**Ben Firshman** [12:23]
And that was the working name, and it just stuck.

**Alessio** [12:25]
Got it. Got it.

**Ben Firshman** [12:27]
Um, yeah.

**Alessio** [12:28]
A- and then from there, tell us more about why you got distracted, right?

**Ben Firshman** [12:32]
Mm-hmm.

**Alessio** [12:32]
So Replicate maybe feels like an overnight success to a lot of people. Um, but you've been building this since 2019. Um-

**Ben Firshman** [12:39]
Yeah

**Alessio** [12:39]
... so what, what prompted the, the start?

**Ben Firshman** [12:41]
And we've been collaborating for even longer. So we created ArxivVanity in 2017. So in some sense, we've been doing this almost, like, six, seven years now. A classic seven-year overnight success.

**Alessio** [12:51]
Overnight success.

**Ben Firshman** [12:53]
Yeah. Uh, yeah, so we did ArxivVanity, and then worked on a bunch of, like, surrounding projects. I was still, like, really interested in science publishing at that point. Um, and I'm trying to remember, 'cause I tell a lot of, like, the condensed story to people, 'cause I can't really tell, like, a seven-year history, so I'm trying to figure out, like, the right -

**Alessio** [13:09]
Oh, we got room

**Ben Firshman** [13:09]
... the right, the right length to-

**Alessio** [13:11]
We wanna nail the, the definitive Replicate story here.

**Ben Firshman** [13:13]
One thing that's really interesting about these machine learning papers is that these machine learning papers are published on A- on the Arxiv, and a lot of them are actual fundamental research, so, like, should be, like, prose describing a theory.

But a lot of them are just running pieces of software that, like, a machine learning researcher made that did something. Um, uh, you know, it was like an image classification model or something, and they managed to make an image classification model that was better than the states of-- the existing state of the art.

And they've made an actual running piece of software that, that, that does image segmentation. And then what they had to do is they then had to take that piece of software and write it up as prose and math in a PDF.

Um, and what's frustrating about that is, like, if, if you wanna... So this was, like, Andreas's-- Andreas was a machine learning engineer at Spotify, and some of his job was, like, he did pure research as well. Like, he did a PhD, and he was doing a lot of stuff internally, but part of his job was also being an engineer and taking some of these existing things that people have made and published and trying to apply them to actual problems at Spotify.

And he was like, you know, you get given a p- a, a paper which, like, describes roughly how the model works. It's probably listing lots of crucial information. There's sometimes code on GitHub. More and more there's code on GitHub, but back in, back in, back then it was kind of relatively rare.

But it was quite often just, like, scrappy research code, and didn't actually run. Um, and you know, there was maybe the weights that were on Google Drive, but they accidentally deleted the weights off Google Drive, you know? And it was, like, really hard to, like, take this stuff and actually use it for real things.

And we just started talking together about, about, like, his problems at Spotify, and I connected this back to my work at, at Docker as well, and was like, "Oh, this is what we created containers for." You know, we solved this problem for normal software by putting the thing inside a container so that you could ship it around and it kept on running.

So we were, we were sort of hypothesizing about, like, "Hmm, what if we put machine learning models inside containers so that they could actually be shipped around, and they could be defined in, like, some production, production-ready format, and other researchers could run them to generate baselines, and you could-- people who wanted to actually apply them to real problems in the world could just pick up the container and run it," you know?

Um, and we then thought, this is probably where the, it gets... Normally, normally in this part of the story, I skip forward to be like, "And then we created Cog, this container standard- ... for, for machine learning models, and we created Replicate, the place for people to publish these machine learning models."

**Alessio** [15:51]
Yeah, exactly.

**Ben Firshman** [15:51]
But there's actually, like, two or three years between that.

**Alessio** [15:52]
Two years in between. Yeah.

**Ben Firshman** [15:54]
The, the thing we then got dialed into was Andreas was like, "What if there was a CI system for machine learning?" 'Cause, like, one of the things he really struggled with, with as a researcher is generating baselines.

**Alessio** [16:05]
Hmm.

**Ben Firshman** [16:05]
So when, like, he's writing a paper, he needs to, like, get, like, five other models that are existing work and get them running.

**Alessio** [16:13]
On the same evals.

**Ben Firshman** [16:14]
On the sa- exactly, on the same eval, so you can compare apples to apples-

**Alessio** [16:17]
Yeah

**Ben Firshman** [16:17]
... 'cause you can't trust the numbers in the paper.

**Alessio** [16:19]
Yeah.

**Ben Firshman** [16:19]
So, um-

**Alessio** [16:20]
Or you can be Google and just publish them anyway.

**Ben Firshman** [16:24]
Um, so he was like, "What if, what if you could..." I think this was coming from the thinking of, like, there should be containers for machine learning, but why are people gonna use that? Okay, maybe we can create a supply of containers by, like, creating this useful tool for researchers, and the useful tool was like, let's get researchers to package up their models and push them to this central place where we run a standard set of benchmarks across the models so that, um, you can trust those results, and you can compare these models apples to apples.

And for, like, a researcher, for Andreas, like, doing a new piece of research, he could trust those numbers, and he could, like, pull down those, pull down those, those models, con-confirm it on his machine, use the standard benchmark to then measure his model, and, you know, all this kind of stuff.

Um, and so we started building that. That's what we applied to YC with. Um, we got into YC, and we started sort of building a prototype of this. And then this is, like, where it all starts to fall apart.

We were like, "Okay, that sounds great." And we talked to a bunch of researchers, and they really wanted that, and that sounds brilliant. That's a great way to create a supply of, like, models on this research platform. But how the hell is this a business?

You know? Like, how are we even gonna make any money out of this? And we're like, "Oh, shit, that's, like, the-- that's the real unknown here of, like, what the business is." So we, um, we thought it would be a really good idea to like, okay, before we get too deep into this, let's try and, like, um, reduce the risk of this turning into a business.

So let's try and fi- like, research what the business could be for this, for this, uh, for, you know, for this research tool, effectively. So we went and talked to a bunch of companies trying to, trying to sell them something which didn't exist.

So we were like, "Hey, do you want a way to share research inside your company so that other researchers or say, like, the product manager-

**Alessio** [18:09]
Mm

**Ben Firshman** [18:09]
... can test out the machine learning model?" And they're like, "Uh, maybe." Um, and we were like, "Do you want-" A, like, a deployment platform for deploying models? Like, do you want, like, a central place for versioning models?

Like, we're trying to think of, like, lots of different, like, products we could sell that were, like, related to this thing. Um, and terrible idea. Like, we're not salespeople- ... and, like, people don't wanna buy something that doesn't exist.

**Swyx** [18:34]
Mm.

**Ben Firshman** [18:35]
Um, I think some people can pull this off, but we were just like, you know, a bunch of product people, product and engineer people, and we just, like, couldn't pull this off. Um, so we then got halfway through our YC batch.

We didn't have-- We hadn't built a product. We had no users. We had no idea what our business was gonna be, 'cause we couldn't get anybody to, like, buy something which didn't exist. Um, and actually, this was quite a way through our...

I think it was, like, two-thirds of the way through our YC batch or something, and we're like, "Okay, well, we're kinda screwed now, um, 'cause we don't have anything to show at demo day." And then we then, like, tried to figure out, okay, what can we build in, like, two weeks that'll be something?

So we, like, desperately tried to... I can't remember what we tried to build at that point. Um, and then two weeks before demo day, I just remember this, um, um... I remember it was all, it was all-- We were going down to Mountain View every week for dinners, and we got called onto, like, an all-hands Zoom call, which was super weird.

We're like, "What's going on?" And they were like, "Don't come to dinner tomorrow." Um, and we realized, we kind of looked at the news and we were like, "Oh, there's a pandemic going on." We were, like, so deep in our startup, we were just, like, completely oblivious to what was going on around us.

**Swyx** [19:44]
Was this-

**Ben Firshman** [19:45]
Um-

**Swyx** [19:45]
... Jan or Feb 2020?

### YC & Pivot

**Ben Firshman** [19:47]
This was March 2020.

**Swyx** [19:48]
March 2020.

**Ben Firshman** [19:49]
2020, yeah.

**Swyx** [19:49]
'Cause I remember Silicon Valley at the time was early to COVID.

**Ben Firshman** [19:53]
Yep.

**Swyx** [19:53]
Like-

**Ben Firshman** [19:53]
Yeah

**Swyx** [19:54]
... they started locking down a lot faster than the rest of US.

**Ben Firshman** [19:55]
Yeah, exactly. And I remember, yeah, soon after that, like, there was the San Francisco lockdowns, and then, like, the YC batch just, like, stopped. There wasn't demo day. Um, and it was in a, in a sense a blessing for us, 'cause we just kind of-

... couldn't raise money anyway. Um-

**Swyx** [20:12]
In, in the normal course of events, you can-- you're actually allowed to defer-

**Ben Firshman** [20:15]
Yeah, exactly

**Swyx** [20:15]
... to a future demo day.

**Ben Firshman** [20:16]
Yep.

**Swyx** [20:16]
Yeah.

**Ben Firshman** [20:17]
So we didn't even take any defer 'cause it just-

**Swyx** [20:18]
Yeah

**Ben Firshman** [20:18]
... kinda didn't happen, you know? So, um, so-

**Swyx** [20:22]
So was YC helpful?

**Ben Firshman** [20:24]
Yes. We completely screwed up the batch, and that was our fault.

**Swyx** [20:27]
Okay.

**Ben Firshman** [20:27]
I think the thing that YC has become incredibly valuable for us has been after YC. Um, and I think, I think reason-- You could-- There was a reasonable argument that we sh- couldn't, didn't need to do YC to start with because we were quite experienced.

We had done some startups before. We were kind of well connected with VCs. You know, it was relatively easy to raise money 'cause we were, like, a n-unknown quantity. You know, if you go to a VC and be like, "Hey, I made this piece of, piece of-"

**Swyx** [20:53]
It's Docker Compose for AI. AI.

**Ben Firshman** [20:55]
Exa-exactly, yeah. And, and, and like, you know, people can pattern match like that, and they can sort of have some trust you know what you're doing. Um, whereas it's much harder for people straight out of college, and that's where, like, YC's sweet spot is, like, helping people straight out of college who are super promising, like, figure out how to do that.

**Swyx** [21:08]
Yeah, no credentials.

**Ben Firshman** [21:09]
Yeah, exactly.

**Swyx** [21:10]
Yeah.

**Ben Firshman** [21:10]
So in some sense, we didn't need that, but the thing that's been incredibly useful for us since YC has been... This was actually, I think-- So Docker was, Docker was a YC company, and Solomon, the founder of Docker, I think, told me this.

He was like, um, "A lot of people underestimate the value of YC after you finish the batch." And he, his biggest regret was, like, not staying in touch with YC. I might be misattributing this, but I think it was him.

And so we made a point of that, and we just stayed in touch with our batch partner, who, um, Jared at YC, who's been fantastic.

**Swyx** [21:42]
Jared Harris?

**Ben Firshman** [21:43]
Um, Jared Friedman.

**Swyx** [21:44]
Friedman.

**Ben Firshman** [21:45]
And all of, like, the team at YC, like, there was the growth team at YC when they, when they were still there, and they've been super helpful. Um, and, um, two things have been super helpful about that is, like, raising money.

Like, they just know exactly how to raise money, and they've been super helpful during that process in all of our rounds. Like, we've done three rounds since we did YC, and they've been super helpful during the whole process.

Um, and also just, like, reaching a ton of customers. So, like, the magic of YC is that you have all of-- Like, there's thousands of YC companies, I think. Like, on the-

**Swyx** [22:14]
Thousands

**Ben Firshman** [22:15]
... order of thousands, I think.

**Swyx** [22:15]
Yeah, yeah.

**Ben Firshman** [22:15]
Um, and they're all of your first customers. And they're, like, super helpful, super receptive, really want to, like, try out new things. Um, you have, like, a warm intro to every, every one of them, basically, and there's this mailing list where you can post about updates to your, um, to your product, um, which is, like, really receptive, and that's just been fantastic for us.

Like, we've, we've just, like, got so many of our, of our users and customers through, through YC. Um-

**Swyx** [22:42]
Yeah, well, so the classic criticism or the sort of, you know, pushback is people don't buy you because, um, you are both from YC, but at least they'll open the email.

**Ben Firshman** [22:53]
Yeah.

**Swyx** [22:53]
Right? Like, that's the-

**Ben Firshman** [22:54]
Yeah.

**Swyx** [22:54]
Okay.

**Ben Firshman** [22:55]
Yeah, effectively. Um, and

yeah. Yeah, so that's been a really, really positive experience for us.

**Swyx** [23:02]
Mm-hmm. And, and sorry, I interrupted, uh, with the YC question. Like, you were-- You were making-- You just made it out of the YC-

**Ben Firshman** [23:07]
Oh, yeah

**Swyx** [23:07]
... survived the pandemic. Um, and you, yeah.

**Ben Firshman** [23:11]
I'll try and condense this a little bit. Then we, then we started building tools for COVID, weirdly. We were like, "Okay, we don't have a startup. We haven't figured out anything. What's the, what's the most useful thing we could be doing right now?"

### Community Launch

**Swyx** [23:21]
Save lives.

**Ben Firshman** [23:22]
So yeah, let's try and, let's try and save lives. I think we failed at that as well. We had a bunch of projects that didn't really go anywhere. Um, uh, we kind of worked on, yeah, a bunch of stuff like contact tracing, which turned out didn't really to be a useful thing.

Um, sort of, uh, Andreas worked on a, like, a, um, like a, a DoorDash for, like, people delivering food to people who are vulnerable. Uh, what else did we do? The meta problem of, like, helping people direct their efforts to what was most useful.

Um, and a few other things like that. Didn't really go anywhere. So we're like, "Okay, this is not really working either." Um, we, we were considering actually just, like, doing, like, work for COVID. We had, like, this decision document early on in our company, which is like, should we become a, like, government app contracting shop, you know?

**Swyx** [24:06]
Mm.

**Ben Firshman** [24:07]
Um, we decided no.

**Swyx** [24:08]
Because you also did, uh, work for the US, uh, for the gov.uk.

**Ben Firshman** [24:11]
Yeah, exactly. We had experience, like, doing some, like, uh-

**Swyx** [24:15]
And The Guardian and, you know, that

**Ben Firshman** [24:16]
... yeah, for, like, government stuff. Um, and we were just, like, really good at building stuff. Like, we were just, like, product people. Like, I was, like, the front-end product side, and Andreas was the back-end side. So we were just, like, a, a product-- And we were working with a designer at the time, um, a guy called Mark, who did our early designs for Replicate.

And we were like, "Hey, what if we just team up and, like, become and, and build stuff?" And But yeah, we gave up on that in the end for-- I can't remember the details. Um, so we c- went back to machine learning, and then we were like, "Oh, well, we're not really sure if this is gonna work."

And one of my most painful experiences from previous startups is shutting them down. Like, when you realize it's not really working and having to shut it down, it's like a ton of work, and it's-- people hate you, and it's just sort of, you know...

Um, so we were like, "How can we make something we don't have to shut down? And even better, how can we make something that won't page us in the middle of the night?"

Uh, so we made an open source project. We made a thing which was an open source weights and biases, um, 'cause we had this theory that, like, weights and-- that, that, like, people want open source tools. There should be, like, an open source, like, version control experiment tracking, like, thing.

And it was intuitive to us in that we're like, "Oh, we're software developers, and we like command line tools." Like, everyone likes command line tools and open source stuff. But machine learning researchers just really didn't care. Mm-hmm. Like, they just wanted to click on buttons.

They didn't mind that it was a cloud service. Like, it was all very visual as well, that you needed lots of graphs and, and charts and stuff like this. So it just didn't-- it wasn't right. Like, it was right-- We were actually rebuilding something that Andreas made at Spotify for just, like, saving experiments to cloud storage automatically, but other people didn't really want this.

So we kinda gave up on that, and then we-- That was actually originally called Replicate, and we renamed that out the way, so it's now called Keepsake, and I think some people still use it. Then we sort of came back-- We looped back to our original idea.

So we were like, "Oh, maybe there was a thing in that thing we were originally sort of thinking about of, like, researchers sharing their work and containers for machine learning models." So we just built that, and at that point, we were kind of running out of the YC money, so we were like, "Okay, this, like, feels good though.

Let's, like, give this a shot." So that was the point we raised a seed round. We raised, um, seed round- Pre-launch. We raised pre-launch. Pre-launch and pre-team. Um, it was an idea, basically. We had a little prototype. It was just an idea and a team.

Um, but we were like, "Okay," like, you know, when-- "bootstrapping this thing is getting hard, so let's actually raise some money." Um, and then we made Cog and Replicates. Mm. It initially didn't have APIs, interestingly. It was just the bit that I was talking about before of helping researchers share their work.

So it was a way for researchers to, to put their work on a webpage such that other people could try it out, uh, and so that you could download the Docker container. So that, like, we didn't have-- We cut the benchmarks thing of it 'cause we thought that was just, like, too complicated.

But it had a Docker container that, like, you know, Andreas in a past life could download and run with his benchmark, and you could compare all these models apples to apples. So that was, like, the theory behind it.

Um, and that kind of started to work. It was, like, still when, like, you know, it was pre-- long time pre-AI hype, and there was lots of interesting stuff going on, but it was, it was very much in, like, the classic deep learning era.

So sort of image segmentation models and sentiment analysis and all these kind of things, you know, that people were using, uh, that were using deep learning models for. And we were very much building for research 'cause all of this stuff was happening in research institutions.

You know, there's some people who'd be publishing to Arxiv. So we were, we were creating an accompanying material for their models, basically. You know, they wanted a demo for their models, and we were creating accompanying material for it.

Um, and they were like-- What was funny about that is they were, like, not very good users. Like, they were, they were doing great work obviously, but, but the way that research worked is that they, they just made, like, one thing every six months, and they just fired and forget it, forgot it.

Yeah. Like, they, they published this piece of paper, and, like, done. I'm, I've, I've published it. Um, so they, like, output it to Replicate, and then they just stopped using Replicate. Yeah. You know? They were, like, once every six monthly users.

And that wasn't great for us. Um, but we stumbled across this early community. This was early 2021 when

people started-- OpenAI created this-- created Clip, and people started smushing Clip and GANs together to produce image generation models. And this started with, um, you know, it was just a bunch of, like, tinkerers on Discord, basically. Um, it was, um-- There was an early model called Big Sleep by AdvadNaun, and then there was VQGAN-CLIP, which was, like, a bit more popular, by RiversHaveWings.

And it was all just people, like, tinkering on stuff in Colabs- Right ... and it was very dynamic, and it was people just making copies of Colabs and playing around with things and forking and... And to me, this-- I saw this and I was like, "Oh, this feels like open source software."

Like, so much more than the research world- Mm-hmm ... where, like, people are publishing these papers. Yeah, you don't know their real names, and then it's just, like, a Discord ID. Yeah, exactly. But crucially, it was like people were tinkering and forking, and people were- Yeah ...

things were moving really fast, and- Yeah ... um, it just felt like this creative, dynamic, collaborative community in a way that research wasn't really. Like, it was still stuck in this kind of six-month publication cycle. So we just kinda latched onto that and started building for this, this community.

Um, and you know, a lot of those early models were published on Replicates. No-- I think the first one that was really primarily on Replicates was one called Pixray, which was sort of, sort of mid-2021. Um, and it had a really cool, like, pixel art output, but it also just, like, produced-- They weren't, like, crisp in images, but they were quite aesthetically pleasing- Oh ...

like some of these early image generation models. And, um, um, you know, that was, like, published primarily on Replicates, and then a few other models around that were, like, published on Replicates. And that's where we really started to find our early community- Okay ...

and, like, where we really found, like, oh, we've actually built a thing that people want. Um, and they were great users as well, and people really wanna try out these models. Lots of people were, like, running the models on Replicate.

We still didn't have APIs though. Interestingly- ... and this is like another like really complicated part of the story. We had no idea what our business model was still at this point. I don't think you- people could even pay for it.

You know, it's just like these web forms where people could run the model. Um, and-

**Swyx** [30:47]
Just before this API bit, uh, continue. Uh, just for-

**Ben Firshman** [30:49]
Yeah

**Swyx** [30:49]
... historical interests, uh, which Discords were they, and how did you find them? Was this the Lion Discord?

**Ben Firshman** [30:54]
Yeah, Lion-

**Swyx** [30:54]
Was this Luther?

**Ben Firshman** [30:55]
Luther, yeah. It was the Luther one-

**Swyx** [30:56]
These two, right?

**Ben Firshman** [30:57]
Luther I particularly remember. There was a channel where, where VQGAN-CLIP-- This was early 2021, where VQGAN-CLIP was set up as a, as a Discord bot, and I just remember being completely just like captivated by this thing. I was just like playing around with it all afternoon, and like the sort of thing where-

**Swyx** [31:15]
In Discord

**Ben Firshman** [31:15]
... you're like, "Oh, shit, it's 2:00 AM," you know .

**Swyx** [31:16]
Yeah. This is the beginnings of Midjourney.

**Ben Firshman** [31:18]
Yeah, exactly. And it was-

**Swyx** [31:19]
And, and stability

**Ben Firshman** [31:20]
... it was the start of-- It was the start of Midjourney, and, you know, it's where that kind of user interface came from. Like what's beautiful how the user interface is like you could see what other people are doing.

**Swyx** [31:29]
Yeah.

**Ben Firshman** [31:29]
And that you could, you could riff off other people's ideas, and it was just so much fun to just like play around with this in like a channel full of 100 people. Uh, and yeah, that just like completely captivated me, and I'm like, "Okay, this, this is like s- this is something," you know?

So like we should get these things on Replicate. Um, and yeah, that's, that's where that, that all came from. Yeah.

**Swyx** [31:49]
Okay. Sorry, uh, and I just wanted to capture that moment.

**Ben Firshman** [31:52]
Yeah, yeah.

**Swyx** [31:53]
Um, um, and then you moved on to-- So was it APIs next, or was it Stable Diffusion next?

**Ben Firshman** [31:57]
It was APIs next, and the APIs happened because one of our users-- Our web form had like an internal API for making the web form work, like with, uh, an API that was called from JavaScript. And somebody like reverse engineered that to start generating images with a script.

### API & Growth

**Ben Firshman** [32:15]
You know, they did like-

**Swyx** [32:16]
Mm-hmm

**Ben Firshman** [32:16]
... you know, web inspector copy-

**Swyx** [32:18]
Some kind of copyright thing

**Ben Firshman** [32:18]
... as Carl, like figured out what the-

**Swyx** [32:19]
Oh, I see. I see

**Ben Firshman** [32:20]
... API request was.

**Swyx** [32:21]
Got it. Yep.

**Ben Firshman** [32:22]
Um, and it wasn't secured or anything. Um-

**Swyx** [32:25]
Of course not

**Ben Firshman** [32:25]
... and they started generating a bunch of images, and like we got tons, tons of traffic, and we're like, "What's going on?" Um, and I think like a s- a sort of usual reaction to that would be like, "Hey, you're abusing our API," and to shut them down.

Instead, we were like, "Oh, this is interesting. Like, people wanna run these models." Um, so we documented the API in a Notion document, like our internal API in a Notion document, and like messaged this person being like, "Hey, you seem to have found our API."

"Um, here's the documentation. That'll be like 1,000 bucks a month, please" With a stripe form, like that we just click some buttons to make. Um, and they were like, "Sure, that sounds great." So that was our first customer .

Um-

**Swyx** [33:09]
1,000 bucks a month?

**Ben Firshman** [33:10]
Uh, it was, it was a surprising amount of money, yeah.

**Swyx** [33:12]
That's not-

**Ben Firshman** [33:12]
It was on the order-

**Swyx** [33:13]
... casual

**Ben Firshman** [33:13]
... it was on the order of 1,000 bucks a month.

**Swyx** [33:15]
So was he-- Was it a business? Like what-

**Ben Firshman** [33:17]
It was the creator of PixRay. Like it was-

**Swyx** [33:20]
Oh

**Ben Firshman** [33:21]
... he generated NFT art, and so he like made a bunch of art with these models-

**Swyx** [33:27]
Mm

**Ben Firshman** [33:27]
... and, um, was, was, you know, selling these NFTs effectively. And I think p- lots of people in his community were doing similar things, and like he then referred us to other people who were also generating, uh, generating NFTs using generative models, and that was like the start of like, uh-- That was the start of, start of our API business, yeah.

And then we, then we like made an official API and actually like added some, some billing to it. Uh, so it wasn't just like a fixed fee and yeah.

**Swyx** [33:53]
And now people think of you as the hosted models API business.

**Ben Firshman** [33:56]
Yep, exactly. And, and, and-- But that just turned out to be our business. You know-

**Swyx** [33:59]
Yeah

**Ben Firshman** [34:00]
... but what, but what, what ended up being beautiful about this is it was really fulfilling like the original goal of what we wanted to do, is that we wanted to make this research that people were making accessible to like other people, and for it to be used in the real world.

And this was like the-- just like ultimately the right way to do it because all of these people making these generative models could publish them to Replicate, and they wanted a place to publish it. And software engineers, you know, like myself, like I'm not a machine learning expert, but I wanna use this stuff, uh, could just run these models with a single line of code.

And we thought, "Ah, maybe the Docker image is enough," but it's actually super hard to get the Docker image running on a GPU and stuff. So it really needed to be the hosted API for this to work and to make it accessible to software engineers, and we just like w- wound our way to this, this-

**Swyx** [34:46]
Yeah, two years to the first paying customer

**Ben Firshman** [34:48]
... this solution. Yeah, exactly. Um-

**Swyx** [34:50]
D- did you ever think about becoming Midjourney during that time? You have like-

**Ben Firshman** [34:54]
Mm

**Swyx** [34:54]
... so much interest in image generation-

**Ben Firshman** [34:55]
What could have been-

**Swyx** [34:56]
... it's

**Ben Firshman** [34:57]
Yeah.

**Swyx** [34:57]
I mean, you're doing fine d- for the record, but you know. It was right there. You were playing with it.

**Ben Firshman** [35:04]
Yeah, I don't, I don't think it was our expertise.

**Swyx** [35:07]
Okay.

**Ben Firshman** [35:07]
Like I think our expertise was dev tools rather than-- Like Midjourney's almost like a consumer product, you know?

**Swyx** [35:11]
It is, yeah.

**Ben Firshman** [35:12]
Um, so I don't think it was our expertise. Uh, it certainly occurred to us. Um, I think at the time we were thinking about like, "Oh, maybe we could hire some of these commun- people in this community and make great models," and stuff like this, but just ended up our-

**Swyx** [35:24]
Mm-hmm

**Ben Firshman** [35:24]
... we, we ended up more being at the tooling. Like I think-

**Swyx** [35:27]
Yeah

**Ben Firshman** [35:27]
... like before I was saying like I'm not really a researcher, but I'm more like the tool builder-

**Swyx** [35:30]
Mm-hmm

**Ben Firshman** [35:30]
... like behind the scenes, and I think both me and Andreas are like that. Yeah.

**Swyx** [35:33]
Yeah.

**Alessio** [35:33]
Yeah.

**Swyx** [35:33]
I, I think this is a-- also like a illustration of the tool builder philosophy, something where you, you're very-- you, you latch onto in dev tools, which is when you see people behaving weird, it's not your f- it's not their fault, it's yours.

Like you, you-- And, and you wanna pave the cow paths is what they say, right? Like the, the unofficial paths that people are making, like make it official and make it easy for them, and then maybe charge a bit of money.

**Ben Firshman** [35:52]
Mm-hmm.

**Alessio** [35:53]
Yep.

**Ben Firshman** [35:53]
Yeah.

**Alessio** [35:54]
Um, and now fast-forward a couple of years, you have two million developers using Replicate. Maybe more. That, that was the last public number that I found.

**Ben Firshman** [36:02]
Two million. I think that got mangled actually by-- It's two million users. Not all those people are developers, but a lot of them are developers, yeah.

**Alessio** [36:09]
Um, and then 30,000 paying customers was the number. Um, that's, that's awesome. Uh, Latent Space runs on Replicate.

**Ben Firshman** [36:17]
Right. Nice.

**Alessio** [36:17]
So we have a small podcaster, and we host, uh, whisper-

**Swyx** [36:19]
We do a transcription on Replicate

**Alessio** [36:20]
... whisper diarization on, on Replicate.

**Swyx** [36:23]
Cool.

**Alessio** [36:23]
Um, so-- And we're paying, so we're-- Latent Space- ... is in the 30,000.

**Swyx** [36:27]
Nice. Thank you.

**Alessio** [36:27]
Um, you raised a $40 million Series B. Um, I would say that maybe the Stable Diffusion time, August '22, was like really when the company started to-

**Ben Firshman** [36:38]
Yeah

**Alessio** [36:38]
... to break out. Um, tell us a bit about that and the community that came out, and I know now you're expanding beyond just, uh, image generation.

**Ben Firshman** [36:46]
Yeah. This-- Like, I think we kind of set ourselves-- Like, we saw there was this really interesting ge-image, generative image world going on, so we kind of, you know, like, we're, we're building the tools for that community already, really.

And we knew Stable Diffusion was coming out. We knew it was a really exciting thing. You know, it was the best, the best generative image model so far. I think the thing we didn't-- we underestimated was just, like, what an inflection point it would be, where it was an-- It was-- I think, I think Simon Willison put it this way, where he said something along the lines of, it was a model that was open source and tinkerable, and, like, good enough that it was just, like...

It was, it was, you know, it was just good enough and open source and tinkerable, such that it just kind of took off-

**Swyx** [37:35]
Mm-hmm

**Ben Firshman** [37:35]
... in a way that none of the models had before. And, like, what was really neat about Stable Diffusion is, is it was open source, so you could, like-- Compared to, like, DALL-E, for example, which was, like, sort of equivalent qua-quality, you-- it was open source, so you could fork it and tinker on it.

And, like, the first week, we saw, like, people making animation models out of it. We saw people make, like, game texture models that, like, use circular convolutions to make repeatable textures. We saw-- What else did we see? Um, you know, a few weeks later, like, people were fine-tuning it, so you could make-- put your face in these models and, um, all of these other-

**Swyx** [38:10]
Yeah, textual inversion.

**Ben Firshman** [38:11]
Yep. Yeah, exactly. That happened a bit before that. And all of this sort of innovation was happening all of a sudden, and people were publishing it on Replicate because you could just, like, publish arbitrary models on Replicate, so we had this sort of supply of, like, interesting stuff being built.

But because it was a sufficiently good model, um, there was also just, like, a ton of people building with it. They were like, "Oh, we can build products with this thing." And this was, like, about the time where people were starting to get really interested in AI, so, like, a ton of the product builders wanted to build stuff with it.

And we were just, like, sitting in there in the middle as, like, the interface layer between, like, all these people who wanted to build and all these, like, machine learning experts who were building cool models. Um, and that's, like, really where it took off.

We were just sort of credible supply and credible demand, and we were just, like, in the middle. Um, and then, yeah, since then, we've just kind of grown and grown, really. And we, we, you know, have been building a lot for, like, the indie hacker community, these, like, individual tinkerers, but also startups and a lot of large companies as well who are sort of exploring and building AI things.

And then kind of the same thing happened, like, middle of last year with language models and Llama 2, where the same kind of Stable Diffusion effect happened with, with Llama. And Llama 2 was, like, our biggest week of growth ever because, like, tons of people wanted to tinker with it and run it.

And, you know, since then, we've just been seeing a ton of growth in language models as well as image models. And, uh, yeah, we're just kind of riding a lot of the, the interest that's going on in AI and all the people building in AI, you know?

**Swyx** [39:37]
That's, uh-- Yeah, kudos. Right place, right time, but also, you know, took a while to position for the, for the right, uh, place before the wave came. Um, I, I'm, I'm curious if, like, um, you have any insights on these different markets.

Um, so Peter Levels, notably a very loud person, uh, very picky about his tools. Um, I wasn't sure actually if he used you. He does.

**Ben Firshman** [40:00]
He does, yeah.

**Swyx** [40:00]
Because you cited him, you cited him on your series V blog post, and Danny Postma as well, his competitor-

**Ben Firshman** [40:04]
Yeah

**Swyx** [40:05]
... um, all, all in that wave. Um, what are their needs versus, um, you know, the more enterprise or B2B type needs? Did, did you, did you come to a decision point where you're like, "Okay, you know, how serious are these indie hackers versus, like, the actual businesses that are bigger and perhaps better customers because they're less churny?"

**Ben Firshman** [40:24]
They're surprisingly similar.

**Swyx** [40:25]
Okay.

**Ben Firshman** [40:26]
Because I think a lot of people right now want to use and build with AI, but they're not AI experts, and they're not infrastructure experts either. So they wanna be able to use this stuff without having to, like, figure out all the internals of the models and, you know, like, touch PyTorch and whatever.

And they also don't wanna be, like, setting up and booting up servers. Um, and that's the same all the way from, like, indie hackers just getting started, because, like, obviously you just wanna get started as quickly as possible, all the way through to, like, large companies who wanna be able to use this stuff but don't have, like, all of the experts on staff, you know?

Um, like, I think some, some companies are quite, you know, there are companies, big companies like Google and so on that do actually have a lot of experts on staff, but the vast majority of companies don't. And they're all software engineers who wanna be able to use this AI stuff, but they just don't know how to use it.

And it's like, you really need to be an expert, and it takes a long time to, like, learn the skills to be able to use that. So they're surprisingly similar in that sense. Um, and I think, I think it's kind of also unfair of, like, the indie community.

Like, surpris- They're not churning, surprisingly, or churny or spiky, surprisingly. Like, they're building real established businesses, which is like kudos to them, like, of, like, building these, like, really, like, large, sustainable businesses, often just as, like, solo developers.

Uh, and it's kind of remarkable how they can do that, actually, and it's a credit to a lot of their, like, their product skills and, you know, we're just, like, there to help them, being like their machine learning team, effectively, uh, to help them use all of this stuff.

Um, so we're actually making some-- Like, like, a lot of these indie hackers are some of our largest customers, like alongside- ... some of our biggest customers that you would think would be-

**Swyx** [42:08]
Yeah

**Ben Firshman** [42:08]
... would be, would be, uh, would be, uh, you know, spending a lot more money than them, but yeah.

**Swyx** [42:14]
Uh, and we should name some of these. You have them on your landing page. You have BuzzFeed, you have Unsplash, uh, Character AI. Um, how-- Like, what do they power? What, what can you say about their, their usage?

**Ben Firshman** [42:24]
Yeah, totally. It's, it's kind of, uh, v-various things. I'm trying to think. Um,

let me actually think. What can I say about what customers?

**Swyx** [42:36]
Well, I, I mean, I'm naming them because they're on your landing page-

**Ben Firshman** [42:38]
Yeah

**Swyx** [42:38]
... so, so you have logo rights.

**Ben Firshman** [42:39]
Yeah.

**Swyx** [42:40]
Um, it's u- It's useful for people to who are-- Like, I, I'm not imaginative. I, I see... Monkey see, monkey do, right? Like if, if I see-

**Ben Firshman** [42:46]
Yeah, yeah

**Swyx** [42:46]
... someone doing something that I wanna do, then I'm like, "Okay, Replicate's great for that."

**Ben Firshman** [42:49]
Yeah, yeah, yeah.

**Swyx** [42:50]
So s- that's what I think about case studies on company landing pages is that it's just a way of explaining, like, "Yep, we, we-- this is something that we are good for."

**Ben Firshman** [42:58]
Yeah, totally. It-- I mean, it's- These companies are doing things all the way up and down the stack at different levels of sophistication. So like Unsplash, for example, they, they actually, they actually publicly posted this story on Twitter where they're using, uh, BLIP to annotate all of the images in their catalog.

So, you know, they have lots of images in the catalog, and they wanna create a text des- description of it so you can search for it. Um, and they're annotating the images with, you know, off-the-shelf open source model.

You know, we have this big library of open source models that you can run, and, you know, we've got lots of people who are running these open source models off the shelf. And then, you know, most of our larger customers are doing more sophisticated stuff, so they're like fine-tuning the models, they're running completely custom models on us.

And so a lot of these, a lot of these larger companies are, like, using us for a lot of their, their, you know, inference, but it's like a lot of custom models and them, like, writing the Python themselves 'cause they've got machine learning experts on team, on the team, and they're using us for like, you know, their inference infrastructure effectively.

Um, so it's like lots of different levels of sophistication where like some people are using these off-the-shelf models, some people are fine-tuning models. So like level, Peter Level is a great example where a lot of his products are based off like fine-tuning, fine-tuning image models, for example.

And then we've also got like larger customers who are just like using us as infrastructure effectively as, as servers. Um, so yeah, it's like all things up and down, up and down the stack.

**Alessio** [44:29]
Yeah. Um, let's talk a bit about Cog and the, the technical layer. So there are a lot of, uh, GPU clouds, uh, I think people at different pricing points, and I think everybody tries to offer a different developer experience on top of it, which then lets you charge a premium.

### Cog & Standards

**Alessio** [44:46]
Why did you wanna create Cog? What were some of the... You worked at Docker. What were some of the issues with traditional container runtimes? Um, and maybe, yeah, what, what were you surprised with as you built it?

**Ben Firshman** [44:57]
Cog came right from the start actually, when we were thinking about this, this, you know, evaluation, this sort of benchmarking system for machine learning researchers, where we wanted researchers to publish their models in a standard format that was guaranteed to keep on running, that you could replicate the results of, like that's where the name came from.

**Alessio** [45:19]
Mm-hmm.

**Ben Firshman** [45:20]
And we realized that we needed something like Docker to make that work, you know. Um, and I think it was just like natural from my point of view of like, obviously, that should be open source, that we should try and like create some kind of open standard here that people can share, because if more people use this format, then that's great for everyone involved.

Um, you know, I think, I think the magic of Docker is not really in the software.

**Alessio** [45:42]
Mm-hmm.

**Ben Firshman** [45:42]
It's just like the standard that people have agreed on, like, here are a bunch of keys for a JSON document-

**Alessio** [45:48]
Right. Yeah.

**Ben Firshman** [45:48]
... basically. And, um, you know, that was the magic of like the metaphor of real containerization as well. It's not the containers that are interesting. It's just like the size and shape of the damn box, you know?

**Alessio** [45:58]
Mm-hmm. Right. Yeah.

**Ben Firshman** [45:59]
Um, and it's similar thing here, where really we just wanted to get people to agree on like, this is what a machine learning model is. This is, this is how a prediction works. This is what the inputs are.

This is what the outputs are. So Cog is really just a Docker container that attaches to a CUDA device if it needs a GPU, that has a OpenAPI specification as a label on the Docker image.

**Alessio** [46:22]
Right.

**Ben Firshman** [46:22]
And the OpenAPI specification defines the interface for the machine learning model, like the, the, um, the inputs and outputs effectively, or the, the params in machine learning terminology. Um, and you know, we just tried, wanted to get people to kind of agree on this thing, and it's like general purpose enough.

Like we weren't saying like some of the existing things were like at the graph level.

**Alessio** [46:45]
Mm-hmm.

**Ben Firshman** [46:45]
But we really wanted something general purpose enough that you could just put anything inside this, and it was like future compatible, and it was just like arbitrary software, and, you know, be future compatible with like future inference servers and future machine learning model formats and all this kind of stuff.

**Alessio** [46:57]
Yeah.

**Ben Firshman** [46:57]
Um, so that was the intent behind it. And, you know, for... It just came naturally that we wanted to define this format and, and that's been really working for us. Like a bunch of people have been using Cog outside of Replicates, which is kind of our original intention.

**Alessio** [47:13]
Oh, wonderful.

**Ben Firshman** [47:13]
Like this should be how machine learning models are packaged and how people should use it. Like it's common to use Cog in situations where like maybe they can't use the SaaS service because-

**Alessio** [47:23]
Uh-huh

**Ben Firshman** [47:23]
... I don't know, they're in a big company and they're not allowed to like, you know, not allowed to use a SaaS, SaaS service, but they can use Cog internally still, and like they can download the models from Replicates and run them internally in, in their org, which we've been seeing happen.

That works really well. Um, people who wanna build like custom inference pipelines but don't wanna like reinvent the world, so they can use Cog off the shelf and use it as like a component in their inference pipelines. Um, we've been seeing tons of, tons of usage like that.

Um, and it's just been kind of happening organically. We haven't really been trying, you know, but it's like there if people want it, and we've been seeing people use it, so that's great. Um-

**Alessio** [47:57]
Yeah

**Ben Firshman** [47:57]
... and, uh, yeah, so a lot of it's just sort of philosophical of just like, this is, this is how it should work from my experience at Docker, you know?

**Alessio** [48:03]
Yeah.

**Ben Firshman** [48:03]
And there's just a lot of value from like the core being open, I think, and that other people can share it, and it's like an integration point. So, you know, if, if Replicate, for example, wanted to, wanted to work with a testing system, like a CI system or whatever, um, y- we can just like interface at the Cog level.

Like-

**Alessio** [48:20]
Mm-hmm

**Ben Firshman** [48:21]
... that, that system just needs to pull Cog models, and then you can like test your models on that CI system before they get deployed to Replicate, and it's just like a format that everyone-- we can get everyone to agree on.

Yeah.

**Alessio** [48:30]
W- what do you think, I guess, Docker got wrong? Because if I look at a Docker Compose and a Cog definition, first of all, the, the Cog is kinda like the Docker file plus the Compose-

**Ben Firshman** [48:40]
Yeah

**Alessio** [48:40]
... versus, and Docker Compose are just exposing the services. And also Docker Compose is very like, uh, ports driven versus-

**Ben Firshman** [48:48]
Mm-hmm

**Alessio** [48:48]
... you have like the actual, you know, predict this is what you have to run. Yeah, any learnings and maybe tips for other people building container-based runtimes? Like how, how much you should just separate the API services versus the, the image building or how much you wanna build them together?

**Ben Firshman** [49:07]
I think it was coming from two sides. We were thinking about the design from the point of view of user needs, like what do users-- what are their problems and what, what problems can we solve for them, but also what the interface should be for a machine learning model, and it's sort of the combination of two things that led us to this design.

So the thing I talked about before was a little bit of, like, the interface around the machine learning model. So we realized that we want it to be general purpose. We want it to be at the, like, the JSON, like, human readable things rather than the, the tensor level.

Um, so it's like an OpenAPI specification that wrapped a Docker container. That's where that design came from, and it's really just a wrapper around Docker, so we're kind of building on, standing on, on shoulders there. But we-- Docker's too low level, so it's just like arbitrary software.

**Alessio** [49:56]
Mm-hmm.

**Ben Firshman** [49:57]
So we ne- we wanted to be able to, like, have a OpenAPI specification there that defined the function effectively that is the machine learning model, but also, like, how that function is written, how that function is run, which is all defined in code and stuff like that.

So it's like a bunch of abstraction on top of Docker to make that work, and that's where that design came from. But the core problems we were solving for users was that, was that Docker's really hard to use and, and productionizing machine learning models is really hard.

**Alessio** [50:31]
Right.

**Ben Firshman** [50:32]
So on the first part of that, uh, we knew we couldn't use Docker files. Like, Docker files are hard enough for software developers-

**Alessio** [50:40]
They are

**Ben Firshman** [50:40]
... to write.

**Alessio** [50:41]
Yeah.

**Ben Firshman** [50:41]
I'm saying this with love as somebody who works on Docker and, like, works on Dock- on, on Docker files. Um, but it's really hard to use, and you need to know a bunch about Linux basically 'cause you're running a bunch of CLI commands.

You need to know a bunch about Linux and best practices and, like, how apt works and all this kind of stuff. So we're like, "Okay, we can't, we can't get to that level. We need to... We need something that machine learning researchers will be able to understand, like people who are used to, like, Colab notebooks."

**Alessio** [51:03]
Mm-hmm.

**Ben Firshman** [51:03]
And what they understand is they're like, "I need this version of Python, I need these Python packages, and somebody told me to apt-get install something." You know?

**Alessio** [51:12]
And throw a sudo in there why not?

**Ben Firshman** [51:13]
And I don't really know.

**Alessio** [51:14]
Right.

**Ben Firshman** [51:14]
And I don't really know what that means. Um, so we tried to create a format that was at that level, and that's what Cog.YAML is. And we're really kind of trying to imagine, like, what is that machine learning researcher gonna understand, you know, and trying to build for them.

And then the productionizing machine learning models thing is like, okay, how can we package up all of the complexity of, like, productionizing machine learning models, like picking CUDA versions-

**Alessio** [51:39]
Mm

**Ben Firshman** [51:39]
... like hooking it up to GPUs, writing an inference server, um, defining a schema, doing batching, um, all of these just, like, really gnarly things that everyone does again and again, and just, like, you know, provide that as a tool.

Uh, and that's where, that's where that side of it came from. So it's like combining those user needs with, you know, the, the sort of world need of needing, like, a-

**Alessio** [52:05]
Mm-hmm

**Ben Firshman** [52:05]
... a common standard for, like, what a machine learning model is, and that's, that's how we thought about the design. I don't know whether that answers the question.

**Alessio** [52:11]
Yeah. So your idea was like, hey, you really want what Docker, uh, stands for in terms of standard, but you actually don't want people to do all the work-

**Ben Firshman** [52:20]
Yeah

**Alessio** [52:20]
... that goes into Docker.

**Ben Firshman** [52:21]
It needs to be higher level, you know?

**Alessio** [52:23]
Mm-hmm.

**Swyx** [52:24]
Um, so I want to, for the listener, um, you're not the only standard that is out there. As with any standard, there must be fourteen of them.

**Ben Firshman** [52:31]
Yeah.

**Swyx** [52:31]
Um, you are very surprisingly friendly with Ollama-

**Ben Firshman** [52:34]
Yeah

**Swyx** [52:34]
... who is your former colleagues from Docker, uh, who came out with the model file. Uh, Mozilla came out with the Llama file.

**Ben Firshman** [52:41]
Yep.

**Swyx** [52:41]
And then, um, I don't know if this is in the same category even, but I'm just gonna throw it in there. Like, Hugging Face has the Transformers and Diffusers library, which is a way of disseminating models that-

**Ben Firshman** [52:49]
Yep

**Swyx** [52:49]
... obviously people use. Um, how would you compare your... contrast your approach of Cog versus all these?

**Ben Firshman** [52:55]
It's kind of complementary actually, which is kind of neat, in that a lot of... Like, Transformers, for example, is lower level than Cog, so it's, you know, a Python library effectively, but you still need to, like-

**Swyx** [53:08]
Expose them.

**Ben Firshman** [53:08]
Yeah. You still need to turn that into an inference server. You still need to, like, install all the Python packages and that kind of thing. So lots of Replicate models are Transformers models, uh, and Diffusers models inside, inside Cog.

You know, so that's, like, the level that that sits. So it's very complementary in some sense, and, you know, we're kind of working on integration with Hugging Face such that you can, like, deploy models from Hugging Face and into Cog models and stuff like that-

**Swyx** [53:31]
Oh

**Ben Firshman** [53:31]
... into Replicate. Um, and, um, so some, some of these things like, uh, Llama File and what Ollama are working on are also very complementary in that they're, they're, they're doing a lot of the sort of running these things locally on laptops, which is not a thing that works very well with Cog.

Like, Cog is really designed around servers and attaching to CUDA devices and, and Nvidia GPUs and this kind of thing. So,

like, we're trying to figure out, we're actually, like, you know, figuring out ways that, like, we can... those things can be interoperable, 'cause, 'cause, you know, they should be. And, um, I think they, they are quite complementary in that you should be able to, like, take a model and replicate it and run it on your local machine.

You should be able to take a model on your local machine and run it in the cloud. Uh, so yeah.

**Swyx** [54:19]
Now, is the base layer something like, um, like is it, is it at the, like, the GGUF level? Which, uh, by the way, I, I need to get a primer on, like, the, the different formats that have emerged.

Uh, or is it at the star.file level, which is model file, Llama file, whatever, whatever? Um, or is it at the Cog level?

**Ben Firshman** [54:37]
I don't know, to be honest.

**Swyx** [54:38]
Yeah.

**Ben Firshman** [54:38]
And I think this is something we still have to, still have to, still have to figure out. Um, I think there's a lot, there's a lot, yeah. Like, exactly where those lines are drawn, don't know exactly. And I think this is something we're trying to figure out ourselves.

But, uh, but I think there's certainly a lot of promise about these systems interoperating. I think we, we just want, we just want things to work together. You know, we wanna try and reduce the number of standards, so the more, the more these things can interoperate and, you know, convert between each other and that kind of stuff at the manner.

**Alessio** [55:01]
A-Andreas comes out of Spotify. Um, Eric from Modo also comes out of S-Spotify. Um, you worked at Docker, and the Ollama guys worked at Docker. Um, where...

**Swyx** [55:13]
Did you know that these ideas were in ... Did both you and Andreas know that there was somebody else you worked with that had a kinda like similar, not similar idea, but like was interested in, in the same thing, or did you then just see, "Oh, I know those people.

They're doing something very similar"?

**Ben Firshman** [55:28]
We learn, we learn about both early on, actually. Yeah. Uh, 'cause we know, we know them both quite well. And it's funny how I think we're all seeing the same problems, and just like applying, you know, trying to fix the same problems that we're all seeing.

**Swyx** [55:40]
Mm-hmm.

**Ben Firshman** [55:41]
I think the Ollama, Ollama one's particularly funny because, um, I joined Docker through my startup. Funnily, actually, the thing which worked for my startup was Compose, but we were actually working on another thing, which was a bit like EC2 for Docker.

**Swyx** [55:56]
Mm.

**Ben Firshman** [55:56]
So we were working on, like, productionizing Docker containers, and Ollama was working on a thing called Co- Kitematic, which was a bit like a, a desktop app for Docker. Um, so ... And our companies both got bought by Docker at the same time.

And, you know, Kitematic turned into Docker Desktop, and then, you know, our thing then turned into Compose. Uh, and it's funny how we're both applying our ... Like, the things we saw at Docker to the AI world.

**Swyx** [56:25]
Yeah.

**Ben Firshman** [56:25]
Where they're building, like, the local environment for us, and we're building, like, the cloud for it. Um, and yeah, so that's just, like, really pleasing, and I think, you know, we're, we're collaborating closely 'cause there's just so much, so much opportunity for working there.

Um-

**Swyx** [56:40]
When you have a hammer, everything's a nail.

**Ben Firshman** [56:42]
Yeah, exactly. Exactly. So I, I think a lot of, a lot of ... This is, I mean, where we're coming from a lot f- with AI is we're taking a lot of things that ... 'Cause we're all kind of, on the Replicate team, we're all kind of people who have built developer tools in the past.

So we've got a team ... Like, I worked at Docker. We've got people who worked at Heroku and GitHub and, like, the iOS ecosystem, and all this kind of thing. Like, the previous generation of, of developer tools, where we, like, figured out a bunch of stuff, and then, like, AI's come along, and we just don't yet have those tools and abstractions, like, to make it easy to use.

So we're trying to, like, take the lessons that we learnt from the previous generation of stuff and apply it to this new generation of stuff. And obviously, there's a bit of nuance there, 'cause the trick is to take, like, the right lessons and do new stuff where it makes sense.

You can't just, like, c- cut and paste, you know?

**Swyx** [57:36]
Mm.

**Ben Firshman** [57:36]
Uh, but that's, like, how we're approaching this, is we're trying to, like, as much as possible, like, take some of those lessons we learned from, like, you know, how Heroku and GitHub was built, for example, and apply them to, apply them to AI.

### Compute & Waste

**Swyx** [57:49]
Excellent. Um, we should, um, also talk a little bit about, um, your compute av- uh, availability. We're trying to ask this of all ... You know, it's Compute Provider Month. Um, do you own your own GPUs? How many, uh, do you have access to?

Do you f- what do you feel about the tightness of the GPU market?

**Ben Firshman** [58:06]
We don't own our own GPUs. We've got a few that we play around with, but not, not for production workloads. And we are primarily built on just public cloud, so primarily GCP and CoreWeave, and, like, some smatterings elsewhere.

And

...

**Swyx** [58:22]
N- none, none from Nvidia, which is your newest investor?

**Ben Firshman** [58:24]
We work with Nvidia, so, you know-

**Swyx** [58:27]
Yeah

**Ben Firshman** [58:27]
... they're, they're, they're kind of helping us get GPU availability. Um, I think GPUs are hard to get hold of if you ... Like, if you go to AWS and ask for one A100, they won't give you an A100.

But if you go to AWS and say, "I would like 100 A100s for two years," they're like, "Sure. We've got some." Um, and I think the problem, the problem is, is the cloud providers, the cloud providers ... Like, that, that makes sense from their point of view.

They want just, like, reliable, sustained usage. They don't want, like, spiky usage and, like, wastage in their infrastructure, which makes total sense. But that makes it really hard for startups, you know, who are wanting to just, like, get hold of GPUs.

I think we're in a fortunate position where we can aggregate demand, so we can make commits to cloud providers. Um, and then, you know, we actually have good availability. Like, it's not, it's not, um, it's not ... You know, we don't have infinite availability, obviously, but, you know, if you want an A100 from Replicate, you can get it.

Um, uh, but, you know, we're seeing other, other companies pop up as well. Like, SF Compute's a great example of this, where they're doing the same idea for training almost, where, you know, a lot of startups need to be able to train a model, but they can't get hold of GPUs from large cloud providers.

So SF Compute are, uh, like, letting people rent, you know, 10 H100s for two days, which is just impossible otherwise.

**Swyx** [59:46]
Yeah.

**Ben Firshman** [59:46]
And, you know, what they're effectively doing there is they're aggregating demand such that they can make a big commit to the cloud provider, and then let people use smaller chunks of it. And that's kinda what we're doing for Replicate as well, where we're make- we're aggregating demand such that we make big commits to the cloud providers, and, you know, then people can get, can run, like, a, a 100 millisecond API request on an A100.

**Swyx** [1:00:05]
Coming from a finance background, this sounds surprisingly similar to banks.

**Ben Firshman** [1:00:09]
Mm.

**Swyx** [1:00:09]
Where the, the job of a bank is, um, maturity transformation is, is, is what you call it. You, you take short-term deposits, which can, which technically can be withdrawn at any time, and you turn that into long-term loans, uh, for mortgages and stuff, and you pocket the difference in interest, and that's, that's the bank.

**Ben Firshman** [1:00:24]
Yep. That's e- that's exactly what we're doing.

**Swyx** [1:00:26]
So you run a bank.

**Ben Firshman** [1:00:27]
Yeah, a GPU bank.

**Swyx** [1:00:28]
Right, yeah. And it, it's, it's supp- so much a finance problem as well, because we have to, we have to make bets on the future demand-

**Ben Firshman** [1:00:35]
We have to do forecasting

**Swyx** [1:00:36]
... for value of GPUs. Yeah. Um-

**Ben Firshman** [1:00:38]
What, what are you ... Okay. I, I, I don't know how much you can disclose, but w- what are you forecasting?

**Swyx** [1:00:43]
Um-

**Ben Firshman** [1:00:44]
Up, down? Up a lot?

**Swyx** [1:00:46]
Yeah. Um-

**Ben Firshman** [1:00:46]
Up 10X? Up-

**Swyx** [1:00:47]
I can't really, can't really ... It, it ... We're projecting our growth with some educated guesses about what kind of models are gonna come out, and what kind of models these will run, you know?

**Ben Firshman** [1:00:54]
Okay.

**Swyx** [1:00:55]
So we, we need to, we need to, we need to bet that, like, okay, maybe language models are getting larger, so we need to, like, have GPUs with a lot of RAM, or, like, multi-GPU nodes, or maybe models are getting smaller and we actually need smaller GPUs.

You know, we have to make some educated guesses about that kind of stuff, yeah.

**Ben Firshman** [1:01:08]
Yeah. Speaking of which, um, the mixture of experts models are, must be throwing a spanner in, into the planning. Um, not so much. I mean, we've got, we've got, we've got s- like multi-node A100 machines which can run this, and multi-node H100 machines which can run this no problem.

So-

**Swyx** [1:01:24]
Okay

**Ben Firshman** [1:01:24]
... we, uh, we, uh, we, we, we, we, uh, we're set up for that, for that, for that world, yeah.

**Swyx** [1:01:30]
Um, okay. Right. I, I didn't, I didn't expect it to be so easy. Um, I mean, the-- my impression was that the amount of RAM per model was increasing a lot, um, s- especially on a sort of per parameter basis.

**Ben Firshman** [1:01:42]
Mm-hmm.

**Swyx** [1:01:42]
Per active parameter basis. Um, i mean, g- going from like Mi- Mixtral being eight experts, um, to like the DeepSeek MOE models, I don't know if you saw them-

**Ben Firshman** [1:01:50]
Mm

**Swyx** [1:01:50]
... being like 30, 60 experts, and you can see it, it keep going up, I guess. I don't know.

**Ben Firshman** [1:01:56]
Yeah. I think we might run into problems at some point. Um, and yeah, I don't know exactly, exactly what's going on there. Um, I think something, something that we're finding which is kind of interesting, like I don't know this in depth, um, but, um, you know, we're certainly seeing a lot of good results from, from, from, uh, lower precision models.

So like, you know, 90% of the performance with just like much less RAM required. Um, and you know, that's, that's-- that means that we can run them on GPUs we have available, and it's good for customers as well because, 'cause it runs faster and like they, they want that trade-off, you know, where, where it's just slightly worse, but like way faster and cheaper.

Yeah.

**Swyx** [1:02:40]
Do you see a lot of, uh, GPU waste in terms of people running the thing on a GPU that is like too advanced? I think we use a T4, uh, to run Whisper, so we- we're at the bottom end of it.

Um, yeah, any thoughts? I think, uh, uh, one of the hackathons we were at, people were like, "Oh, how do I get access to like H100s?" And it's like, you need to run like Stable Diffusion.

**Ben Firshman** [1:03:01]
Dude, you don't need an H100.

**Swyx** [1:03:01]
It's like you don't need an H100.

**Ben Firshman** [1:03:02]
Yeah. Yeah. Well, if you want low latency, you like sure, like spend a lot of money on an H100. Um, uh, yeah, we see a ton of that kind of stuff, and it's surprisingly, it's surprisingly hard to optimize these models right now.

So a lot of people are just running like really unoptimized models. We're doing the same, honestly. Like where a lot of models on Replicate have just been like not been optimized very well. Um, so something we want to like be able to help people with is optimizing those models.

Like either, either we, you know, show people how to with guides, or we make it easier to use some of these more optimized inference servers, or we show people how to compile the models, or we do that automatically, or something like that.

But that's certainly something we're exploring, 'cause yeah, there's, there's so much wastage. Like it's not just wasting the GPUs, it's also like a bad experience and the models run slow, you know?

**Swyx** [1:03:55]
Right.

**Ben Firshman** [1:03:56]
So like a lot of, a lot of the models on Replicate, some of the most popular models on Replicate we have-- So the, the models on Replicate are, are almost all pushed by our community, like people have pushed those models themselves.

But like it's like a big-headed distribution where there's like a long tail of lots of models that people have pushed, and then like a big head of like the models most people run.

**Swyx** [1:04:16]
Mm-hmm.

**Ben Firshman** [1:04:17]
So models like Llama 2, like Stable Diffusion, we, um, we, you know, we work with Meta and Stability to like maintain those models, and we've done a ton of optimization to work those, make those really fast. So, um, yeah, those models are optimized, but the long tail is not, and there's like a lot of, a lot of wastage there.

**Swyx** [1:04:35]
Yeah. And going into the... Well, it's already the new year. Um, do you see the customer demand and the GPU like hardware demand kind of like staying together? Because I think a lot of people are saying, "Oh, there's like hundreds of thousands of GPUs being shipped this year, like the, the crunch is gonna be over."

But you also have like millions of people that now care about using AI. You know, uh, uh, h- how do you see the two lines progressing? Are you seeing customer demand that's gonna outpace the GPU growth? Do you see them together?

Do you see maybe a lot of this like model improvement work kind of helping alleviate that?

**Ben Firshman** [1:05:09]
From our point of view, demand is not outpacing supply of GPUs. Like we have enough, from our point of view, we have enough GPUs to go around, but that might change for sure.

**Swyx** [1:05:18]
Yeah. That's a very, um, nicely put way as a startup founder to respond. Like... Yeah. Uh, I'll maybe-

**Ben Firshman** [1:05:27]
Yeah

**Swyx** [1:05:27]
... get into a little bit of this on the... You s- you said optimizing models. Actually, so like when Alessio said-- talked about GPU waste, he was more... Oh, that you.

**Ben Firshman** [1:05:34]
Sorry.

**Swyx** [1:05:35]
Uh, just-

**Ben Firshman** [1:05:36]
One second. Yeah.

**Swyx** [1:05:36]
Yeah, it is getting a little bit warm in here. Just some greenhouse gas effect. Um, so, so Al- Alessio framed it more as like sort of picking the wrong box model, whereas yours is more about, um, m- maybe the inference stack, if you can call it.

Were you referencing vLLM? Um, w- what, what other sort of techniques are you referencing? And also keeping in mind that when I talk to your competitors, I d- and, and I don't know if, um, we don't have to name any of them, but they are working on trying to optimize the kinds of models.

Like they basically, they'll, they'll quantize their models for you with their special stack. So you, you basically use their versions of Llama 2, you use their versions of Mistral, and that's one way to, to approach it. Uh, I don't see it as the Replicate DNA to do that because that would be like sort of you would have to slap the Rep- Replicate house brand on something, which...

I mean, just comment on any of that. Like what, what do you mean when you say optimize models?

**Ben Firshman** [1:06:28]
Yeah, I mean, you know, things like, I mean, quantizing the models. You can imagine a way that we could help people quantize their models if we want to. Um, we've, um, we've had success using inference servers like vLLM and TRT-LLM, um, and we're using those kind of things to serve language models.

We've had success with things like AI templates which compile, compile the models, um, all of those kind of things. And there's like some even really just boring things of just like, um, making the code more efficient. Like some people, like when they're just writing- ...

some Python code, it's really easy to just write, write inefficient Python code, you know? Um, there's like really boring things like that as well. Um, but it's like a whole smattering of things like that. Um, um, and-

**Swyx** [1:07:15]
So you will do that for a customer? Like you, you'll look at their code and-

**Ben Firshman** [1:07:18]
We-- yeah, we've certainly helped some of our customers be able to do, do that some of the stuff.

**Swyx** [1:07:21]
Wow.

**Ben Firshman** [1:07:21]
That some stuff, yeah. And a lot of the models on, like, the popular models on Replicate, we've, like, rewritten them to use that stuff as well.

**Swyx** [1:07:28]
Okay.

**Ben Firshman** [1:07:29]
Um, and like, like the stable diffusion that we run, for example, is compiled with AI template to make it super fast. And, you know, it's all open source that you can see all of this stuff on GitHub if you wanna, if you wanna like see how we do it.

Um, but you can imagine ways that we could help people, you know, it's almost like built into the Cog layer maybe, where we could help people, like, use these fast inference servers or use AI template to compile their models to make it faster, whether it's like manual, semi-manual, or automatic, we're not really sure.

You know, but that's something we want to explore 'cause, you know, that benefits everyone.

**Swyx** [1:07:59]
Yeah. Awesome. Yeah, and then on the competitive piece, um, there was a price war on Mixtral last year, last year, this last December. Um, as far as I can tell, you guys did not enter that war. Um, you have, you have Mixtral, but you, you know, you-- it's just regular pricing.

Um, I, I think also some of these com- some of these players are probably losing money, um, on, on their pricing. Um, you know, you don't have to say anything, but it's, you know, it's somewhere be- the break-even is somewhere between fifty to seventy-five cents per million tokens, uh, served.

Um, how are you thinking about like the just the overall competitiveness in the market? How, how should people choose when everyone's an API?

**Ben Firshman** [1:08:37]
We actually for our-- So for Llama Two and Mistral, I think not Mixtral, but I can't remember exactly, we have, you know, similar performance and similar price to some of these other ser- other services. We're not like bargain basement like to some of the others 'cause to your point, like we don't wanna like burn tons of money.

Um, but we're, you know, pricing it sensib- sensibly and sustainably, um, to a point where we think it's, we think, you know, it's competitive with other people such that... You know, the thing we don't want-- We-- Like, we, we want developers using Replicate, and we don't wanna, we don't wanna like price it such that it's like only affordable by big companies.

You know, we wanna make it, we wanna make it cheap enough such that the developers can afford it, but we also don't like want the super cheap prices 'cause then like it's almost, it's almost like then your customers are hostile, you know?

**Swyx** [1:09:27]
Mm-hmm.

**Ben Firshman** [1:09:27]
And the, like, the more customers you get, the worse it gets, you know. So we're, we're pricing it sensibly, but still to the point where, you know, uh, where hopefully it's cheap enough to build, build on. Um, and I think the thing we really care about, like we want to, we want to, like obviously we want, you know, models on Replicate to be comparable to other people.

**Swyx** [1:09:47]
Mm-hmm.

**Ben Firshman** [1:09:48]
Um, but I think the really crucial thing about Replicate and the way I think we think about it is that it's not just the API for the-- particularly in open source, it's not just the API for the model that is the important bit.

It's-- Because quite often with open source models, like the whole point of open source is that you can tinker on it, and you can customize it, and you can fine-tune it, and you can like smush it together with another model, like, like Lava, for example.

**Swyx** [1:10:13]
Mm-hmm.

**Ben Firshman** [1:10:13]
Um, and you can't do that if it's just like a hosted API 'cause it's just like it's, you know, it's, it's, you know, you can't, you can't touch the code. Um, so

that's... Like what we wanna do with Replicate is build a platform that's actually open. So like we've got all of these models where the performance and price is on par with everything else. But if you wanna customize it, you can fine-tune it.

You can go to GitHub and get the source code for it and edit the source code and push up your own custom version and this kind of thing. Because that's like the, that's the crucial thing for open source mach- machine learning, is being able to tinker on it and customizing it.

Um, and we think, we think, we think that's really important for, for, for, you know, to make open source AI work.

**Swyx** [1:10:58]
Um, you mentioned open source. How do you think about levels of openness? When Llama Two came out, uh, I wrote a post about this, about it's like open source, and there's open weights, then there's restricted weights. It was on the front page of Hacker News, so there, it, there was like all sort of comments from, from people.

### Open Source & Future

**Swyx** [1:11:14]
So I'm always curious to hear your thoughts. Like what do you think is okay for people to license? What's okay for people to r- not release? Um, yeah.

**Ben Firshman** [1:11:25]
Yeah, I was saying, I mean, you know, before it was just like closed source, big models, open source, little models. You know, purely open source stuff. And we're now seeing like lots of variations where, you know, model companies, uh, putting restrictive licenses on their models.

Um, you know, that means it can only be used for non-commercial use, you know. And a lot of the, you know, open source crowd is complaining it's not true open source, you know, and all this kind of thing.

And yeah, I think a lot of that is coming from philosophy, you know, of like the sort of free software movement kind of philosophy. And I don't think it's necessarily a bad thing. Like it's-- I think it's good that model companies can make money out of their models.

You know, that's like how-- it's what will incentivize people to make more models and this kind of thing. And I think it's totally fine if like somebody made something to ask for some money in return if you're making money out of it, and I think that's, that's totally okay.

And I think there's some really interesting like midpoints as well, where people are releasing the codes. You can still tinker on it-

**Swyx** [1:12:21]
Mm-hmm

**Ben Firshman** [1:12:22]
... but the person who trained the model still wants to get a cut of it if like you're making a bunch of money out of it, and I think that's, that's good, and that's gonna make like the ecosystem more, more sustainable.

And I think we're just gonna see-- I don't think anybody's really figured it out yet, and we're gonna see like more experimentation with this and more people like try to figure out like, "Hmm, what are the business models around building models, and how can I make money out of this?"

And we'll just see where it ends up, and I think it's something we want to support as Replicate as well 'cause we're-- we believe in open source. We think it's great, but there's also gonna be lots of models which are closed source as well, and these companies might not be-- There's probably gonna be a long tail of a bunch of people building models that don't have the reach that OpenAI have, and, you know, hopefully as Replicate, we can help those people find developers and, and help them make money and that kind of thing.

**Swyx** [1:13:13]
Yeah. I, I think the computer requirements of AI kind of change the thing. I, I started an open source company. I'm a big open source fan, and before it was kind of man-hours was really all that went into open source.

It wasn't much monetary investment.

**Alessio** [1:13:27]
Well, not that man-hours are not worth a lot, but if you think about Llama 2, it's like, it's like twenty-five million dollars, you know, like all in. It's like you can't just spin up a Discord and like spend twenty-five million dollars.

So I think it's net positive for everybody that Llama 2 is open source. And, uh, well, is the open source-- You know, is the open source term-- I, I think people, like you're saying, it's like they kind of argue on the semantics of it.

But like all we care about is that Llama 2 is open. Because if Llama 2 wasn't open source today, like the-- if Mistral was not open source, we would be in a bad spot, you know? So-

**Ben Firshman** [1:14:03]
And I think the nuance here is making sure that these models are still tinkerable, because the beautiful thing about Llama 2 as a base model is that, like, yeah, it costs twenty-five million dollars to train to start with, but then you can fine-tune it for like fifty bucks.

**Alessio** [1:14:17]
Right.

**Ben Firshman** [1:14:18]
And that's what's so beautiful about the open source ecosystem, and something I think is really surprising as well. It completely surprised me. Like, I think a lot of people assumed that, um, uh, like it's not gonna be-- open source machine learning is just not gonna be practical because it's so expensive to train these models.

But like fine-tuning is unreasonably effective, and people are getting really good results out of it, and it's really cheap. So people can effectively create open source models, um, really cheaply, and there's gonna be like this sort of ecosystem of tons of models being made.

And I think the risk there from a licensing point of view is we need to make sure that the licenses let people do that.

**Alessio** [1:14:56]
Mm-hmm.

**Ben Firshman** [1:14:56]
Because if you release a big model under a non-commercial license and people can't fine-tune it, you've lost the magic of it being open. And I'm sure there are ways to structure that such that the person paying twenty-five million dollars feels like they're compensated somehow, and they can feel like they can-- you know, they should keep on training models, um, and people can keep on fine-tuning it.

But I guess we just have to figure out exactly how that plays out.

**Swyx** [1:15:20]
Yeah. Excellent. Um, so just wanted to round it out. Uh, you've been, you've been an excellent, a very open guest so far. Um, I actually kind of-- I, I should have started this-- started my intro with this, but I feel like you found the sort of AI engineer crew before I did.

And, uh, you know, something that really resonated with you in sort of, sort of the Series B announcement was that, um, you put in some stats here about how there are two orders of magnitude more software engineers than there are machine learning engineers, about thirty million software engineers and five hundred thousand machine learning engineers.

Um, you can maybe plus or minus one of those orders of magnitude, but it's around that ballpark. And so obviously, there will be a lot more AI engineers than there will be ML engineers. Um, how do you see this group?

Like, is it all software engineers? Are they going to specialize? Um, what would you advise someone trying to become an AI engineer? Is this a legitimate career path?

**Ben Firshman** [1:16:14]
Yeah, absolutely. I mean, it's very clear that AI is gonna be a large part of how we build software in the future now. It's a bit like being a software developer in the nineties and ignoring the internet, you know?

You just need to-- You need to learn about this stuff, and you need to figure this stuff out. I don't think it needs to be-- You don't need to be like super low level. You don't need to be like...

You know, the metaphor here is, is like you don't need to be digging down into like, uh, this sort of PyTorch level if you don't want to. In the same way as a software engineer in the nineties, you don't need to be like understanding how network stacks work to be able to build a website, you know?

But you need to understand the shape of this thing and how to hold it and what it's good at and what it's not, and, and, uh, that's really important. So yeah, certainly just advise people to like just start playing around with it, get a feel of like how language models work, get a feel of like how these diffusion models work, get a feel of like what fine-tuning is and how it works, because some of your job might be building datasets.

You know? Get a feeling of how prompting works, 'cause some of your job might be writing a prompt. And, uh, those are just all really important skills to, skills to, to sort of figure out.

**Swyx** [1:17:29]
Yeah. Well, thanks for building the definitive platform for doing all that.

**Ben Firshman** [1:17:34]
Yeah, of course.

**Alessio** [1:17:35]
Um, any final call to actions? Who should come work at Replicate? Um, yeah, anything for the, for the audience?

**Ben Firshman** [1:17:42]
Yeah. Well, I mean, we're, we're hiring. If you, if you click on jobs at the bottom of our, of replicate.com, there's, there's some jobs. Uh, and I just encourage you to like just like try out AI, even if you don't-- even if you think you're not smart enough.

Like the whole reason I started this company is because I was looking at the cool stuff that Andreas was making. Like Andreas is like a proper machine learning person with a PhD, you know? And I was like, just like a, you know, a sort of lonely software engineer, and I was like, "You're doing really cool stuff, and I wanna be able to do that."

Um, and by us working together, you know, we've now made it, made it accessible to dummies like me. And I just encourage anyone who's like wants to try this stuff out, just give it a try. And I think I would also encourage people who are tool builders.

Like the, the, the limiting factor now on AI is not like the technology. Like the technology's made incredible advances, and there's just so many incredible machine learning models that can do a ton of stuff. The limiting factor is just like making that accessible to people who build products, 'cause it's really hard to use this stuff right now.

And obviously, we're building some of that stuff as Replicate, but there's just like a ton of other tooling and abstractions that need to be built out to make this stuff usable. So I just encourage people who like, like building developer tools to just like get stuck into it as well, 'cause that's gonna make this stuff accessible to everyone.

**Swyx** [1:18:59]
Yeah. I, I especially wanna highlight you have a hacker-in-residence job opening available, which not every company has, which means just join you and hack stuff. I think Charlie Holtz is doing a fantastic job of that.

**Ben Firshman** [1:19:10]
Yep. Effectively, like most of our-- a lot of our job is just like showing people how to use AI.

**Swyx** [1:19:16]
Mm-hmm.

**Ben Firshman** [1:19:16]
So we've just got a team of like software developers and people who've kind of figured this stuff out, who are writing about it, who are, uh, you know, uh, making, making videos about it, who are making example applications just to like show people what you can do with this stuff.

**Swyx** [1:19:29]
Yeah. In, in my world, that used to be called DevRel. But now, now it's hacker-in-residence, and that's, uh...

**Ben Firshman** [1:19:35]
Yeah, this, this came, this came from, um-- Zeke is another one of our-

**Swyx** [1:19:39]
Zeke, yes

**Ben Firshman** [1:19:39]
... of our hackers. Um-

**Swyx** [1:19:41]
Tell me this came from Chroma, 'cause, 'cause I... To start that one.

**Ben Firshman** [1:19:44]
We developed-- Like they, they-- Anton actually was like, "Hey, we came up with that first," but I think we came up with it independently.

**Swyx** [1:19:49]
Oh, yeah, I made that page, yeah.

**Ben Firshman** [1:19:50]
I think we came up with it independently, because the story behind this is, is we originally called it the DevRel team.

**Swyx** [1:19:57]
Yeah.

**Ben Firshman** [1:19:57]
And, um-

**Swyx** [1:19:58]
DevRel's cursed now. Everyone who works with us in the DevRel-

**Ben Firshman** [1:20:01]
Zeke was like, Zeke was like, "That sounds so boring." "I don't wanna go to someone and say I'm a Dev-De-Developer Relations person."

**Swyx** [1:20:05]
I wanna be a hacker man.

**Ben Firshman** [1:20:08]
Or a developer advocate or something. So we were like, "Okay, what's like the way we can make this sound the most fun? All right, you're a hacker."

**Swyx** [1:20:15]
Yeah. Um, I would say like that, that is consistently the vibe I get from Replicate. Every-everyone on, on your team I interact with. When I go to your, your San Francisco office, like that's the vibe that you're generating.

Like it's, it's a hacker space more than an office. Um, and you hold, uh, fantastic meet up, meet ups there, and I think you're a really positive presence in our community. So thank you for doing all that, and it's instilling the hacker vibe and culture into AI.

**Ben Firshman** [1:20:37]
Oh, I'm really glad that, I'm really glad that's working.

**Swyx** [1:20:39]
Yeah.

**Alessio** [1:20:39]
Cool. That's a wrap, I think. Thank you so much for coming on, man.

**Ben Firshman** [1:20:43]
Thank you. Yeah, of course. Thank you. This was a lot of fun.

---

This library is powered by PodHood (https://podhood.com), the podcast website platform.
