LALatent SpaceFeb 28, 2024· 1:21:09

A Brief History of the Open Source AI Hacker - with Ben Firshman of Replicate

Ben Firshman, CEO of Replicate, explains how the inference platform grew from a research reproducibility tool into a 2M-user API business by embracing the generative image community and treating open-source AI as a hacker-friendly ecosystem. They accidentally discovered their API when a user reverse-engineered their web form, leading to their first $1k/month customer. Cog, their container standard for ML models, was born from lessons at Docker and the need to make models tinkerable. Ben argues that fine-tuning's low cost makes open-source models sustainable, and that AI engineers (orders of magnitude more than ML engineers) just need to start playing with models. He also discusses GPU scarcity, preferring sustainable pricing over price wars, and reveals that demand is not outpacing supply thanks to aggregating demand.

  1. 0:00Intro
  2. 6:47Arxiv Vanity
  3. 19:47YC & Pivot
  4. 23:11Community Launch
  5. 31:58API & Growth
  6. 44:30Cog & Standards
  7. 57:49Compute & Waste
  8. 1:10:58Open Source & Future

Powered by PodHood

Transcript

Intro0:00

Alessio0:00

Hey everyone, welcome to the Latent Space Podcast. This is Alessio, partner and CTO in residence at Decibel Partners, and I'm joined by my co-host Swyx, founder of Small AI.

Swyx0:09

Hey, and today we have Ben Firshman in the studio. Welcome, Ben.

Ben Firshman0:13

Hey, good to be here.

Swyx0:15

Uh, Ben, you're a co-founder and CEO, uh, CEO of Replicate. Uh, before that you were most notably cr- uh, creator of Fig, or founder of Fig, which got... Which became Docker Compose. Um, you also did a couple other things, um, before that, but, uh, that's what a lot of people know you for.

Um, what should people know about you that, you know, outside of your, your sort of LinkedIn- ... profile?

Ben Firshman0:36

Yeah, good question. I think I'm a builder and tinkerer, like in a very broad sense, and I love using my hands to make things. So like I work on, you know, things maybe a bit closer to tech, like electronics, but I also like build things out of wood and I like, uh, fix cars, and I fix my bike and build bicycles and all this kind of stuff.

And there's so much I think I've learnt from transferable skills from just like working in the real world to building things, building things, uh, in software. Um, and you know, so much about being a builder both in real life and, and in software that's, uh, that, uh, crosses over.

Swyx1:14

Is there a real world analogy that you use often when you're thinking about like a code ar- architecture or, or problem?

Ben Firshman1:22

I like to build software tools as if they were a, um, as if they were something real. I like to imagine, uh... So I wrote this thing called the Command Line Interface Guidelines, which was a bit like sort of-

Swyx1:38

Yes

Ben Firshman1:39

... the Mac Human Interface Guidelines, but for command line interfaces. I did it with, um, uh, the guy I created Docker Compose with, um, and, and a few other people. And I think something in there, I think I described that your command line interface should feel like a big iron machine, where you pull a lever and it goes clunk.

And like things should respond within like 50 milliseconds, um, as if it was like a real life thing. Um, and like another analogy here is like in the real life, you know when you press a button on an electronic device and it's like a soft switch, and you press it and nothing happens, and there's no physical feedback-

Swyx2:18

Mm-hmm

Ben Firshman2:18

... of anything happening, and then like half a second later something happens? Like that's how a lot of software feels, but instead, like software should feel more like something that's real, where you touch, you pull a physical lever and the physical lever moves, you know?

And I've taken that lesson of kind of human interface, um, to, to software a ton. You know, it's all about kind of low latency, it feeling, things feeling really solid and robust, um, both the command lines and, and uh, and user interfaces as well.

Swyx2:44

And how did you operationalize that for Fig or Docker?

Ben Firshman2:49

Uh, a lot of it's just low latency. Actually, we didn't do it very well for, for Fig and Compose- ... in the first place. We used Python, which was a big mistake, where Python's really hard to get booting up fast, because you have to load up the whole Python runtime before it can run anything.

Swyx3:02

Okay.

Ben Firshman3:03

Um, Go is much better at this, where like Go just instantly starts, um...

Swyx3:08

So, so what's the... Like you have to be under 500 milliseconds to start up?

Ben Firshman3:12

Yeah, effectively.

Swyx3:13

Yeah. Okay.

Ben Firshman3:13

I mean, I mean, you know, human, perception of human things being immediate is, you know, something like 100 milliseconds.

Swyx3:18

Okay.

Ben Firshman3:19

Uh, so any- anything like that is, is, is yeah, good enough.

Swyx3:24

Yeah. Um, um, also I should mention since we're talking about your side projects, uh... Well, one thing is, um, I am maybe one of a few fellow people who have actually written something about CLI design principles, uh, because I was, uh, in charge of the Netlify CLI back in the day, um, and, uh, had many thoughts.

Uh, but one of my fun thoughts I'll just share in case you have thoughts is, um, I think CLIs are effectively starting points for scripts that are, that are then run, and the moment one of the script's preconditions are fa- are not fulfilled, typically they end.

So the, the CLI developer will just, will just exit the-

Ben Firshman3:57

Mm

Swyx3:58

... the program. Um, and the way that I desi- I, I really wanted to, to create the Netlify dev workflow was for it to be kind of a state machine that would resolve the, itself. Um, if it detected a precondition wasn't fulfilled, it would actually delegate to a subprogram that would then fulfill that pre- precondition, asking for more info or waiting until a condition is fulfilled, then it would go back to the original flow and continue, continue that.

Ben Firshman4:20

Mm.

Swyx4:20

Uh, don't know if you like that was ever tried or is there a more formal definition of it, because I just came up with it randomly. But it felt like the beginnings of AI, in the sense that when you run a CLI command, you have an intent to do something, and you may not have given the CLI all the things that it needs to do-

Ben Firshman4:36

Mm-hmm

Swyx4:36

... to execute that intent.

Ben Firshman4:37

Mm-hmm.

Swyx4:38

So that was my two cents.

Ben Firshman4:40

Yeah, that reminds me of a thing we f- we sort of, um, thought about when writing the CLI guidelines, where CLIs were designed in a wor- world where the CLI was really a programming environment, and it was primarily designed for like machines, like to use all of these commands and scripts.

Whereas over time, the CLI has evolved to humans. It was, you know, it's back in a world where like the primary way of like using an... computers was like writing shell scripts effectively. Um, and, uh, we've, we've transitioned to a world where actually humans are using CLI programs much more than they used to, and, um, and the current, the current sort of best practices about how Unix was designed, like, you know, there's, there's lots of sort of design documents about Unix from, from the '70s and '80s, where they, they say things like, "Command line commands should not output s- anything on success.

It should be completely silent." And which makes sense if you're using it in a shell script.

Swyx5:42

Yeah.

Ben Firshman5:42

But if a user is using that, it just looks like it's broken.

Swyx5:45

Yeah.

Ben Firshman5:45

Like if you type copy and it just doesn't say anything-

Swyx5:48

Yeah

Ben Firshman5:48

... you assume that it didn't work as a new user. Um, and yeah, so I, I think what's really interesting about the CLI is for, is that it's actually a really good, to your point, it's a really good user interface Where it can be like a conversation, where it feels like you're, instead of just like you telling the computer to do this thing, and either silently succeeding or saying, "No, you did-- failed," you know?

It can, like, guide you in the right direction and tell you what your intent might be and, and that kind of thing, in a way that's actually -- it's almost more natural to a CLI than it is in a graphical user interface, 'cause it feels like this back and forth with the computer.

Swyx6:31

Yeah.

Ben Firshman6:32

Um, almost funnily like, like a language model. Um, uh, so I think there's some, some interesting intersection of like CLIs and language models actually being very, very sort of, uh, uh, you know, closely related and a good fit for each other.

Swyx6:47

Yeah. I'll, I'll say, uh, one of the surprises from last year, um, you know, I worked on a coding agent, uh, but I think the most successful coding agent of my cohort was Open Interpreter, which was a CLI implementation.

Arxiv Vanity6:47

Swyx6:57

And, uh, I have chronically, even as a CLI person, I have chronically underestimated the CLI as a useful interface.

Ben Firshman7:04

Yeah. Yeah, totally.

Swyx7:05

Um, you also developed Archive Vanity-

Ben Firshman7:06

Yep

Swyx7:06

... which you recently retired after a glorious seven years of-

Ben Firshman7:10

Something like that, yeah

Swyx7:11

... something like that. Um, which is nice, and I guess HTML PDFs.

Ben Firshman7:16

Yep. That, that was actually the, the start of where Replicate came from.

Swyx7:21

Okay.

Ben Firshman7:21

Um-

Swyx7:21

We can tell that story

Ben Firshman7:22

... which, uh... So when I quit Docker, I got really interested in, um, science infrastructure, just as like a problem area, um, because it is... Like, science has created so much progress in the world. The fact that we're, you know, can talk to each other on a podcast and we use computers, and the fact that we're alive is probably thanks to medical research, you know?

But science is just like completely archaic and broken, and it's like 19th century processes that just happen to be, you know, copied to the internet rather than take into account that, you know, we can transfer information at the speed of light now.

Um, and the whole way science is funded, and all this kind of thing is all kind of very broken. Um, there's just so much potential for making science work better, and I realized that I wasn't a scientist, and I didn't really have any time to go and get a PhD and become a researcher, but I'm a tool builder, and I could make existing scientists better at their job.

And if I could make like a bunch of scientists a little bit better at their job, maybe, you know, that's the kind of equivalent of being a researcher. Um, so, um, one particular thing I dialed in on is just how science is disseminated, in that, um, uh, it's all of these, um, PDFs quite often behind paywalls, you know, on the internet.

Um, but-

Swyx8:40

And that's a whole thing.

Ben Firshman8:41

That's-

Swyx8:41

Because it's funded by national grants.

Ben Firshman8:43

Yep.

Swyx8:44

Government grants that are then put behind paywalls.

Ben Firshman8:46

Yeah, exactly. That's, that's like a whole... Yeah, I could talk for hours about that. But the particular thing we got, we got dialed in on was, um, or I, I got kind of, um, but interestingly, these PDFs are also, there's a bunch of open science that happens as well.

So maths, physics, computer science, machine learning notably, is all published on the Arxiv, which is, um, actually a surprisingly old institution.

Swyx9:11

Some random Cornell side project.

Ben Firshman9:12

Yeah, it was just like somebody in Cornell who started a mailing list in the '80s. And then when the web was invented, they built a web interface around it. Like it's super old. Um, and it-

Swyx9:23

It, it's like kind of like a use- user group thing, right? That's why there are all these like numbers and stuff.

Ben Firshman9:26

Yeah, exactly.

Swyx9:27

Uh-

Ben Firshman9:27

Like it's, uh, it's a bit like-

Swyx9:29

The special interest group

Ben Firshman9:29

... or something. Yeah. Um, and, um, that's where all, basically all of maths, physics, and computer science happens. Um, but it's still PDFs published to this thing.

Swyx9:39

Yeah.

Ben Firshman9:39

You know? Which is just so infuriating. Um, so, um, you know, like the, the, the, the web was invented at CERN, a physics institution, to share academic writing.

Swyx9:52

Mm.

Ben Firshman9:52

Like there are these, there are figure tags, there are like author tags, there are heading tags, there are cite tags. You know, hyperlinks are effectively citations-

Swyx10:00

Yes

Ben Firshman10:00

... because you want to link to another academic paper. But instead you have to like copy and paste these things and try and get around paywalls. Like it's absurd, you know? Um, like the, and now, now we have like social media and things, but still like academic papers as PDFs, you know?

It's just like why? This is not what the web was for. Um, so anyway, I got really frustrated with that, and I went on vacation with my old friend Andreas. So we were, we used to work together in London on a startup, at somebody else's startup, and we were just on vacation in Greece for fun, and he was like trying to read a machine learning paper on his phone.

You know, like we had to like zoom in and like scroll line by line on the PDF. And he was like, "This is fucking stupid." So- ... and I was like, "I know." Like this is something, we discovered our mutual hatred for, for this, you know?

And, uh, we spent our vacation sitting by the pool like making LaTeX to HTML like converters and making the first version of Archive Vanity. Um, anyway, that was, uh, been a whole thing, and, um, the story, we shut it down recently because I- they caught, caught the eye of Arxiv who were like, "Oh, this is great.

We just haven't had the time to work on this." And what's tragic about the Arxiv is it is, um, it is like this department of, it's like this project of Cornell that's like they can barely scrounge together enough money to survive.

I think it might be better funded now than it was when we were, we were collaborating with them. Um, and compared to these like scientific journals, it's just like this is actually where the work happens.

Swyx11:23

Yeah.

Ben Firshman11:23

But they just have a fraction of the money that like the, these, these big scientific journals have, which is just so tragic. Um, but anyway, they were like, "Yeah, this is great. We can't afford to like do it, but do you wanna like as a volunteer integrate Archive Vanity into Arxiv?"

Um-

Swyx11:37

Oh, you did the work.

Ben Firshman11:38

We didn't do the work.

Swyx11:39

Oh.

Ben Firshman11:39

We started doing the work.

Swyx11:40

Okay.

Ben Firshman11:40

We did some. I think we worked on this for like a few months to actually get it integrated into Arxiv.

Swyx11:44

Huh.

Ben Firshman11:44

Um, and then we got like distracted by Replicate.

Swyx11:49

Ah.

Ben Firshman11:49

So, um, a guy called Dayan picked up the work and made it happen. Um, like somebody who works on one of the, the piece, the libraries that powers Archive Vanity, um-

Swyx11:58

Okay

Ben Firshman11:59

... but yeah.

Swyx11:59

And the relationship with Archive Sanity?

Ben Firshman12:02

Um, none

Alessio12:03

Did, did you predate them? I, I actually don't know-

Ben Firshman12:05

We-

Alessio12:05

... the, the lineage

Ben Firshman12:06

... we were after-- We both were both users of ArxivSanity-

Alessio12:08

Okay

Ben Firshman12:08

... which is like a sort of archive an- aggregates by-

Alessio12:10

Which is Andreas', uh-

Ben Firshman12:11

I'm sure Andreas Babathi, yeah

Alessio12:12

... like Rexus on top of Arxiv.

Ben Firshman12:13

Yeah, yeah. And we were both users of that.

Alessio12:15

Yeah.

Ben Firshman12:15

And I think we were trying to come up with a working name for Arxiv.

Alessio12:18

Okay. All right.

Ben Firshman12:18

And Andreas just, like, cracked a joke of like, "Oh, let's call it ArxivVanity, 'cause it's making the papers look nice."

Alessio12:23

Yeah, yeah.

Ben Firshman12:23

And that was the working name, and it just stuck.

Alessio12:25

Got it. Got it.

Ben Firshman12:27

Um, yeah.

Alessio12:28

A- and then from there, tell us more about why you got distracted, right?

Ben Firshman12:32

Mm-hmm.

Alessio12:32

So Replicate maybe feels like an overnight success to a lot of people. Um, but you've been building this since 2019. Um-

Ben Firshman12:39

Yeah

Alessio12:39

... so what, what prompted the, the start?

Ben Firshman12:41

And we've been collaborating for even longer. So we created ArxivVanity in 2017. So in some sense, we've been doing this almost, like, six, seven years now. A classic seven-year overnight success.

Alessio12:51

Overnight success.

Ben Firshman12:53

Yeah. Uh, yeah, so we did ArxivVanity, and then worked on a bunch of, like, surrounding projects. I was still, like, really interested in science publishing at that point. Um, and I'm trying to remember, 'cause I tell a lot of, like, the condensed story to people, 'cause I can't really tell, like, a seven-year history, so I'm trying to figure out, like, the right -

Alessio13:09

Oh, we got room

Ben Firshman13:09

... the right, the right length to-

Alessio13:11

We wanna nail the, the definitive Replicate story here.

Ben Firshman13:13

One thing that's really interesting about these machine learning papers is that these machine learning papers are published on A- on the Arxiv, and a lot of them are actual fundamental research, so, like, should be, like, prose describing a theory.

But a lot of them are just running pieces of software that, like, a machine learning researcher made that did something. Um, uh, you know, it was like an image classification model or something, and they managed to make an image classification model that was better than the states of-- the existing state of the art.

And they've made an actual running piece of software that, that, that does image segmentation. And then what they had to do is they then had to take that piece of software and write it up as prose and math in a PDF.

Um, and what's frustrating about that is, like, if, if you wanna... So this was, like, Andreas's-- Andreas was a machine learning engineer at Spotify, and some of his job was, like, he did pure research as well. Like, he did a PhD, and he was doing a lot of stuff internally, but part of his job was also being an engineer and taking some of these existing things that people have made and published and trying to apply them to actual problems at Spotify.

And he was like, you know, you get given a p- a, a paper which, like, describes roughly how the model works. It's probably listing lots of crucial information. There's sometimes code on GitHub. More and more there's code on GitHub, but back in, back in, back then it was kind of relatively rare.

But it was quite often just, like, scrappy research code, and didn't actually run. Um, and you know, there was maybe the weights that were on Google Drive, but they accidentally deleted the weights off Google Drive, you know? And it was, like, really hard to, like, take this stuff and actually use it for real things.

And we just started talking together about, about, like, his problems at Spotify, and I connected this back to my work at, at Docker as well, and was like, "Oh, this is what we created containers for." You know, we solved this problem for normal software by putting the thing inside a container so that you could ship it around and it kept on running.

So we were, we were sort of hypothesizing about, like, "Hmm, what if we put machine learning models inside containers so that they could actually be shipped around, and they could be defined in, like, some production, production-ready format, and other researchers could run them to generate baselines, and you could-- people who wanted to actually apply them to real problems in the world could just pick up the container and run it," you know?

Um, and we then thought, this is probably where the, it gets... Normally, normally in this part of the story, I skip forward to be like, "And then we created Cog, this container standard- ... for, for machine learning models, and we created Replicate, the place for people to publish these machine learning models."

Alessio15:51

Yeah, exactly.

Ben Firshman15:51

But there's actually, like, two or three years between that.

Alessio15:52

Two years in between. Yeah.

Ben Firshman15:54

The, the thing we then got dialed into was Andreas was like, "What if there was a CI system for machine learning?" 'Cause, like, one of the things he really struggled with, with as a researcher is generating baselines.

Alessio16:05

Hmm.

Ben Firshman16:05

So when, like, he's writing a paper, he needs to, like, get, like, five other models that are existing work and get them running.

Alessio16:13

On the same evals.

Ben Firshman16:14

On the sa- exactly, on the same eval, so you can compare apples to apples-

Alessio16:17

Yeah

Ben Firshman16:17

... 'cause you can't trust the numbers in the paper.

Alessio16:19

Yeah.

Ben Firshman16:19

So, um-

Alessio16:20

Or you can be Google and just publish them anyway.

Ben Firshman16:24

Um, so he was like, "What if, what if you could..." I think this was coming from the thinking of, like, there should be containers for machine learning, but why are people gonna use that? Okay, maybe we can create a supply of containers by, like, creating this useful tool for researchers, and the useful tool was like, let's get researchers to package up their models and push them to this central place where we run a standard set of benchmarks across the models so that, um, you can trust those results, and you can compare these models apples to apples.

And for, like, a researcher, for Andreas, like, doing a new piece of research, he could trust those numbers, and he could, like, pull down those, pull down those, those models, con-confirm it on his machine, use the standard benchmark to then measure his model, and, you know, all this kind of stuff.

Um, and so we started building that. That's what we applied to YC with. Um, we got into YC, and we started sort of building a prototype of this. And then this is, like, where it all starts to fall apart.

We were like, "Okay, that sounds great." And we talked to a bunch of researchers, and they really wanted that, and that sounds brilliant. That's a great way to create a supply of, like, models on this research platform. But how the hell is this a business?

You know? Like, how are we even gonna make any money out of this? And we're like, "Oh, shit, that's, like, the-- that's the real unknown here of, like, what the business is." So we, um, we thought it would be a really good idea to like, okay, before we get too deep into this, let's try and, like, um, reduce the risk of this turning into a business.

So let's try and fi- like, research what the business could be for this, for this, uh, for, you know, for this research tool, effectively. So we went and talked to a bunch of companies trying to, trying to sell them something which didn't exist.

So we were like, "Hey, do you want a way to share research inside your company so that other researchers or say, like, the product manager-

Alessio18:09

Mm

Ben Firshman18:09

... can test out the machine learning model?" And they're like, "Uh, maybe." Um, and we were like, "Do you want-" A, like, a deployment platform for deploying models? Like, do you want, like, a central place for versioning models?

Like, we're trying to think of, like, lots of different, like, products we could sell that were, like, related to this thing. Um, and terrible idea. Like, we're not salespeople- ... and, like, people don't wanna buy something that doesn't exist.

Swyx18:34

Mm.

Ben Firshman18:35

Um, I think some people can pull this off, but we were just like, you know, a bunch of product people, product and engineer people, and we just, like, couldn't pull this off. Um, so we then got halfway through our YC batch.

We didn't have-- We hadn't built a product. We had no users. We had no idea what our business was gonna be, 'cause we couldn't get anybody to, like, buy something which didn't exist. Um, and actually, this was quite a way through our...

I think it was, like, two-thirds of the way through our YC batch or something, and we're like, "Okay, well, we're kinda screwed now, um, 'cause we don't have anything to show at demo day." And then we then, like, tried to figure out, okay, what can we build in, like, two weeks that'll be something?

So we, like, desperately tried to... I can't remember what we tried to build at that point. Um, and then two weeks before demo day, I just remember this, um, um... I remember it was all, it was all-- We were going down to Mountain View every week for dinners, and we got called onto, like, an all-hands Zoom call, which was super weird.

We're like, "What's going on?" And they were like, "Don't come to dinner tomorrow." Um, and we realized, we kind of looked at the news and we were like, "Oh, there's a pandemic going on." We were, like, so deep in our startup, we were just, like, completely oblivious to what was going on around us.

Swyx19:44

Was this-

Ben Firshman19:45

Um-

Swyx19:45

... Jan or Feb 2020?

YC & Pivot19:47

Ben Firshman19:47

This was March 2020.

Swyx19:48

March 2020.

Ben Firshman19:49

2020, yeah.

Swyx19:49

'Cause I remember Silicon Valley at the time was early to COVID.

Ben Firshman19:53

Yep.

Swyx19:53

Like-

Ben Firshman19:53

Yeah

Swyx19:54

... they started locking down a lot faster than the rest of US.

Ben Firshman19:55

Yeah, exactly. And I remember, yeah, soon after that, like, there was the San Francisco lockdowns, and then, like, the YC batch just, like, stopped. There wasn't demo day. Um, and it was in a, in a sense a blessing for us, 'cause we just kind of-

... couldn't raise money anyway. Um-

Swyx20:12

In, in the normal course of events, you can-- you're actually allowed to defer-

Ben Firshman20:15

Yeah, exactly

Swyx20:15

... to a future demo day.

Ben Firshman20:16

Yep.

Swyx20:16

Yeah.

Ben Firshman20:17

So we didn't even take any defer 'cause it just-

Swyx20:18

Yeah

Ben Firshman20:18

... kinda didn't happen, you know? So, um, so-

Swyx20:22

So was YC helpful?

Ben Firshman20:24

Yes. We completely screwed up the batch, and that was our fault.

Swyx20:27

Okay.

Ben Firshman20:27

I think the thing that YC has become incredibly valuable for us has been after YC. Um, and I think, I think reason-- You could-- There was a reasonable argument that we sh- couldn't, didn't need to do YC to start with because we were quite experienced.

We had done some startups before. We were kind of well connected with VCs. You know, it was relatively easy to raise money 'cause we were, like, a n-unknown quantity. You know, if you go to a VC and be like, "Hey, I made this piece of, piece of-"

Swyx20:53

It's Docker Compose for AI. AI.

Ben Firshman20:55

Exa-exactly, yeah. And, and, and like, you know, people can pattern match like that, and they can sort of have some trust you know what you're doing. Um, whereas it's much harder for people straight out of college, and that's where, like, YC's sweet spot is, like, helping people straight out of college who are super promising, like, figure out how to do that.

Swyx21:08

Yeah, no credentials.

Ben Firshman21:09

Yeah, exactly.

Swyx21:10

Yeah.

Ben Firshman21:10

So in some sense, we didn't need that, but the thing that's been incredibly useful for us since YC has been... This was actually, I think-- So Docker was, Docker was a YC company, and Solomon, the founder of Docker, I think, told me this.

He was like, um, "A lot of people underestimate the value of YC after you finish the batch." And he, his biggest regret was, like, not staying in touch with YC. I might be misattributing this, but I think it was him.

And so we made a point of that, and we just stayed in touch with our batch partner, who, um, Jared at YC, who's been fantastic.

Swyx21:42

Jared Harris?

Ben Firshman21:43

Um, Jared Friedman.

Swyx21:44

Friedman.

Ben Firshman21:45

And all of, like, the team at YC, like, there was the growth team at YC when they, when they were still there, and they've been super helpful. Um, and, um, two things have been super helpful about that is, like, raising money.

Like, they just know exactly how to raise money, and they've been super helpful during that process in all of our rounds. Like, we've done three rounds since we did YC, and they've been super helpful during the whole process.

Um, and also just, like, reaching a ton of customers. So, like, the magic of YC is that you have all of-- Like, there's thousands of YC companies, I think. Like, on the-

Swyx22:14

Thousands

Ben Firshman22:15

... order of thousands, I think.

Swyx22:15

Yeah, yeah.

Ben Firshman22:15

Um, and they're all of your first customers. And they're, like, super helpful, super receptive, really want to, like, try out new things. Um, you have, like, a warm intro to every, every one of them, basically, and there's this mailing list where you can post about updates to your, um, to your product, um, which is, like, really receptive, and that's just been fantastic for us.

Like, we've, we've just, like, got so many of our, of our users and customers through, through YC. Um-

Swyx22:42

Yeah, well, so the classic criticism or the sort of, you know, pushback is people don't buy you because, um, you are both from YC, but at least they'll open the email.

Ben Firshman22:53

Yeah.

Swyx22:53

Right? Like, that's the-

Ben Firshman22:54

Yeah.

Swyx22:54

Okay.

Ben Firshman22:55

Yeah, effectively. Um, and

yeah. Yeah, so that's been a really, really positive experience for us.

Swyx23:02

Mm-hmm. And, and sorry, I interrupted, uh, with the YC question. Like, you were-- You were making-- You just made it out of the YC-

Ben Firshman23:07

Oh, yeah

Swyx23:07

... survived the pandemic. Um, and you, yeah.

Ben Firshman23:11

I'll try and condense this a little bit. Then we, then we started building tools for COVID, weirdly. We were like, "Okay, we don't have a startup. We haven't figured out anything. What's the, what's the most useful thing we could be doing right now?"

Community Launch23:11

Swyx23:21

Save lives.

Ben Firshman23:22

So yeah, let's try and, let's try and save lives. I think we failed at that as well. We had a bunch of projects that didn't really go anywhere. Um, uh, we kind of worked on, yeah, a bunch of stuff like contact tracing, which turned out didn't really to be a useful thing.

Um, sort of, uh, Andreas worked on a, like, a, um, like a, a DoorDash for, like, people delivering food to people who are vulnerable. Uh, what else did we do? The meta problem of, like, helping people direct their efforts to what was most useful.

Um, and a few other things like that. Didn't really go anywhere. So we're like, "Okay, this is not really working either." Um, we, we were considering actually just, like, doing, like, work for COVID. We had, like, this decision document early on in our company, which is like, should we become a, like, government app contracting shop, you know?

Swyx24:06

Mm.

Ben Firshman24:07

Um, we decided no.

Swyx24:08

Because you also did, uh, work for the US, uh, for the gov.uk.

Ben Firshman24:11

Yeah, exactly. We had experience, like, doing some, like, uh-

Swyx24:15

And The Guardian and, you know, that

Ben Firshman24:16

... yeah, for, like, government stuff. Um, and we were just, like, really good at building stuff. Like, we were just, like, product people. Like, I was, like, the front-end product side, and Andreas was the back-end side. So we were just, like, a, a product-- And we were working with a designer at the time, um, a guy called Mark, who did our early designs for Replicate.

And we were like, "Hey, what if we just team up and, like, become and, and build stuff?" And But yeah, we gave up on that in the end for-- I can't remember the details. Um, so we c- went back to machine learning, and then we were like, "Oh, well, we're not really sure if this is gonna work."

And one of my most painful experiences from previous startups is shutting them down. Like, when you realize it's not really working and having to shut it down, it's like a ton of work, and it's-- people hate you, and it's just sort of, you know...

Um, so we were like, "How can we make something we don't have to shut down? And even better, how can we make something that won't page us in the middle of the night?"

Uh, so we made an open source project. We made a thing which was an open source weights and biases, um, 'cause we had this theory that, like, weights and-- that, that, like, people want open source tools. There should be, like, an open source, like, version control experiment tracking, like, thing.

And it was intuitive to us in that we're like, "Oh, we're software developers, and we like command line tools." Like, everyone likes command line tools and open source stuff. But machine learning researchers just really didn't care. Mm-hmm. Like, they just wanted to click on buttons.

They didn't mind that it was a cloud service. Like, it was all very visual as well, that you needed lots of graphs and, and charts and stuff like this. So it just didn't-- it wasn't right. Like, it was right-- We were actually rebuilding something that Andreas made at Spotify for just, like, saving experiments to cloud storage automatically, but other people didn't really want this.

So we kinda gave up on that, and then we-- That was actually originally called Replicate, and we renamed that out the way, so it's now called Keepsake, and I think some people still use it. Then we sort of came back-- We looped back to our original idea.

So we were like, "Oh, maybe there was a thing in that thing we were originally sort of thinking about of, like, researchers sharing their work and containers for machine learning models." So we just built that, and at that point, we were kind of running out of the YC money, so we were like, "Okay, this, like, feels good though.

Let's, like, give this a shot." So that was the point we raised a seed round. We raised, um, seed round- Pre-launch. We raised pre-launch. Pre-launch and pre-team. Um, it was an idea, basically. We had a little prototype. It was just an idea and a team.

Um, but we were like, "Okay," like, you know, when-- "bootstrapping this thing is getting hard, so let's actually raise some money." Um, and then we made Cog and Replicates. Mm. It initially didn't have APIs, interestingly. It was just the bit that I was talking about before of helping researchers share their work.

So it was a way for researchers to, to put their work on a webpage such that other people could try it out, uh, and so that you could download the Docker container. So that, like, we didn't have-- We cut the benchmarks thing of it 'cause we thought that was just, like, too complicated.

But it had a Docker container that, like, you know, Andreas in a past life could download and run with his benchmark, and you could compare all these models apples to apples. So that was, like, the theory behind it.

Um, and that kind of started to work. It was, like, still when, like, you know, it was pre-- long time pre-AI hype, and there was lots of interesting stuff going on, but it was, it was very much in, like, the classic deep learning era.

So sort of image segmentation models and sentiment analysis and all these kind of things, you know, that people were using, uh, that were using deep learning models for. And we were very much building for research 'cause all of this stuff was happening in research institutions.

You know, there's some people who'd be publishing to Arxiv. So we were, we were creating an accompanying material for their models, basically. You know, they wanted a demo for their models, and we were creating accompanying material for it.

Um, and they were like-- What was funny about that is they were, like, not very good users. Like, they were, they were doing great work obviously, but, but the way that research worked is that they, they just made, like, one thing every six months, and they just fired and forget it, forgot it.

Yeah. Like, they, they published this piece of paper, and, like, done. I'm, I've, I've published it. Um, so they, like, output it to Replicate, and then they just stopped using Replicate. Yeah. You know? They were, like, once every six monthly users.

And that wasn't great for us. Um, but we stumbled across this early community. This was early 2021 when

people started-- OpenAI created this-- created Clip, and people started smushing Clip and GANs together to produce image generation models. And this started with, um, you know, it was just a bunch of, like, tinkerers on Discord, basically. Um, it was, um-- There was an early model called Big Sleep by AdvadNaun, and then there was VQGAN-CLIP, which was, like, a bit more popular, by RiversHaveWings.

And it was all just people, like, tinkering on stuff in Colabs- Right ... and it was very dynamic, and it was people just making copies of Colabs and playing around with things and forking and... And to me, this-- I saw this and I was like, "Oh, this feels like open source software."

Like, so much more than the research world- Mm-hmm ... where, like, people are publishing these papers. Yeah, you don't know their real names, and then it's just, like, a Discord ID. Yeah, exactly. But crucially, it was like people were tinkering and forking, and people were- Yeah ...

things were moving really fast, and- Yeah ... um, it just felt like this creative, dynamic, collaborative community in a way that research wasn't really. Like, it was still stuck in this kind of six-month publication cycle. So we just kinda latched onto that and started building for this, this community.

Um, and you know, a lot of those early models were published on Replicates. No-- I think the first one that was really primarily on Replicates was one called Pixray, which was sort of, sort of mid-2021. Um, and it had a really cool, like, pixel art output, but it also just, like, produced-- They weren't, like, crisp in images, but they were quite aesthetically pleasing- Oh ...

like some of these early image generation models. And, um, um, you know, that was, like, published primarily on Replicates, and then a few other models around that were, like, published on Replicates. And that's where we really started to find our early community- Okay ...

and, like, where we really found, like, oh, we've actually built a thing that people want. Um, and they were great users as well, and people really wanna try out these models. Lots of people were, like, running the models on Replicate.

We still didn't have APIs though. Interestingly- ... and this is like another like really complicated part of the story. We had no idea what our business model was still at this point. I don't think you- people could even pay for it.

You know, it's just like these web forms where people could run the model. Um, and-

Swyx30:47

Just before this API bit, uh, continue. Uh, just for-

Ben Firshman30:49

Yeah

Swyx30:49

... historical interests, uh, which Discords were they, and how did you find them? Was this the Lion Discord?

Ben Firshman30:54

Yeah, Lion-

Swyx30:54

Was this Luther?

Ben Firshman30:55

Luther, yeah. It was the Luther one-

Swyx30:56

These two, right?

Ben Firshman30:57

Luther I particularly remember. There was a channel where, where VQGAN-CLIP-- This was early 2021, where VQGAN-CLIP was set up as a, as a Discord bot, and I just remember being completely just like captivated by this thing. I was just like playing around with it all afternoon, and like the sort of thing where-

Swyx31:15

In Discord

Ben Firshman31:15

... you're like, "Oh, shit, it's 2:00 AM," you know .

Swyx31:16

Yeah. This is the beginnings of Midjourney.

Ben Firshman31:18

Yeah, exactly. And it was-

Swyx31:19

And, and stability

Ben Firshman31:20

... it was the start of-- It was the start of Midjourney, and, you know, it's where that kind of user interface came from. Like what's beautiful how the user interface is like you could see what other people are doing.

Swyx31:29

Yeah.

Ben Firshman31:29

And that you could, you could riff off other people's ideas, and it was just so much fun to just like play around with this in like a channel full of 100 people. Uh, and yeah, that just like completely captivated me, and I'm like, "Okay, this, this is like s- this is something," you know?

So like we should get these things on Replicate. Um, and yeah, that's, that's where that, that all came from. Yeah.

Swyx31:49

Okay. Sorry, uh, and I just wanted to capture that moment.

Ben Firshman31:52

Yeah, yeah.

Swyx31:53

Um, um, and then you moved on to-- So was it APIs next, or was it Stable Diffusion next?

Ben Firshman31:57

It was APIs next, and the APIs happened because one of our users-- Our web form had like an internal API for making the web form work, like with, uh, an API that was called from JavaScript. And somebody like reverse engineered that to start generating images with a script.

API & Growth31:58

Ben Firshman32:15

You know, they did like-

Swyx32:16

Mm-hmm

Ben Firshman32:16

... you know, web inspector copy-

Swyx32:18

Some kind of copyright thing

Ben Firshman32:18

... as Carl, like figured out what the-

Swyx32:19

Oh, I see. I see

Ben Firshman32:20

... API request was.

Swyx32:21

Got it. Yep.

Ben Firshman32:22

Um, and it wasn't secured or anything. Um-

Swyx32:25

Of course not

Ben Firshman32:25

... and they started generating a bunch of images, and like we got tons, tons of traffic, and we're like, "What's going on?" Um, and I think like a s- a sort of usual reaction to that would be like, "Hey, you're abusing our API," and to shut them down.

Instead, we were like, "Oh, this is interesting. Like, people wanna run these models." Um, so we documented the API in a Notion document, like our internal API in a Notion document, and like messaged this person being like, "Hey, you seem to have found our API."

"Um, here's the documentation. That'll be like 1,000 bucks a month, please" With a stripe form, like that we just click some buttons to make. Um, and they were like, "Sure, that sounds great." So that was our first customer .

Um-

Swyx33:09

1,000 bucks a month?

Ben Firshman33:10

Uh, it was, it was a surprising amount of money, yeah.

Swyx33:12

That's not-

Ben Firshman33:12

It was on the order-

Swyx33:13

... casual

Ben Firshman33:13

... it was on the order of 1,000 bucks a month.

Swyx33:15

So was he-- Was it a business? Like what-

Ben Firshman33:17

It was the creator of PixRay. Like it was-

Swyx33:20

Oh

Ben Firshman33:21

... he generated NFT art, and so he like made a bunch of art with these models-

Swyx33:27

Mm

Ben Firshman33:27

... and, um, was, was, you know, selling these NFTs effectively. And I think p- lots of people in his community were doing similar things, and like he then referred us to other people who were also generating, uh, generating NFTs using generative models, and that was like the start of like, uh-- That was the start of, start of our API business, yeah.

And then we, then we like made an official API and actually like added some, some billing to it. Uh, so it wasn't just like a fixed fee and yeah.

Swyx33:53

And now people think of you as the hosted models API business.

Ben Firshman33:56

Yep, exactly. And, and, and-- But that just turned out to be our business. You know-

Swyx33:59

Yeah

Ben Firshman34:00

... but what, but what, what ended up being beautiful about this is it was really fulfilling like the original goal of what we wanted to do, is that we wanted to make this research that people were making accessible to like other people, and for it to be used in the real world.

And this was like the-- just like ultimately the right way to do it because all of these people making these generative models could publish them to Replicate, and they wanted a place to publish it. And software engineers, you know, like myself, like I'm not a machine learning expert, but I wanna use this stuff, uh, could just run these models with a single line of code.

And we thought, "Ah, maybe the Docker image is enough," but it's actually super hard to get the Docker image running on a GPU and stuff. So it really needed to be the hosted API for this to work and to make it accessible to software engineers, and we just like w- wound our way to this, this-

Swyx34:46

Yeah, two years to the first paying customer

Ben Firshman34:48

... this solution. Yeah, exactly. Um-

Swyx34:50

D- did you ever think about becoming Midjourney during that time? You have like-

Ben Firshman34:54

Mm

Swyx34:54

... so much interest in image generation-

Ben Firshman34:55

What could have been-

Swyx34:56

... it's

Ben Firshman34:57

Yeah.

Swyx34:57

I mean, you're doing fine d- for the record, but you know. It was right there. You were playing with it.

Ben Firshman35:04

Yeah, I don't, I don't think it was our expertise.

Swyx35:07

Okay.

Ben Firshman35:07

Like I think our expertise was dev tools rather than-- Like Midjourney's almost like a consumer product, you know?

Swyx35:11

It is, yeah.

Ben Firshman35:12

Um, so I don't think it was our expertise. Uh, it certainly occurred to us. Um, I think at the time we were thinking about like, "Oh, maybe we could hire some of these commun- people in this community and make great models," and stuff like this, but just ended up our-

Swyx35:24

Mm-hmm

Ben Firshman35:24

... we, we ended up more being at the tooling. Like I think-

Swyx35:27

Yeah

Ben Firshman35:27

... like before I was saying like I'm not really a researcher, but I'm more like the tool builder-

Swyx35:30

Mm-hmm

Ben Firshman35:30

... like behind the scenes, and I think both me and Andreas are like that. Yeah.

Swyx35:33

Yeah.

Alessio35:33

Yeah.

Swyx35:33

I, I think this is a-- also like a illustration of the tool builder philosophy, something where you, you're very-- you, you latch onto in dev tools, which is when you see people behaving weird, it's not your f- it's not their fault, it's yours.

Like you, you-- And, and you wanna pave the cow paths is what they say, right? Like the, the unofficial paths that people are making, like make it official and make it easy for them, and then maybe charge a bit of money.

Ben Firshman35:52

Mm-hmm.

Alessio35:53

Yep.

Ben Firshman35:53

Yeah.

Alessio35:54

Um, and now fast-forward a couple of years, you have two million developers using Replicate. Maybe more. That, that was the last public number that I found.

Ben Firshman36:02

Two million. I think that got mangled actually by-- It's two million users. Not all those people are developers, but a lot of them are developers, yeah.

Alessio36:09

Um, and then 30,000 paying customers was the number. Um, that's, that's awesome. Uh, Latent Space runs on Replicate.

Ben Firshman36:17

Right. Nice.

Alessio36:17

So we have a small podcaster, and we host, uh, whisper-

Swyx36:19

We do a transcription on Replicate

Alessio36:20

... whisper diarization on, on Replicate.

Swyx36:23

Cool.

Alessio36:23

Um, so-- And we're paying, so we're-- Latent Space- ... is in the 30,000.

Swyx36:27

Nice. Thank you.

Alessio36:27

Um, you raised a $40 million Series B. Um, I would say that maybe the Stable Diffusion time, August '22, was like really when the company started to-

Ben Firshman36:38

Yeah

Alessio36:38

... to break out. Um, tell us a bit about that and the community that came out, and I know now you're expanding beyond just, uh, image generation.

Ben Firshman36:46

Yeah. This-- Like, I think we kind of set ourselves-- Like, we saw there was this really interesting ge-image, generative image world going on, so we kind of, you know, like, we're, we're building the tools for that community already, really.

And we knew Stable Diffusion was coming out. We knew it was a really exciting thing. You know, it was the best, the best generative image model so far. I think the thing we didn't-- we underestimated was just, like, what an inflection point it would be, where it was an-- It was-- I think, I think Simon Willison put it this way, where he said something along the lines of, it was a model that was open source and tinkerable, and, like, good enough that it was just, like...

It was, it was, you know, it was just good enough and open source and tinkerable, such that it just kind of took off-

Swyx37:35

Mm-hmm

Ben Firshman37:35

... in a way that none of the models had before. And, like, what was really neat about Stable Diffusion is, is it was open source, so you could, like-- Compared to, like, DALL-E, for example, which was, like, sort of equivalent qua-quality, you-- it was open source, so you could fork it and tinker on it.

And, like, the first week, we saw, like, people making animation models out of it. We saw people make, like, game texture models that, like, use circular convolutions to make repeatable textures. We saw-- What else did we see? Um, you know, a few weeks later, like, people were fine-tuning it, so you could make-- put your face in these models and, um, all of these other-

Swyx38:10

Yeah, textual inversion.

Ben Firshman38:11

Yep. Yeah, exactly. That happened a bit before that. And all of this sort of innovation was happening all of a sudden, and people were publishing it on Replicate because you could just, like, publish arbitrary models on Replicate, so we had this sort of supply of, like, interesting stuff being built.

But because it was a sufficiently good model, um, there was also just, like, a ton of people building with it. They were like, "Oh, we can build products with this thing." And this was, like, about the time where people were starting to get really interested in AI, so, like, a ton of the product builders wanted to build stuff with it.

And we were just, like, sitting in there in the middle as, like, the interface layer between, like, all these people who wanted to build and all these, like, machine learning experts who were building cool models. Um, and that's, like, really where it took off.

We were just sort of credible supply and credible demand, and we were just, like, in the middle. Um, and then, yeah, since then, we've just kind of grown and grown, really. And we, we, you know, have been building a lot for, like, the indie hacker community, these, like, individual tinkerers, but also startups and a lot of large companies as well who are sort of exploring and building AI things.

And then kind of the same thing happened, like, middle of last year with language models and Llama 2, where the same kind of Stable Diffusion effect happened with, with Llama. And Llama 2 was, like, our biggest week of growth ever because, like, tons of people wanted to tinker with it and run it.

And, you know, since then, we've just been seeing a ton of growth in language models as well as image models. And, uh, yeah, we're just kind of riding a lot of the, the interest that's going on in AI and all the people building in AI, you know?

Swyx39:37

That's, uh-- Yeah, kudos. Right place, right time, but also, you know, took a while to position for the, for the right, uh, place before the wave came. Um, I, I'm, I'm curious if, like, um, you have any insights on these different markets.

Um, so Peter Levels, notably a very loud person, uh, very picky about his tools. Um, I wasn't sure actually if he used you. He does.

Ben Firshman40:00

He does, yeah.

Swyx40:00

Because you cited him, you cited him on your series V blog post, and Danny Postma as well, his competitor-

Ben Firshman40:04

Yeah

Swyx40:05

... um, all, all in that wave. Um, what are their needs versus, um, you know, the more enterprise or B2B type needs? Did, did you, did you come to a decision point where you're like, "Okay, you know, how serious are these indie hackers versus, like, the actual businesses that are bigger and perhaps better customers because they're less churny?"

Ben Firshman40:24

They're surprisingly similar.

Swyx40:25

Okay.

Ben Firshman40:26

Because I think a lot of people right now want to use and build with AI, but they're not AI experts, and they're not infrastructure experts either. So they wanna be able to use this stuff without having to, like, figure out all the internals of the models and, you know, like, touch PyTorch and whatever.

And they also don't wanna be, like, setting up and booting up servers. Um, and that's the same all the way from, like, indie hackers just getting started, because, like, obviously you just wanna get started as quickly as possible, all the way through to, like, large companies who wanna be able to use this stuff but don't have, like, all of the experts on staff, you know?

Um, like, I think some, some companies are quite, you know, there are companies, big companies like Google and so on that do actually have a lot of experts on staff, but the vast majority of companies don't. And they're all software engineers who wanna be able to use this AI stuff, but they just don't know how to use it.

And it's like, you really need to be an expert, and it takes a long time to, like, learn the skills to be able to use that. So they're surprisingly similar in that sense. Um, and I think, I think it's kind of also unfair of, like, the indie community.

Like, surpris- They're not churning, surprisingly, or churny or spiky, surprisingly. Like, they're building real established businesses, which is like kudos to them, like, of, like, building these, like, really, like, large, sustainable businesses, often just as, like, solo developers.

Uh, and it's kind of remarkable how they can do that, actually, and it's a credit to a lot of their, like, their product skills and, you know, we're just, like, there to help them, being like their machine learning team, effectively, uh, to help them use all of this stuff.

Um, so we're actually making some-- Like, like, a lot of these indie hackers are some of our largest customers, like alongside- ... some of our biggest customers that you would think would be-

Swyx42:08

Yeah

Ben Firshman42:08

... would be, would be, uh, would be, uh, you know, spending a lot more money than them, but yeah.

Swyx42:14

Uh, and we should name some of these. You have them on your landing page. You have BuzzFeed, you have Unsplash, uh, Character AI. Um, how-- Like, what do they power? What, what can you say about their, their usage?

Ben Firshman42:24

Yeah, totally. It's, it's kind of, uh, v-various things. I'm trying to think. Um,

let me actually think. What can I say about what customers?

Swyx42:36

Well, I, I mean, I'm naming them because they're on your landing page-

Ben Firshman42:38

Yeah

Swyx42:38

... so, so you have logo rights.

Ben Firshman42:39

Yeah.

Swyx42:40

Um, it's u- It's useful for people to who are-- Like, I, I'm not imaginative. I, I see... Monkey see, monkey do, right? Like if, if I see-

Ben Firshman42:46

Yeah, yeah

Swyx42:46

... someone doing something that I wanna do, then I'm like, "Okay, Replicate's great for that."

Ben Firshman42:49

Yeah, yeah, yeah.

Swyx42:50

So s- that's what I think about case studies on company landing pages is that it's just a way of explaining, like, "Yep, we, we-- this is something that we are good for."

Ben Firshman42:58

Yeah, totally. It-- I mean, it's- These companies are doing things all the way up and down the stack at different levels of sophistication. So like Unsplash, for example, they, they actually, they actually publicly posted this story on Twitter where they're using, uh, BLIP to annotate all of the images in their catalog.

So, you know, they have lots of images in the catalog, and they wanna create a text des- description of it so you can search for it. Um, and they're annotating the images with, you know, off-the-shelf open source model.

You know, we have this big library of open source models that you can run, and, you know, we've got lots of people who are running these open source models off the shelf. And then, you know, most of our larger customers are doing more sophisticated stuff, so they're like fine-tuning the models, they're running completely custom models on us.

And so a lot of these, a lot of these larger companies are, like, using us for a lot of their, their, you know, inference, but it's like a lot of custom models and them, like, writing the Python themselves 'cause they've got machine learning experts on team, on the team, and they're using us for like, you know, their inference infrastructure effectively.

Um, so it's like lots of different levels of sophistication where like some people are using these off-the-shelf models, some people are fine-tuning models. So like level, Peter Level is a great example where a lot of his products are based off like fine-tuning, fine-tuning image models, for example.

And then we've also got like larger customers who are just like using us as infrastructure effectively as, as servers. Um, so yeah, it's like all things up and down, up and down the stack.

Alessio44:29

Yeah. Um, let's talk a bit about Cog and the, the technical layer. So there are a lot of, uh, GPU clouds, uh, I think people at different pricing points, and I think everybody tries to offer a different developer experience on top of it, which then lets you charge a premium.

Cog & Standards44:30

Alessio44:46

Why did you wanna create Cog? What were some of the... You worked at Docker. What were some of the issues with traditional container runtimes? Um, and maybe, yeah, what, what were you surprised with as you built it?

Ben Firshman44:57

Cog came right from the start actually, when we were thinking about this, this, you know, evaluation, this sort of benchmarking system for machine learning researchers, where we wanted researchers to publish their models in a standard format that was guaranteed to keep on running, that you could replicate the results of, like that's where the name came from.

Alessio45:19

Mm-hmm.

Ben Firshman45:20

And we realized that we needed something like Docker to make that work, you know. Um, and I think it was just like natural from my point of view of like, obviously, that should be open source, that we should try and like create some kind of open standard here that people can share, because if more people use this format, then that's great for everyone involved.

Um, you know, I think, I think the magic of Docker is not really in the software.

Alessio45:42

Mm-hmm.

Ben Firshman45:42

It's just like the standard that people have agreed on, like, here are a bunch of keys for a JSON document-

Alessio45:48

Right. Yeah.

Ben Firshman45:48

... basically. And, um, you know, that was the magic of like the metaphor of real containerization as well. It's not the containers that are interesting. It's just like the size and shape of the damn box, you know?

Alessio45:58

Mm-hmm. Right. Yeah.

Ben Firshman45:59

Um, and it's similar thing here, where really we just wanted to get people to agree on like, this is what a machine learning model is. This is, this is how a prediction works. This is what the inputs are.

This is what the outputs are. So Cog is really just a Docker container that attaches to a CUDA device if it needs a GPU, that has a OpenAPI specification as a label on the Docker image.

Alessio46:22

Right.

Ben Firshman46:22

And the OpenAPI specification defines the interface for the machine learning model, like the, the, um, the inputs and outputs effectively, or the, the params in machine learning terminology. Um, and you know, we just tried, wanted to get people to kind of agree on this thing, and it's like general purpose enough.

Like we weren't saying like some of the existing things were like at the graph level.

Alessio46:45

Mm-hmm.

Ben Firshman46:45

But we really wanted something general purpose enough that you could just put anything inside this, and it was like future compatible, and it was just like arbitrary software, and, you know, be future compatible with like future inference servers and future machine learning model formats and all this kind of stuff.

Alessio46:57

Yeah.

Ben Firshman46:57

Um, so that was the intent behind it. And, you know, for... It just came naturally that we wanted to define this format and, and that's been really working for us. Like a bunch of people have been using Cog outside of Replicates, which is kind of our original intention.

Alessio47:13

Oh, wonderful.

Ben Firshman47:13

Like this should be how machine learning models are packaged and how people should use it. Like it's common to use Cog in situations where like maybe they can't use the SaaS service because-

Alessio47:23

Uh-huh

Ben Firshman47:23

... I don't know, they're in a big company and they're not allowed to like, you know, not allowed to use a SaaS, SaaS service, but they can use Cog internally still, and like they can download the models from Replicates and run them internally in, in their org, which we've been seeing happen.

That works really well. Um, people who wanna build like custom inference pipelines but don't wanna like reinvent the world, so they can use Cog off the shelf and use it as like a component in their inference pipelines. Um, we've been seeing tons of, tons of usage like that.

Um, and it's just been kind of happening organically. We haven't really been trying, you know, but it's like there if people want it, and we've been seeing people use it, so that's great. Um-

Alessio47:57

Yeah

Ben Firshman47:57

... and, uh, yeah, so a lot of it's just sort of philosophical of just like, this is, this is how it should work from my experience at Docker, you know?

Alessio48:03

Yeah.

Ben Firshman48:03

And there's just a lot of value from like the core being open, I think, and that other people can share it, and it's like an integration point. So, you know, if, if Replicate, for example, wanted to, wanted to work with a testing system, like a CI system or whatever, um, y- we can just like interface at the Cog level.

Like-

Alessio48:20

Mm-hmm

Ben Firshman48:21

... that, that system just needs to pull Cog models, and then you can like test your models on that CI system before they get deployed to Replicate, and it's just like a format that everyone-- we can get everyone to agree on.

Yeah.

Alessio48:30

W- what do you think, I guess, Docker got wrong? Because if I look at a Docker Compose and a Cog definition, first of all, the, the Cog is kinda like the Docker file plus the Compose-

Ben Firshman48:40

Yeah

Alessio48:40

... versus, and Docker Compose are just exposing the services. And also Docker Compose is very like, uh, ports driven versus-

Ben Firshman48:48

Mm-hmm

Alessio48:48

... you have like the actual, you know, predict this is what you have to run. Yeah, any learnings and maybe tips for other people building container-based runtimes? Like how, how much you should just separate the API services versus the, the image building or how much you wanna build them together?

Ben Firshman49:07

I think it was coming from two sides. We were thinking about the design from the point of view of user needs, like what do users-- what are their problems and what, what problems can we solve for them, but also what the interface should be for a machine learning model, and it's sort of the combination of two things that led us to this design.

So the thing I talked about before was a little bit of, like, the interface around the machine learning model. So we realized that we want it to be general purpose. We want it to be at the, like, the JSON, like, human readable things rather than the, the tensor level.

Um, so it's like an OpenAPI specification that wrapped a Docker container. That's where that design came from, and it's really just a wrapper around Docker, so we're kind of building on, standing on, on shoulders there. But we-- Docker's too low level, so it's just like arbitrary software.

Alessio49:56

Mm-hmm.

Ben Firshman49:57

So we ne- we wanted to be able to, like, have a OpenAPI specification there that defined the function effectively that is the machine learning model, but also, like, how that function is written, how that function is run, which is all defined in code and stuff like that.

So it's like a bunch of abstraction on top of Docker to make that work, and that's where that design came from. But the core problems we were solving for users was that, was that Docker's really hard to use and, and productionizing machine learning models is really hard.

Alessio50:31

Right.

Ben Firshman50:32

So on the first part of that, uh, we knew we couldn't use Docker files. Like, Docker files are hard enough for software developers-

Alessio50:40

They are

Ben Firshman50:40

... to write.

Alessio50:41

Yeah.

Ben Firshman50:41

I'm saying this with love as somebody who works on Docker and, like, works on Dock- on, on Docker files. Um, but it's really hard to use, and you need to know a bunch about Linux basically 'cause you're running a bunch of CLI commands.

You need to know a bunch about Linux and best practices and, like, how apt works and all this kind of stuff. So we're like, "Okay, we can't, we can't get to that level. We need to... We need something that machine learning researchers will be able to understand, like people who are used to, like, Colab notebooks."

Alessio51:03

Mm-hmm.

Ben Firshman51:03

And what they understand is they're like, "I need this version of Python, I need these Python packages, and somebody told me to apt-get install something." You know?

Alessio51:12

And throw a sudo in there why not?

Ben Firshman51:13

And I don't really know.

Alessio51:14

Right.

Ben Firshman51:14

And I don't really know what that means. Um, so we tried to create a format that was at that level, and that's what Cog.YAML is. And we're really kind of trying to imagine, like, what is that machine learning researcher gonna understand, you know, and trying to build for them.

And then the productionizing machine learning models thing is like, okay, how can we package up all of the complexity of, like, productionizing machine learning models, like picking CUDA versions-

Alessio51:39

Mm

Ben Firshman51:39

... like hooking it up to GPUs, writing an inference server, um, defining a schema, doing batching, um, all of these just, like, really gnarly things that everyone does again and again, and just, like, you know, provide that as a tool.

Uh, and that's where, that's where that side of it came from. So it's like combining those user needs with, you know, the, the sort of world need of needing, like, a-

Alessio52:05

Mm-hmm

Ben Firshman52:05

... a common standard for, like, what a machine learning model is, and that's, that's how we thought about the design. I don't know whether that answers the question.

Alessio52:11

Yeah. So your idea was like, hey, you really want what Docker, uh, stands for in terms of standard, but you actually don't want people to do all the work-

Ben Firshman52:20

Yeah

Alessio52:20

... that goes into Docker.

Ben Firshman52:21

It needs to be higher level, you know?

Alessio52:23

Mm-hmm.

Swyx52:24

Um, so I want to, for the listener, um, you're not the only standard that is out there. As with any standard, there must be fourteen of them.

Ben Firshman52:31

Yeah.

Swyx52:31

Um, you are very surprisingly friendly with Ollama-

Ben Firshman52:34

Yeah

Swyx52:34

... who is your former colleagues from Docker, uh, who came out with the model file. Uh, Mozilla came out with the Llama file.

Ben Firshman52:41

Yep.

Swyx52:41

And then, um, I don't know if this is in the same category even, but I'm just gonna throw it in there. Like, Hugging Face has the Transformers and Diffusers library, which is a way of disseminating models that-

Ben Firshman52:49

Yep

Swyx52:49

... obviously people use. Um, how would you compare your... contrast your approach of Cog versus all these?

Ben Firshman52:55

It's kind of complementary actually, which is kind of neat, in that a lot of... Like, Transformers, for example, is lower level than Cog, so it's, you know, a Python library effectively, but you still need to, like-

Swyx53:08

Expose them.

Ben Firshman53:08

Yeah. You still need to turn that into an inference server. You still need to, like, install all the Python packages and that kind of thing. So lots of Replicate models are Transformers models, uh, and Diffusers models inside, inside Cog.

You know, so that's, like, the level that that sits. So it's very complementary in some sense, and, you know, we're kind of working on integration with Hugging Face such that you can, like, deploy models from Hugging Face and into Cog models and stuff like that-

Swyx53:31

Oh

Ben Firshman53:31

... into Replicate. Um, and, um, so some, some of these things like, uh, Llama File and what Ollama are working on are also very complementary in that they're, they're, they're doing a lot of the sort of running these things locally on laptops, which is not a thing that works very well with Cog.

Like, Cog is really designed around servers and attaching to CUDA devices and, and Nvidia GPUs and this kind of thing. So,

like, we're trying to figure out, we're actually, like, you know, figuring out ways that, like, we can... those things can be interoperable, 'cause, 'cause, you know, they should be. And, um, I think they, they are quite complementary in that you should be able to, like, take a model and replicate it and run it on your local machine.

You should be able to take a model on your local machine and run it in the cloud. Uh, so yeah.

Swyx54:19

Now, is the base layer something like, um, like is it, is it at the, like, the GGUF level? Which, uh, by the way, I, I need to get a primer on, like, the, the different formats that have emerged.

Uh, or is it at the star.file level, which is model file, Llama file, whatever, whatever? Um, or is it at the Cog level?

Ben Firshman54:37

I don't know, to be honest.

Swyx54:38

Yeah.

Ben Firshman54:38

And I think this is something we still have to, still have to, still have to figure out. Um, I think there's a lot, there's a lot, yeah. Like, exactly where those lines are drawn, don't know exactly. And I think this is something we're trying to figure out ourselves.

But, uh, but I think there's certainly a lot of promise about these systems interoperating. I think we, we just want, we just want things to work together. You know, we wanna try and reduce the number of standards, so the more, the more these things can interoperate and, you know, convert between each other and that kind of stuff at the manner.

Alessio55:01

A-Andreas comes out of Spotify. Um, Eric from Modo also comes out of S-Spotify. Um, you worked at Docker, and the Ollama guys worked at Docker. Um, where...

Swyx55:13

Did you know that these ideas were in ... Did both you and Andreas know that there was somebody else you worked with that had a kinda like similar, not similar idea, but like was interested in, in the same thing, or did you then just see, "Oh, I know those people.

They're doing something very similar"?

Ben Firshman55:28

We learn, we learn about both early on, actually. Yeah. Uh, 'cause we know, we know them both quite well. And it's funny how I think we're all seeing the same problems, and just like applying, you know, trying to fix the same problems that we're all seeing.

Swyx55:40

Mm-hmm.

Ben Firshman55:41

I think the Ollama, Ollama one's particularly funny because, um, I joined Docker through my startup. Funnily, actually, the thing which worked for my startup was Compose, but we were actually working on another thing, which was a bit like EC2 for Docker.

Swyx55:56

Mm.

Ben Firshman55:56

So we were working on, like, productionizing Docker containers, and Ollama was working on a thing called Co- Kitematic, which was a bit like a, a desktop app for Docker. Um, so ... And our companies both got bought by Docker at the same time.

And, you know, Kitematic turned into Docker Desktop, and then, you know, our thing then turned into Compose. Uh, and it's funny how we're both applying our ... Like, the things we saw at Docker to the AI world.

Swyx56:25

Yeah.

Ben Firshman56:25

Where they're building, like, the local environment for us, and we're building, like, the cloud for it. Um, and yeah, so that's just, like, really pleasing, and I think, you know, we're, we're collaborating closely 'cause there's just so much, so much opportunity for working there.

Um-

Swyx56:40

When you have a hammer, everything's a nail.

Ben Firshman56:42

Yeah, exactly. Exactly. So I, I think a lot of, a lot of ... This is, I mean, where we're coming from a lot f- with AI is we're taking a lot of things that ... 'Cause we're all kind of, on the Replicate team, we're all kind of people who have built developer tools in the past.

So we've got a team ... Like, I worked at Docker. We've got people who worked at Heroku and GitHub and, like, the iOS ecosystem, and all this kind of thing. Like, the previous generation of, of developer tools, where we, like, figured out a bunch of stuff, and then, like, AI's come along, and we just don't yet have those tools and abstractions, like, to make it easy to use.

So we're trying to, like, take the lessons that we learnt from the previous generation of stuff and apply it to this new generation of stuff. And obviously, there's a bit of nuance there, 'cause the trick is to take, like, the right lessons and do new stuff where it makes sense.

You can't just, like, c- cut and paste, you know?

Swyx57:36

Mm.

Ben Firshman57:36

Uh, but that's, like, how we're approaching this, is we're trying to, like, as much as possible, like, take some of those lessons we learned from, like, you know, how Heroku and GitHub was built, for example, and apply them to, apply them to AI.

Compute & Waste57:49

Swyx57:49

Excellent. Um, we should, um, also talk a little bit about, um, your compute av- uh, availability. We're trying to ask this of all ... You know, it's Compute Provider Month. Um, do you own your own GPUs? How many, uh, do you have access to?

Do you f- what do you feel about the tightness of the GPU market?

Ben Firshman58:06

We don't own our own GPUs. We've got a few that we play around with, but not, not for production workloads. And we are primarily built on just public cloud, so primarily GCP and CoreWeave, and, like, some smatterings elsewhere.

And

...

Swyx58:22

N- none, none from Nvidia, which is your newest investor?

Ben Firshman58:24

We work with Nvidia, so, you know-

Swyx58:27

Yeah

Ben Firshman58:27

... they're, they're, they're kind of helping us get GPU availability. Um, I think GPUs are hard to get hold of if you ... Like, if you go to AWS and ask for one A100, they won't give you an A100.

But if you go to AWS and say, "I would like 100 A100s for two years," they're like, "Sure. We've got some." Um, and I think the problem, the problem is, is the cloud providers, the cloud providers ... Like, that, that makes sense from their point of view.

They want just, like, reliable, sustained usage. They don't want, like, spiky usage and, like, wastage in their infrastructure, which makes total sense. But that makes it really hard for startups, you know, who are wanting to just, like, get hold of GPUs.

I think we're in a fortunate position where we can aggregate demand, so we can make commits to cloud providers. Um, and then, you know, we actually have good availability. Like, it's not, it's not, um, it's not ... You know, we don't have infinite availability, obviously, but, you know, if you want an A100 from Replicate, you can get it.

Um, uh, but, you know, we're seeing other, other companies pop up as well. Like, SF Compute's a great example of this, where they're doing the same idea for training almost, where, you know, a lot of startups need to be able to train a model, but they can't get hold of GPUs from large cloud providers.

So SF Compute are, uh, like, letting people rent, you know, 10 H100s for two days, which is just impossible otherwise.

Swyx59:46

Yeah.

Ben Firshman59:46

And, you know, what they're effectively doing there is they're aggregating demand such that they can make a big commit to the cloud provider, and then let people use smaller chunks of it. And that's kinda what we're doing for Replicate as well, where we're make- we're aggregating demand such that we make big commits to the cloud providers, and, you know, then people can get, can run, like, a, a 100 millisecond API request on an A100.

Swyx1:00:05

Coming from a finance background, this sounds surprisingly similar to banks.

Ben Firshman1:00:09

Mm.

Swyx1:00:09

Where the, the job of a bank is, um, maturity transformation is, is, is what you call it. You, you take short-term deposits, which can, which technically can be withdrawn at any time, and you turn that into long-term loans, uh, for mortgages and stuff, and you pocket the difference in interest, and that's, that's the bank.

Ben Firshman1:00:24

Yep. That's e- that's exactly what we're doing.

Swyx1:00:26

So you run a bank.

Ben Firshman1:00:27

Yeah, a GPU bank.

Swyx1:00:28

Right, yeah. And it, it's, it's supp- so much a finance problem as well, because we have to, we have to make bets on the future demand-

Ben Firshman1:00:35

We have to do forecasting

Swyx1:00:36

... for value of GPUs. Yeah. Um-

Ben Firshman1:00:38

What, what are you ... Okay. I, I, I don't know how much you can disclose, but w- what are you forecasting?

Swyx1:00:43

Um-

Ben Firshman1:00:44

Up, down? Up a lot?

Swyx1:00:46

Yeah. Um-

Ben Firshman1:00:46

Up 10X? Up-

Swyx1:00:47

I can't really, can't really ... It, it ... We're projecting our growth with some educated guesses about what kind of models are gonna come out, and what kind of models these will run, you know?

Ben Firshman1:00:54

Okay.

Swyx1:00:55

So we, we need to, we need to, we need to bet that, like, okay, maybe language models are getting larger, so we need to, like, have GPUs with a lot of RAM, or, like, multi-GPU nodes, or maybe models are getting smaller and we actually need smaller GPUs.

You know, we have to make some educated guesses about that kind of stuff, yeah.

Ben Firshman1:01:08

Yeah. Speaking of which, um, the mixture of experts models are, must be throwing a spanner in, into the planning. Um, not so much. I mean, we've got, we've got, we've got s- like multi-node A100 machines which can run this, and multi-node H100 machines which can run this no problem.

So-

Swyx1:01:24

Okay

Ben Firshman1:01:24

... we, uh, we, uh, we, we, we, we, uh, we're set up for that, for that, for that world, yeah.

Swyx1:01:30

Um, okay. Right. I, I didn't, I didn't expect it to be so easy. Um, I mean, the-- my impression was that the amount of RAM per model was increasing a lot, um, s- especially on a sort of per parameter basis.

Ben Firshman1:01:42

Mm-hmm.

Swyx1:01:42

Per active parameter basis. Um, i mean, g- going from like Mi- Mixtral being eight experts, um, to like the DeepSeek MOE models, I don't know if you saw them-

Ben Firshman1:01:50

Mm

Swyx1:01:50

... being like 30, 60 experts, and you can see it, it keep going up, I guess. I don't know.

Ben Firshman1:01:56

Yeah. I think we might run into problems at some point. Um, and yeah, I don't know exactly, exactly what's going on there. Um, I think something, something that we're finding which is kind of interesting, like I don't know this in depth, um, but, um, you know, we're certainly seeing a lot of good results from, from, from, uh, lower precision models.

So like, you know, 90% of the performance with just like much less RAM required. Um, and you know, that's, that's-- that means that we can run them on GPUs we have available, and it's good for customers as well because, 'cause it runs faster and like they, they want that trade-off, you know, where, where it's just slightly worse, but like way faster and cheaper.

Yeah.

Swyx1:02:40

Do you see a lot of, uh, GPU waste in terms of people running the thing on a GPU that is like too advanced? I think we use a T4, uh, to run Whisper, so we- we're at the bottom end of it.

Um, yeah, any thoughts? I think, uh, uh, one of the hackathons we were at, people were like, "Oh, how do I get access to like H100s?" And it's like, you need to run like Stable Diffusion.

Ben Firshman1:03:01

Dude, you don't need an H100.

Swyx1:03:01

It's like you don't need an H100.

Ben Firshman1:03:02

Yeah. Yeah. Well, if you want low latency, you like sure, like spend a lot of money on an H100. Um, uh, yeah, we see a ton of that kind of stuff, and it's surprisingly, it's surprisingly hard to optimize these models right now.

So a lot of people are just running like really unoptimized models. We're doing the same, honestly. Like where a lot of models on Replicate have just been like not been optimized very well. Um, so something we want to like be able to help people with is optimizing those models.

Like either, either we, you know, show people how to with guides, or we make it easier to use some of these more optimized inference servers, or we show people how to compile the models, or we do that automatically, or something like that.

But that's certainly something we're exploring, 'cause yeah, there's, there's so much wastage. Like it's not just wasting the GPUs, it's also like a bad experience and the models run slow, you know?

Swyx1:03:55

Right.

Ben Firshman1:03:56

So like a lot of, a lot of the models on Replicate, some of the most popular models on Replicate we have-- So the, the models on Replicate are, are almost all pushed by our community, like people have pushed those models themselves.

But like it's like a big-headed distribution where there's like a long tail of lots of models that people have pushed, and then like a big head of like the models most people run.

Swyx1:04:16

Mm-hmm.

Ben Firshman1:04:17

So models like Llama 2, like Stable Diffusion, we, um, we, you know, we work with Meta and Stability to like maintain those models, and we've done a ton of optimization to work those, make those really fast. So, um, yeah, those models are optimized, but the long tail is not, and there's like a lot of, a lot of wastage there.

Swyx1:04:35

Yeah. And going into the... Well, it's already the new year. Um, do you see the customer demand and the GPU like hardware demand kind of like staying together? Because I think a lot of people are saying, "Oh, there's like hundreds of thousands of GPUs being shipped this year, like the, the crunch is gonna be over."

But you also have like millions of people that now care about using AI. You know, uh, uh, h- how do you see the two lines progressing? Are you seeing customer demand that's gonna outpace the GPU growth? Do you see them together?

Do you see maybe a lot of this like model improvement work kind of helping alleviate that?

Ben Firshman1:05:09

From our point of view, demand is not outpacing supply of GPUs. Like we have enough, from our point of view, we have enough GPUs to go around, but that might change for sure.

Swyx1:05:18

Yeah. That's a very, um, nicely put way as a startup founder to respond. Like... Yeah. Uh, I'll maybe-

Ben Firshman1:05:27

Yeah

Swyx1:05:27

... get into a little bit of this on the... You s- you said optimizing models. Actually, so like when Alessio said-- talked about GPU waste, he was more... Oh, that you.

Ben Firshman1:05:34

Sorry.

Swyx1:05:35

Uh, just-

Ben Firshman1:05:36

One second. Yeah.

Swyx1:05:36

Yeah, it is getting a little bit warm in here. Just some greenhouse gas effect. Um, so, so Al- Alessio framed it more as like sort of picking the wrong box model, whereas yours is more about, um, m- maybe the inference stack, if you can call it.

Were you referencing vLLM? Um, w- what, what other sort of techniques are you referencing? And also keeping in mind that when I talk to your competitors, I d- and, and I don't know if, um, we don't have to name any of them, but they are working on trying to optimize the kinds of models.

Like they basically, they'll, they'll quantize their models for you with their special stack. So you, you basically use their versions of Llama 2, you use their versions of Mistral, and that's one way to, to approach it. Uh, I don't see it as the Replicate DNA to do that because that would be like sort of you would have to slap the Rep- Replicate house brand on something, which...

I mean, just comment on any of that. Like what, what do you mean when you say optimize models?

Ben Firshman1:06:28

Yeah, I mean, you know, things like, I mean, quantizing the models. You can imagine a way that we could help people quantize their models if we want to. Um, we've, um, we've had success using inference servers like vLLM and TRT-LLM, um, and we're using those kind of things to serve language models.

We've had success with things like AI templates which compile, compile the models, um, all of those kind of things. And there's like some even really just boring things of just like, um, making the code more efficient. Like some people, like when they're just writing- ...

some Python code, it's really easy to just write, write inefficient Python code, you know? Um, there's like really boring things like that as well. Um, but it's like a whole smattering of things like that. Um, um, and-

Swyx1:07:15

So you will do that for a customer? Like you, you'll look at their code and-

Ben Firshman1:07:18

We-- yeah, we've certainly helped some of our customers be able to do, do that some of the stuff.

Swyx1:07:21

Wow.

Ben Firshman1:07:21

That some stuff, yeah. And a lot of the models on, like, the popular models on Replicate, we've, like, rewritten them to use that stuff as well.

Swyx1:07:28

Okay.

Ben Firshman1:07:29

Um, and like, like the stable diffusion that we run, for example, is compiled with AI template to make it super fast. And, you know, it's all open source that you can see all of this stuff on GitHub if you wanna, if you wanna like see how we do it.

Um, but you can imagine ways that we could help people, you know, it's almost like built into the Cog layer maybe, where we could help people, like, use these fast inference servers or use AI template to compile their models to make it faster, whether it's like manual, semi-manual, or automatic, we're not really sure.

You know, but that's something we want to explore 'cause, you know, that benefits everyone.

Swyx1:07:59

Yeah. Awesome. Yeah, and then on the competitive piece, um, there was a price war on Mixtral last year, last year, this last December. Um, as far as I can tell, you guys did not enter that war. Um, you have, you have Mixtral, but you, you know, you-- it's just regular pricing.

Um, I, I think also some of these com- some of these players are probably losing money, um, on, on their pricing. Um, you know, you don't have to say anything, but it's, you know, it's somewhere be- the break-even is somewhere between fifty to seventy-five cents per million tokens, uh, served.

Um, how are you thinking about like the just the overall competitiveness in the market? How, how should people choose when everyone's an API?

Ben Firshman1:08:37

We actually for our-- So for Llama Two and Mistral, I think not Mixtral, but I can't remember exactly, we have, you know, similar performance and similar price to some of these other ser- other services. We're not like bargain basement like to some of the others 'cause to your point, like we don't wanna like burn tons of money.

Um, but we're, you know, pricing it sensib- sensibly and sustainably, um, to a point where we think it's, we think, you know, it's competitive with other people such that... You know, the thing we don't want-- We-- Like, we, we want developers using Replicate, and we don't wanna, we don't wanna like price it such that it's like only affordable by big companies.

You know, we wanna make it, we wanna make it cheap enough such that the developers can afford it, but we also don't like want the super cheap prices 'cause then like it's almost, it's almost like then your customers are hostile, you know?

Swyx1:09:27

Mm-hmm.

Ben Firshman1:09:27

And the, like, the more customers you get, the worse it gets, you know. So we're, we're pricing it sensibly, but still to the point where, you know, uh, where hopefully it's cheap enough to build, build on. Um, and I think the thing we really care about, like we want to, we want to, like obviously we want, you know, models on Replicate to be comparable to other people.

Swyx1:09:47

Mm-hmm.

Ben Firshman1:09:48

Um, but I think the really crucial thing about Replicate and the way I think we think about it is that it's not just the API for the-- particularly in open source, it's not just the API for the model that is the important bit.

It's-- Because quite often with open source models, like the whole point of open source is that you can tinker on it, and you can customize it, and you can fine-tune it, and you can like smush it together with another model, like, like Lava, for example.

Swyx1:10:13

Mm-hmm.

Ben Firshman1:10:13

Um, and you can't do that if it's just like a hosted API 'cause it's just like it's, you know, it's, it's, you know, you can't, you can't touch the code. Um, so

that's... Like what we wanna do with Replicate is build a platform that's actually open. So like we've got all of these models where the performance and price is on par with everything else. But if you wanna customize it, you can fine-tune it.

You can go to GitHub and get the source code for it and edit the source code and push up your own custom version and this kind of thing. Because that's like the, that's the crucial thing for open source mach- machine learning, is being able to tinker on it and customizing it.

Um, and we think, we think, we think that's really important for, for, for, you know, to make open source AI work.

Swyx1:10:58

Um, you mentioned open source. How do you think about levels of openness? When Llama Two came out, uh, I wrote a post about this, about it's like open source, and there's open weights, then there's restricted weights. It was on the front page of Hacker News, so there, it, there was like all sort of comments from, from people.

Open Source & Future1:10:58

Swyx1:11:14

So I'm always curious to hear your thoughts. Like what do you think is okay for people to license? What's okay for people to r- not release? Um, yeah.

Ben Firshman1:11:25

Yeah, I was saying, I mean, you know, before it was just like closed source, big models, open source, little models. You know, purely open source stuff. And we're now seeing like lots of variations where, you know, model companies, uh, putting restrictive licenses on their models.

Um, you know, that means it can only be used for non-commercial use, you know. And a lot of the, you know, open source crowd is complaining it's not true open source, you know, and all this kind of thing.

And yeah, I think a lot of that is coming from philosophy, you know, of like the sort of free software movement kind of philosophy. And I don't think it's necessarily a bad thing. Like it's-- I think it's good that model companies can make money out of their models.

You know, that's like how-- it's what will incentivize people to make more models and this kind of thing. And I think it's totally fine if like somebody made something to ask for some money in return if you're making money out of it, and I think that's, that's totally okay.

And I think there's some really interesting like midpoints as well, where people are releasing the codes. You can still tinker on it-

Swyx1:12:21

Mm-hmm

Ben Firshman1:12:22

... but the person who trained the model still wants to get a cut of it if like you're making a bunch of money out of it, and I think that's, that's good, and that's gonna make like the ecosystem more, more sustainable.

And I think we're just gonna see-- I don't think anybody's really figured it out yet, and we're gonna see like more experimentation with this and more people like try to figure out like, "Hmm, what are the business models around building models, and how can I make money out of this?"

And we'll just see where it ends up, and I think it's something we want to support as Replicate as well 'cause we're-- we believe in open source. We think it's great, but there's also gonna be lots of models which are closed source as well, and these companies might not be-- There's probably gonna be a long tail of a bunch of people building models that don't have the reach that OpenAI have, and, you know, hopefully as Replicate, we can help those people find developers and, and help them make money and that kind of thing.

Swyx1:13:13

Yeah. I, I think the computer requirements of AI kind of change the thing. I, I started an open source company. I'm a big open source fan, and before it was kind of man-hours was really all that went into open source.

It wasn't much monetary investment.

Alessio1:13:27

Well, not that man-hours are not worth a lot, but if you think about Llama 2, it's like, it's like twenty-five million dollars, you know, like all in. It's like you can't just spin up a Discord and like spend twenty-five million dollars.

So I think it's net positive for everybody that Llama 2 is open source. And, uh, well, is the open source-- You know, is the open source term-- I, I think people, like you're saying, it's like they kind of argue on the semantics of it.

But like all we care about is that Llama 2 is open. Because if Llama 2 wasn't open source today, like the-- if Mistral was not open source, we would be in a bad spot, you know? So-

Ben Firshman1:14:03

And I think the nuance here is making sure that these models are still tinkerable, because the beautiful thing about Llama 2 as a base model is that, like, yeah, it costs twenty-five million dollars to train to start with, but then you can fine-tune it for like fifty bucks.

Alessio1:14:17

Right.

Ben Firshman1:14:18

And that's what's so beautiful about the open source ecosystem, and something I think is really surprising as well. It completely surprised me. Like, I think a lot of people assumed that, um, uh, like it's not gonna be-- open source machine learning is just not gonna be practical because it's so expensive to train these models.

But like fine-tuning is unreasonably effective, and people are getting really good results out of it, and it's really cheap. So people can effectively create open source models, um, really cheaply, and there's gonna be like this sort of ecosystem of tons of models being made.

And I think the risk there from a licensing point of view is we need to make sure that the licenses let people do that.

Alessio1:14:56

Mm-hmm.

Ben Firshman1:14:56

Because if you release a big model under a non-commercial license and people can't fine-tune it, you've lost the magic of it being open. And I'm sure there are ways to structure that such that the person paying twenty-five million dollars feels like they're compensated somehow, and they can feel like they can-- you know, they should keep on training models, um, and people can keep on fine-tuning it.

But I guess we just have to figure out exactly how that plays out.

Swyx1:15:20

Yeah. Excellent. Um, so just wanted to round it out. Uh, you've been, you've been an excellent, a very open guest so far. Um, I actually kind of-- I, I should have started this-- started my intro with this, but I feel like you found the sort of AI engineer crew before I did.

And, uh, you know, something that really resonated with you in sort of, sort of the Series B announcement was that, um, you put in some stats here about how there are two orders of magnitude more software engineers than there are machine learning engineers, about thirty million software engineers and five hundred thousand machine learning engineers.

Um, you can maybe plus or minus one of those orders of magnitude, but it's around that ballpark. And so obviously, there will be a lot more AI engineers than there will be ML engineers. Um, how do you see this group?

Like, is it all software engineers? Are they going to specialize? Um, what would you advise someone trying to become an AI engineer? Is this a legitimate career path?

Ben Firshman1:16:14

Yeah, absolutely. I mean, it's very clear that AI is gonna be a large part of how we build software in the future now. It's a bit like being a software developer in the nineties and ignoring the internet, you know?

You just need to-- You need to learn about this stuff, and you need to figure this stuff out. I don't think it needs to be-- You don't need to be like super low level. You don't need to be like...

You know, the metaphor here is, is like you don't need to be digging down into like, uh, this sort of PyTorch level if you don't want to. In the same way as a software engineer in the nineties, you don't need to be like understanding how network stacks work to be able to build a website, you know?

But you need to understand the shape of this thing and how to hold it and what it's good at and what it's not, and, and, uh, that's really important. So yeah, certainly just advise people to like just start playing around with it, get a feel of like how language models work, get a feel of like how these diffusion models work, get a feel of like what fine-tuning is and how it works, because some of your job might be building datasets.

You know? Get a feeling of how prompting works, 'cause some of your job might be writing a prompt. And, uh, those are just all really important skills to, skills to, to sort of figure out.

Swyx1:17:29

Yeah. Well, thanks for building the definitive platform for doing all that.

Ben Firshman1:17:34

Yeah, of course.

Alessio1:17:35

Um, any final call to actions? Who should come work at Replicate? Um, yeah, anything for the, for the audience?

Ben Firshman1:17:42

Yeah. Well, I mean, we're, we're hiring. If you, if you click on jobs at the bottom of our, of replicate.com, there's, there's some jobs. Uh, and I just encourage you to like just like try out AI, even if you don't-- even if you think you're not smart enough.

Like the whole reason I started this company is because I was looking at the cool stuff that Andreas was making. Like Andreas is like a proper machine learning person with a PhD, you know? And I was like, just like a, you know, a sort of lonely software engineer, and I was like, "You're doing really cool stuff, and I wanna be able to do that."

Um, and by us working together, you know, we've now made it, made it accessible to dummies like me. And I just encourage anyone who's like wants to try this stuff out, just give it a try. And I think I would also encourage people who are tool builders.

Like the, the, the limiting factor now on AI is not like the technology. Like the technology's made incredible advances, and there's just so many incredible machine learning models that can do a ton of stuff. The limiting factor is just like making that accessible to people who build products, 'cause it's really hard to use this stuff right now.

And obviously, we're building some of that stuff as Replicate, but there's just like a ton of other tooling and abstractions that need to be built out to make this stuff usable. So I just encourage people who like, like building developer tools to just like get stuck into it as well, 'cause that's gonna make this stuff accessible to everyone.

Swyx1:18:59

Yeah. I, I especially wanna highlight you have a hacker-in-residence job opening available, which not every company has, which means just join you and hack stuff. I think Charlie Holtz is doing a fantastic job of that.

Ben Firshman1:19:10

Yep. Effectively, like most of our-- a lot of our job is just like showing people how to use AI.

Swyx1:19:16

Mm-hmm.

Ben Firshman1:19:16

So we've just got a team of like software developers and people who've kind of figured this stuff out, who are writing about it, who are, uh, you know, uh, making, making videos about it, who are making example applications just to like show people what you can do with this stuff.

Swyx1:19:29

Yeah. In, in my world, that used to be called DevRel. But now, now it's hacker-in-residence, and that's, uh...

Ben Firshman1:19:35

Yeah, this, this came, this came from, um-- Zeke is another one of our-

Swyx1:19:39

Zeke, yes

Ben Firshman1:19:39

... of our hackers. Um-

Swyx1:19:41

Tell me this came from Chroma, 'cause, 'cause I... To start that one.

Ben Firshman1:19:44

We developed-- Like they, they-- Anton actually was like, "Hey, we came up with that first," but I think we came up with it independently.

Swyx1:19:49

Oh, yeah, I made that page, yeah.

Ben Firshman1:19:50

I think we came up with it independently, because the story behind this is, is we originally called it the DevRel team.

Swyx1:19:57

Yeah.

Ben Firshman1:19:57

And, um-

Swyx1:19:58

DevRel's cursed now. Everyone who works with us in the DevRel-

Ben Firshman1:20:01

Zeke was like, Zeke was like, "That sounds so boring." "I don't wanna go to someone and say I'm a Dev-De-Developer Relations person."

Swyx1:20:05

I wanna be a hacker man.

Ben Firshman1:20:08

Or a developer advocate or something. So we were like, "Okay, what's like the way we can make this sound the most fun? All right, you're a hacker."

Swyx1:20:15

Yeah. Um, I would say like that, that is consistently the vibe I get from Replicate. Every-everyone on, on your team I interact with. When I go to your, your San Francisco office, like that's the vibe that you're generating.

Like it's, it's a hacker space more than an office. Um, and you hold, uh, fantastic meet up, meet ups there, and I think you're a really positive presence in our community. So thank you for doing all that, and it's instilling the hacker vibe and culture into AI.

Ben Firshman1:20:37

Oh, I'm really glad that, I'm really glad that's working.

Swyx1:20:39

Yeah.

Alessio1:20:39

Cool. That's a wrap, I think. Thank you so much for coming on, man.

Ben Firshman1:20:43

Thank you. Yeah, of course. Thank you. This was a lot of fun.