LALatent SpaceMay 20, 2026· 1:29:54

The Agent-Native Cloud: 3M Users, 100K Signups/Wk, Data Centers, & Death PRs — Jake Cooper, Railway

Jake Cooper, founder of Railway, argues that the next era of software infrastructure requires an agent-native cloud built on own-metal data centers, with primitives for version control, observability, and orchestration at 1000x scale. Railway grew from a slow six-year grind to 3 million users and 100,000 weekly signups with only 35 people, after surviving a $500,000/month loss on its free tier and rebuilding the business. Cooper explains that building its own data centers yields a three-month payback period and 70% margins, while using cloud bursting (AWS, GCP, Oracle) for overflow, and that data center debt is a better tool than venture debt for infra startups. He details how agents need CLIs with many flags, safe production forks, feature flags, and incremental rollouts, and why the pull request is dying in favor of prompt requests. Railway relies on Temporal for orchestration but may build its own workflow engine, and uses an internal tool called Central Station to aggregate customer feedback and incidents. Cooper advocates using agents to generate and review code instead of writing it by hand, and says focus on primitives—not GPUs—is key for now.

  1. 0:00Intro
  2. 7:34The Slow Grind
  3. 14:40Bare Metal
  4. 26:45Agent Needs
  5. 38:00Central Station
  6. 43:36SRE Agents
  7. 52:41Serverless Shift
  8. 59:36Temporal Deep Dive
  9. 1:05:27Railpack
  10. 1:07:14Death PRs
  11. 1:20:50Solo Founder
  12. 1:28:00Future Cloud

Powered by PodHood

Transcript

Intro0:00

Jake Cooper0:00

If you're writing code by hand, you are doing this wrong, right? Like, their code- the tools are good enough at this point, like, that you can move extremely, extremely quickly. And like, yes, there are issues and pain points and all these other things, um, but you should be reviewing the code that you are writing instead of trying to go in and write it by hand.

Like, all of those architectural patterns, all of those other things, like-

Alessio0:19

Yeah

Jake Cooper0:19

... you're not- just don't- like, you're not gonna throw them in the garbage or whatever. Actually, they matter more now than it- than, than any other time. But you just shouldn't spend your time generating code that you would write.

Like, if you know how to go in and write it, just, like, ask the agent to go in and write it, and then reconcile it until it looks like you would've written it.

Alessio0:35

Before we get into today's episode, I just have a small message for listeners. Thank you. We would not be able to bring you the AI engineering, science, and entertainment content that you so clearly want if you didn't choose to also click in and tune into our content.

We've been approached by sponsors on an almost daily basis. But fortunately, enough of you actually subscribe to us to keep all this sustainable without ads, and we wanna keep it that way. But I just have one favor to ask all of you.

The single most powerful, completely free thing you can do is to click that subscribe button. It's the only thing I'll ever ask of you, and it means absolutely everything to me and my team that works so hard to bring The Latent Space to you each and every week.

If you do it, I promise you, we'll never stop working to make the show even better. Now let's get into it.

Hey, everyone. Welcome to the Latent Space Podcast. This is Alessio, founder of Kernel Labs, and I'm joined by swyx, editor of Latent Space.

Jake Cooper1:31

Hey, hey, hey, and today we're in the studio with Jake Cooper of Railway.

Alessio1:34

Conductor-

Jake Cooper1:35

Very ex-

Alessio1:35

... of Railway.

Jake Cooper1:35

Conductor at Railway, yeah.

Alessio1:37

Choo choo.

Jake Cooper1:37

Choo choo. Uh, very excited-

Alessio1:38

Do you actually have that, like, anywhere on, like, your business card?

Jake Cooper1:40

Well, we, like, we roughly call, like, people inter- well, I don't have a business card. Um, we're not, we're not that big yet. Um, at some point I will. I got handed a nice business card from the Supermicro folks, and I was like, "Damn, this is actually, like-

Alessio1:49

Mm.

Jake Cooper1:49

... this is pretty official." Um-

Alessio1:50

They're coming back-

Jake Cooper1:51

It's sick

Alessio1:51

... business cards.

Jake Cooper1:52

Yeah, yeah. They're, they're cool. Um, they're hip, they're jiggy. Um, but yeah, the, the whole conductor thing, like, we call some of our volunteer moderators conductors, you know. Yeah, so it's a good one. It's a good one. Like, we were trying to figure out what we wanna call each other internally, and there's, like, varying levels of, um, thought.

Some people are like, "Oh, it's super cringe." Like, just don't... Like, you don't need a name for, like, you know, people internally, and some people are like, "Oh yeah, we wanna call, like, each other, like, this thing or whatever."

I was like, we still don't have a really good one, you know? We've got, like, we've got, like, New Railcrews. We've got, like, Trainiacs. We've got, like... Nothing's, like, really stuck yet.

Alessio2:20

I like Trainiac. Trainiac sounds-

Jake Cooper2:22

Yeah

Alessio2:22

... good, sounds good.

Jake Cooper2:22

Yeah.

Alessio2:22

Railwayians. Um, okay. So well, for those who don't know, what is Railway? Let's give people a crisp definition up front.

Jake Cooper2:29

Yeah. Railway is the easiest way to ship anything. You just go to the canvas, or you talk with Claude, um, and you say deploy Postgres instance, deploy my GitHub repository, run this code, et cetera, right? Um, and you'll just be up and away to the races, right?

Um, you got-

Alessio2:43

Yeah, you have nice animation on the landing page.

Jake Cooper2:45

Oh. Well, thank you.

Alessio2:45

Yeah.

Jake Cooper2:45

Um, none of my work, by the way. Um, they, they don't let me touch any of the design stuff anymore. But yeah, we wanna make it trivially easy for not just to, like, deploy things, but for you to almost, like, evolve applications over time.

Like, we believe that most of the tooling right now is kind of, like, stacked up. Like, you're stacking entropy on top of entropy on top of entropy, right? So you have, like, Docker and Kube and then, like, Ansible scripts and all of these other things, right?

Um, and if we can kind of, like, version all of your software for you and keep track of all of the changes, then we can make it actually trivial for you to clone environments, you know, fork into a parallel universe, get copies of, like, production data, get copies of, like, any of your services, make those changes, validate those changes, collapse it in, uh, without kind of having to just, like, reproduce everything across a, you know, a staging environment or all of those other things, right?

So...

Alessio3:27

Yeah. Amazing. Uh, one thing I, I was looking at your background, right? Like, um, Bloomberg, Uber. There's nothing immediately that stands out to me as like, okay, this guy's gonna found, like, the next great platform as a service.

Uh, what prepared you for Railway?

Jake Cooper3:41

It's almost like a curiosity to just, like, ever go deeper, right? Um, and so like, you know, started out on, like, front end stuff, you know, like working on the, like, Wolfram, like, web Mathematica-

Alessio3:50

Yes, Wolfram

Jake Cooper3:50

... and, like, porting it over there. And then, you know, briefly moving to Bloomberg, um, and then moving towards Uber and, like, distributed systems and kind of like taking all the Jump Bikes, uh, kind of systems and, and moving them over to a distributed, uh, system built on top of, uh, Cadence, like, uh, the pre-

Alessio4:04

Temporal

Jake Cooper4:05

... yeah, the te- pre-Temporal, Temporal. Um-

Alessio4:06

Which, by the way, I'm happy to talk about pros and cons.

Jake Cooper4:09

Yeah. I think, like-

Alessio4:10

But, uh-

Jake Cooper4:10

... so it's, it's like-

Alessio4:10

We, we-

Jake Cooper4:10

It's my-

Alessio4:11

But let's, let's do the, uh-

Jake Cooper4:12

Totally

Alessio4:12

... the, the Railway story.

Jake Cooper4:13

A- and so, like, it's just been a, a s- a continual step of, like, I, I want this experience, whether it is, like, walking up to, like, a bike and just unlocking it and, like, having it be, like, frictionless to, like, work or whatever, and then, like, necessitating the, like, depth required to go in and make that happen, right?

Like, it, it... a lot of the work that, that I do and a lot of the team does is, like, it's all in service of that experience, right? And, like, we fundamentally don't care, like, how deep we have to go, whatever.

Like, we will swim to the bottom of the swimming pool to go and get the experience, right? Um, and I think that's what a lot of, of, you know, kind of the trajectory was, right? And so it's not like I have a physics PhD or whatever.

I did, like, an EECS degree. You know, it, it's just, it's always been about just trying to figure out that next step of, like, how do we get there, right? Um, and that's, like, what's led to, you know, starting Railway for that experience and then, like, moving all the way to bare metal data centers, right?

Like, you know, I was adding patches to the kernel this week, right? Just to, like, get the experience there, 'cause I'm like, "I see it, and, like, how much better it can be," right? Um-

Alessio5:10

You added patches to the Linux kernel this week?

Jake Cooper5:11

Yeah. Well, not upstream.

Alessio5:12

That's a flex.

Jake Cooper5:13

Uh, our, our, our, our fork.

Alessio5:14

Railpack? No, this is different.

Jake Cooper5:15

Mm-hmm.

Alessio5:15

This is the op- OS on top of Railpack.

Jake Cooper5:17

Yeah. No, this is like, this is a, a actual kernel, like patches.

Alessio5:19

Yeah.

Jake Cooper5:19

But it, it's, it's always literally just what do we have to do to get that experience and just, like, figure it out, right? Like, anything is figure-out-able, right? Like, you'll just, you'll just figure it out, you know? Um, so...

Alessio5:30

Would you send the patch upstream, or is it just because, like, it doesn't fit for the use cases?

Jake Cooper5:34

Maybe. It's, it's like we have to, we have to work out the experience for us, uh-

Alessio5:37

Yeah

Jake Cooper5:37

... internally. It has to do a, a lot with, um, the, like, storage layer that we're building, um, for some of the agentic stuff. Um, so maybe it'll be useful to, to people upstream, but it's, it's deeply useful for, for us internally.

Alessio5:49

Yeah. Yeah, we're... And I mean, you mentioned open source before, so I'm just kinda curious about, um, how you think about starting from open source and then coding agents let you do a lot more from forks of it.

Jake Cooper5:59

I think the, it's funny 'cause, like, I think GitHub's original sin is that it's, like, almost a series of broken pointers. Um, it's like you have essentially this thing, and then you clone it, and then, okay, great, like, I've just lost that whole upstream, right?

How do we make it trivial for people to modify really, really small pieces of it, right? And you d- like, you think of, of Git almost in this, like, discrete sense of, like, I've either made a change and I've merged upstream or I've, or I haven't, right?

Um, what would it look like if it was, like, percentage based or-

Alessio6:24

Mm-hmm

Jake Cooper6:24

... a little bit more non-deterministic or anything else like that? More of, like, a stream of changes that you kind of, like, traversed as a user, more as kind of like a, a percentage of this is ruled out in general, and it's been ruled all the way up, right?

Um, you know, we have the open source, like, kickback program, um, and allowing you to deploy those templates because we almost wanna make it trivial for people to, like, go and version these shards over time. It solves, like, a really, really large problem in terms of authentication, authorization, security.

Like, you know, NPM has that thing where you can, uh, almost define, hey, don't take any new packages or whatever.

Alessio6:55

Right. Yeah, yeah.

Jake Cooper6:55

Like, the ideal end state is actually, like, you should roll out progressively to the users who have the minimum impact zone for any of these things and just continually roll up, right? Like, JPMorgan or something else like that should probably be the last one on the patch line-

Alessio7:09

Right

Jake Cooper7:09

... uh, for that, right? For all of our sakes, right? Like, because we have all of our, you know, money and livelihood-

Alessio7:14

Livelihood

Jake Cooper7:14

... all of, all of those other things. It's okay if, like, Johnny Vibe Coder gets, like, a broken patch or something else like that because ultimately there's so much entropy in the system that you do have to, you do have to roll...

Like, rubber has to meet road at some point. Like, you have to test at, at varying levels, right? So yeah, a little diversion from wherever we started, but, you know.

The Slow Grind7:34

Alessio7:34

So I just wanted to, like, pull up this glorious chart you say, uh, which is basically your usage or number of-

Jake Cooper7:43

Uh, daily signups, I think.

Alessio7:44

Daily signups?

Jake Cooper7:44

Yeah, yeah.

Alessio7:45

Um, so you started six years ago, and-

Jake Cooper7:48

Yeah

Alessio7:48

... like, a slow grind.

Jake Cooper7:50

Slow grind, yeah.

Alessio7:50

Uh, and now, now obviously you're on a rocket ship. Uh, you say, "Don't doubt your fight and don't quit." But, like, maybe if you wanna pick out, like, certain points that were, like, sort of key inflections to the company, that, that might be fun.

Jake Cooper8:00

Oh, yeah, yeah, yeah. Well, I mean, at the start it's basically like how do you get your first hundred users, like, hell or high water, right?

Alessio8:05

Yeah.

Jake Cooper8:05

Um, and so, like, starting in, uh, you know, we had a website, and we had a support link, and the support link was the Discord channel, and you just showed up there, and I had notifications on. I had two monitors.

I had the monitor I was working on, and then I had the other monitor, and if anybody came in, I was like, "Oh, hey, how's it going?" Like, you know. It was like, and it was, like, super rare or whatever, so trying to get those initial, like, first hundred users to, like, actually kinda come back to it.

Um, and that's, I think, where you can kinda, like, see the really, like, in between January twenty twenty-one and twenty twenty-two, like, probably the middle. Like, there kind of, right? And that's like the, the start. Um, and then you ultimately end up building a consultancy factory of, like, users wanted all of these things in, in general, and so you kinda have to g- go back to the board a little bit and be like, "Well, what is the actual product offering that I wanna build on top of these?"

And I, I think, like, incidentally, uh, it's funny. Like, I think VCs really want, like, charts that, like, always look like this or whatever. Right? Um, but I think i- in reality you actually don't want charts that look like that.

Most companies, I think, or at least for us, there's been periods of, like, expansion, of like, "Oh, okay, we're gonna go and add these features to, like, go in and test these use cases," and then there's been periods of, like, compaction where we're saying, like, "Okay, how do we have...

If the experience we have is really, really good, how do we make it significantly better?" Right? Like, maybe we're even stripping out features that don't, like, fit our ICP anymore. Like, how do we go in and do that?

And I think throughout this whole chart you can see a lot of those things. Like, the, the boom in the, like, twenty twenty-two to twenty twenty-three is like we had a free tier, and, like, everybody under the sun was, like, using it and, and all those other things, right?

Alessio9:30

Yeah, a lot of Reddit bots and stuff.

Jake Cooper9:31

Yeah, right.

Alessio9:32

Discord bots.

Jake Cooper9:32

And, and, and, like, I think there's a, there's a thing that's really, really tough to, like, teach people or tell people about is, like, when you build an open product on the internet where anybody can sign up, the internet is a horrible place that has, like, so many things.

Like-

Alessio9:46

Oh, yeah

Jake Cooper9:46

... where the-

Alessio9:46

Did I tell you about my PC-

Jake Cooper9:47

Yeah, like, we got, you got-

Alessio9:48

PC and triad.

Jake Cooper9:49

Yeah.

Alessio9:49

Foreign crypto nurses.

Jake Cooper9:50

And then crypto miners. You got, like, all these other things, right?

Alessio9:52

Yeah.

Jake Cooper9:52

And so you kinda go through these periods of, like, "Well, how do I reach as many people as possible?" And then, like, "How do I fit in exactly the use case for the people who are really, really gonna matter and, and are gonna be really, really excited about specifically this thing," right?

And we go back and forth internally. And then there's, like, a, what is that? A two-year period of, like, making the actual business work in general.

Alessio10:13

Uh-huh.

Jake Cooper10:13

Right? So, like, free tier era, losing, what? I think half a million dollars a month, and, like, you know, we're making-

Alessio10:20

On, like, a twenty million b-bank account.

Jake Cooper10:22

Yeah. Yeah, on, like, a twenty million bank account with, like, uh, I don't know, like, maybe $50,000 a month in revenue or something else. Like, that is hor-horrible business. I don't know anybody about it. Um, but anyways, you have to kinda go through and, and be like, "Cool," like, we have an experience that people love in general, but, like, the business has to work, right?

And I think there's, there's, like, I guess, two schools of thought, is you can, you can continually run the horrible business all the way up, uh, in, in general and, and have bad margins, um, or you can actually go and go back and kind of make it work, right?

And for us, you know, we've always really wanted to have, like, a super lean team, right? So we're thirty-five people right now. You know, it's very, very small. We have, like, what? Three million-

Alessio10:57

Supporting two mi- uh, three million?

Jake Cooper10:58

Three million, yeah. Yeah, yeah.

Alessio10:59

Holy shit.

Jake Cooper10:59

Well, 'cause we're adding, like, 100,000 users a week, uh, right now, right?

Alessio11:02

Oh, yeah.

Jake Cooper11:02

So it's like, it's growing really fast, right? But we've always wanted to have a really, really lean team. Like, we didn't wanna just, like, add headcount for the sake of headcount, uh, just, like, throw bodies at these problems.

We wanna build, like, systems, right? And it's really, really hard to build systems when you're, you're kind of in that expansion phase 'cause you're just adding stuff to the, to the system in general 'cause people are asking for it or things are breaking in general, right?

We basically were like, "All right," like, you know, "We're gonna, we're gonna cut it for now." Like, "We're just... We can't support this, like, these free users that, like, we want." Like, we wanna reach as many people as possible because we believe that, you know, software is this really, really important thing where if you can kind of, like, create something, it's become really difficult to create things in a physical world, so it's really important to make it really easy for people to build things in a virtual world so that people have access to creation, right?

And so we wanna reach as many people as possible, um, but there's kind of, like, legs on that journey. So we basically had to kinda close off the, the free, free kinda users for a little while, rebuild the business, make sure that it worked in general, right?

And then I think you can kinda, like, see the building of that in general, right? And then I think you, you see kinda some- ... yeah, divots in those charts, right? Like, if you actually follow between, I think, 2025 and 26, 26, it's either summer or winter.

That's basically it, right?

Alessio12:07

Mm.

Jake Cooper12:07

Like, where either people go on holidays with their family or they, uh, go on holiday.

Alessio12:10

Oh, it affects that much?

Jake Cooper12:11

Yeah.

Alessio12:12

Damn.

Jake Cooper12:12

Yeah, yeah. Well, because it's like, it's kind of B2C, it's kind of B2B in general, right? And so you have a lot of these users where, like, they're shipping constantly, and then, you know, they'll kind of, like, stop or whatever, right?

And so maybe for summer or, like, maybe... Like, our, our activation curve is, like, now we see a lot of people, like, activating in the weekday, right? Uh, because we have a lot more, like, business users in general.

So that gets a lot less, um, sheer, so to speak, right? And it kinda, like, smooths out over time, you know?

Alessio12:37

Yeah. Is there any point at which you started prioritizing, um, AI developments or agent development?

Jake Cooper12:45

I think, like, we've... So we've prioritized almost, like, agentic as, like, a, um, top-of-funnel thing. Um, and probably over the last, like, six months, we've probably deep- deeply prioritized, like, agentic as a, as a mechanism to go and build and, and deploy things, uh, just because we, we believe fundamentally, like, the, the curve is so sheer, and, like, that is the way that people are gonna go and build and deploy software.

Um, and it, it, it almost, like, fundamentally doesn't matter if it's, if, like, this is dot-com or not because we're all on the internet now anyways, right? Um, and so if agents are gonna go and deploy a bunch of things and we hit an inference wall at some point, then, like, at some point we will go in and fix those problems.

But, like, that will be kind of the dominant species over the next, like, 10 years is, is we've moved from assembly to C to C++ to JavaScript to now, like, words, right? And you're gonna need to be able to close that loop, right?

Um, but that's, that's where it goes, you know? So.

Alessio13:34

But when you say this is dot-com, do you mean, like, buying the domain or-

Jake Cooper13:37

No. No, no, no

Alessio13:38

... a cloud or something?

Jake Cooper13:38

No. Uh, I, I mean, like, actually just, like, you know, they had a bunch of run-up in the dot-com era for companies-

Alessio13:43

Oh, yeah, yeah

Jake Cooper13:43

... because they were like, "The internet is really, really important." And then you hit kind of like bottlenecks, fundamental laws of physics, math didn't work, all of those other things. And, and everybody kinda like, you know, went back down to earth, right?

Um, but at the end of the day, it didn't matter because the internet is, like, so, so impactful for our lives that if you operate on a long enough time horizon, that you should be, like, you should just build these things anyways because you, you can see where that's going, right?

And that's where I fundamentally believe a lot of the agent stuff is, right? And we can talk about it, a little bit of it later, but you're gonna get to a point where you're running thousands of these agents, like, in parallel, right?

Like, one, what's the inference cost for that? What's the compute cost? How do you go and make that efficient? All of those other things. But, like, two, how do you go and coordinate all this stuff? Like we had...

We have we have issues coordinating humans in, in, in general, right? We don't even have good tooling for that. And now we're starting to, to figure out this like, oh, like, how do you get agents to coordinate? How do you go and, and get them to be able to, like, safely version changes or, or like, for them to know when to, like, put their hand up to get somebody to intervene, right?

Um, otherwise it just becomes like a interrupt factory that's, like, crazy, you know?

Bare Metal14:40

Alessio14:40

Well, so s- maybe we'll go right on the technical side of things.

Jake Cooper14:43

Yeah, yeah.

Alessio14:43

What are the core, like, infrastructure or architectural beliefs of Railway that allow you to do what you do?

Jake Cooper14:50

Yeah. I think the, the primitives matter a lot for us. Um, like a lot, a lot. We need to be able to do network, compute and storage, and orchestration all kind of around it. You kinda need control over a lot of those things.

Like, we've talked a lot about, like, how we don't really use Kube, uh, like Kubernetes, because we want the higher order of control to be able to, like, go in and, and place workloads in very, very specific places, right?

Um, and the reason for that is, like, you know, it's kind of the thing we, we talked about previously, but like, you have to be very, very efficient with these agents, like memory reuse, all of those other things, or you're gonna massively, massively blow up your cost structure, right?

I think also incidentally, being able to rack and stack your own servers and, and, uh, build your own, uh, metal, it unlocks a level of like, uh, performance, one. But, like, two, cost, where you can say, "Oh, those experiences that you wanna offer where you're running 1,000 agents in parallel are not, like, massively cost-prohibitive," right?

'Cause, like, if you look at just, like, token use right now or compute use or anything else like that, those things are blowing up massively, right? Over time, those things are gonna have to get a lot and lot more efficient.

You can get a lot of almost, like, back of the napkin, balance sheet margin, whatever you wanna call it, to, to kind of make those experiences, like, solid by building your own metal, right? And so kind of to the earlier point of, like, we've always tried to go a little bit deeper every time to make that experience.

It's all in the service of offering that, uh, that, uh, differentiated experience to as many people as, like, like, humanly possible, you know?

Alessio16:11

Yeah. You have a data center in Singapore.

Jake Cooper16:13

Yeah. So we've two in every other region now. Singapore, we're adding a second one in Q3. So yep.

Alessio16:19

So, so, like, what's it like? I mean, I, I've, I've never built a data center or like-

Jake Cooper16:22

Yeah. Well, we'll have to, like, go to one or whatever

Alessio16:23

... So just go to, like, Equinox and say, "Hey, I want some spots."

Jake Cooper16:26

Yeah, yeah. So, so, yeah, I mean, I can, I can probably-

Alessio16:28

Equinix.

Jake Cooper16:28

Equinix. Uh-

Alessio16:29

Uh, right

Jake Cooper16:30

... yeah. Yeah, Equinox. I mean, you can go to-

Alessio16:31

Equinox for your body-

Jake Cooper16:32

Yeah, yeah, yeah

Alessio16:32

... Equinix for your software.

Jake Cooper16:33

I mean, you can put a, you can put a data center in the steam room and get nice and hot or whatever. Um, but yeah. Yeah, you basically just go and you say, "Hey, listen, I want, uh, power, and I want a cage."

Uh, and they're like, "Great. Here, this is what it's gonna be." Um, and then you rent the cage for a period of time, and then you have to fill the cage, uh, with racks, servers, and then hook up internet to it, right?

Um, that's realistically all it is.

Alessio16:56

And then you handle everything else, right?

Jake Cooper16:58

Yeah. You just handle everything else, right? Um, so-

Alessio16:59

And, and, like, what's the math versus obviously the clouds-

Jake Cooper17:02

Yeah. Uh-

Alessio17:03

... doing it for you

Jake Cooper17:04

... our, our payback period when we go to, to metal, um, if we rented in the cloud, our payback period is about three months.

Alessio17:10

Which is crazy.

Jake Cooper17:11

It's nuts. Yeah. And, and that's like four years' worth of, like, depreciated hardware, right? Um, and so I think it, it, it's like you're gonna see a lot of this almost, like, compute crunch-

Alessio17:20

Yeah, yeah

Jake Cooper17:20

... so to speak, 'cause a lot of the hyperscalers are, are buying up a lot of stuff. Like, we're working directly with OEMs and ret- like, resellers and, and, like, directly with people who are, like, building these machines, like Supermicro, Dell, all those other things, to go in and get these things, things working.

But, uh, you know, up- upstream, there's, like, a bunch of supply stuff. You know, we... It was funny 'cause when we raised our last round, uh, in between basically b- deploying the capital for the servers and, um, actually, I think even, even now, the amount of money that we've raised, uh, is less than the amount of money that we have in the bank, plus how, what the value of the servers are, because the servers have actually appreciated in value 'cause RAM has gone up in, in general.

Right? Um, so it's kinda nuts just in terms of like- How valuable hardware and all of this stuff is, right? If you look at especially a lot of the, like, hyperscalers, like what they deployed, like $80 billion of like c- capital expenditures, like, this year and, and, like, into next, it's gonna be like more in, in general, right?

There's these massive, massive scale like infrastructure build-outs, and you're looking at that being like, "Wow, that's crazy that they're spending like way more than the Manhattan Project." But, like, again, if you go back to every person is going to run, you know, dozens, hundreds, whatever of agents in, in parallel, like-

Swyx18:29

You should spend more than the Manhattan Project

Jake Cooper18:30

... you have no-- you have-- like, you have no conceptual idea of, like, how much compute is required to go in and make that experience happen. Even if you're deeply efficient, even if you're sharing resources, even if you're doing all of these things correctly, and that doesn't even count inference.

Swyx18:43

How do you plan on the build-out? Like, I mean, the growth chart is so vertical that like-

Jake Cooper18:46

Yeah

Swyx18:47

... you know, like, uh, are you usually 100% utilization rate as soon as you're live with these tracks, or like how far ahead are you planning?

Jake Cooper18:53

Yeah, so, so, um, like we still maintain like cloud presence for like bursting essentially. And so what we can do is, um, you know, we work with AWS and GCP and a few of those, you know, other, other clouds.

Like, we can just rent, and then the moment we kinda get space or power or whatever, you almost just like compact-

Swyx19:10

Yeah

Jake Cooper19:10

... those off the cloud, right? Like we... 'Cause we started on, on the clouds, and then we built a system to allow us to migrate to our own metal. And so there's nothing that says you can't just continually do that again, which is exactly what we do right now, right?

And so we, we never wanna be in a spot where essentially we are, um, you know, compute constrained, right?

Swyx19:29

Yeah.

Jake Cooper19:29

Um, and at the start of the year, like we actually got to a point where we were cons- compute constrained because, uh, the one upstream provider that we were actually working with wasn't able to give us quota at the rate that we needed to, and the hardware was like slower, right?

And so we had to do a bunch of different stuff. I spent a weekend rebuilding our entire, like, network-

Swyx19:46

Mm-hmm

Jake Cooper19:46

... like overlay essentially, so that we could straddle, uh, five different clouds.

Swyx19:51

Right.

Jake Cooper19:52

Yeah. Oracle, AWS, ourselves, GCP, and, and like one other one, right? And we can do more than that now, right? But, you know, we got into a spot where, like, we were just trying to, like, pack instances tight 'cause we couldn't get the amount of compute that we needed, right?

Um, and it was really unfortunate 'cause as a result, like some of, uh, we had a few, like, reliability kind of things-

Swyx20:09

Yeah

Jake Cooper20:09

... which are now kind of past us, but it was all a result of this kind of like... There's a, there was a tweet that I made where like, you know, I, I like got in, in trouble 'cause I was trying to point it out, but I accidentally caught the Superbase, uh, folks in the crossfire.

But, like, the tweet was about there's-- It's really, really difficult and it's gonna become more and more difficult to acquire compute at the rate that these models need to acquire compute, right? Um, and we got bit by it, which is-

Swyx20:30

Mm-hmm

Jake Cooper20:31

... you know, fair and reasonable in the karma scheme of, of me, you know, trying to point it out. Um, so yeah.

Swyx20:36

How do you think about pricing knowing that you might not have your own metal available at all time? Like, are you pricing assuming that you'll need to, like, pay yourself extra margins if you had to end up going in the cloud?

Jake Cooper20:47

Because we've built out our, our metal data centers, like our margins on metal are like quite high, for like 70%. Um, and so we can actually s- deeply subsidize the, the cloud business if we wanna scale at a reasonable rate.

And so we have a few different... Like, it's actually very fun from like an operations perspective 'cause you have a few different levers on how you can go and scale it. You have like the metal, which actually like makes your margins.

You have the cloud burst, et cetera. You have debt you can use to, like, buy servers in general. So it's a very interesting, like, operational, like, problem to basically say like, "Okay, we have this much cash." Oh, and then you have obviously venture capital that you can raise on top of it, right?

Um, and so you have this much cash. How much money should we raise? How-- Like, how quickly can we go and deploy it, et cetera, if we can scale revenue as basically as quickly as we can scale compute, provided we continue to make it trivially easy for people to go and build and deploy in that.

Like, the faster you can close this loop and, and the more operationally excellent you are with the capital, like just the faster your business can... It's, it's just a basically like straight linear kind of like-

Swyx21:37

Yeah

Jake Cooper21:38

... deployment rate o- on some of that stuff, you know?

Swyx21:40

I think infra startups raising debt is a tool that people don't utilize enough or know enough about.

Jake Cooper21:47

Oh, my God. Yeah, I-

Swyx21:47

What, tell-- What, what can you tell us about that?

Jake Cooper21:49

Yeah, I mean, well-

Swyx21:50

Like, is it secured against your g- your CPUs or what ?

Jake Cooper21:53

Yeah, it's just, it's just secured against, uh, against our, uh, like, hardware.

Swyx21:56

Yeah.

Jake Cooper21:56

Right? Um, and so-

Swyx21:57

What, what, what rates do you get? Like, who are the lenders?

Jake Cooper21:59

Oh, we just, we just pay like prime at whatever it is, plus like a, like we, we can refinance any of the debt as it goes down. Um, like the terms are pretty good from that perspective. I think like the, the unfortunate thing is like Twitter has no nuance or whatever, so they're like, "Venture debt bad," or whatever.

It's like, well, no, like, as with all things, like most-

Swyx22:15

It's not venture debt.

Jake Cooper22:16

Yeah, yeah, or whatever.

Swyx22:17

It's data center debt.

Jake Cooper22:17

Yeah, it's data center debt, right? But like, yeah, I think there's, there's specific tools and specific areas where you can be very, very deliberate about not just using one specific tool as a hammer, like venture capital as a hammer for everything.

Um, you just have to kind of like go out and explore it and figure out how it kind of like, like works.

Swyx22:33

Yeah. VC's the most expensive financing you can get in a sense.

Jake Cooper22:35

Yeah. Yeah, yeah, I, I think, uh, incidentally, I think also people think about VC completely wrong from a raising capital perspective. Like-

Swyx22:40

Okay, tell us how is VC is wrong .

Jake Cooper22:42

Yeah, well, I, I think most people are like, "Okay, well, how do I raise as much money as possible from, like, whoever is, like, probably the best that I can get at that point in time?" And I think that's, like, kinda close to right, but I think what you should be doing, or at least what we've tried to go in and do, is, like, try and figure out what almost unfair advantage you can buy with that equity because it's the cheapest equity or, or it's the most expensive kind of equity you're gonna give away at that point in time, assuming your company's gonna get better and better and better, and how do you use that to, like, go in and work with somebody who is stellar and who's going to go in and compliment you,

right? Like, you know. Yeah, like Series A-

Swyx23:11

So lucky.

Jake Cooper23:12

Yeah, right. Like, you know, great. I've, I've never started a company. Ray's, Ray's from Lucky. He's got good advice. I can text him all the time, uh, he's really fast, et cetera. Like, awesome, right? Then you kinda, like, move on and, and you kinda like, you know, worked with, uh, you know, John and, uh-

Swyx23:24

Erica

Jake Cooper23:24

... Jordan, uh, Jordan at Unusual, right? Um-

Swyx23:27

Yeah

Jake Cooper23:27

... and they were like, "Yeah, you roughly know what you're doing in building a product. Like, we're just gonna mostly, like, leave you alone and be totally available for advice." Amazing. Awesome. Get to Series A. Business is a total, you know, operational tire fire, right, because we just don't know how to scale a business, right?

Go and work with Erica, and, you know, Jordan's over at, at Redpoint, so bonus. Like, you know, we get to work with them continually, right? And then now moving into, like, you know, a raise from, uh, TQ and FPV, like we're moving into the enterprises now, right, and, like, feeding into there, right?

So every step of the way, we've kind of moved towards, like, who can we partner at, at this specific time who's gonna help us unlock that ex- section of the journey? Because- Guess what? I just, I don't know enterprise sales.

I can roughly, like, eyeball it and be like, "Yeah, as an engineer, I think these are the kind of features that we're gonna roughly go in and need," and we have some wonderful people who are gonna help us internally.

But you really wanna work with those people who, like, at the boardroom dynamic level, are gonna be like, "Oh yeah, we're all aligned, and that's obviously what we wanna go in and do, and we can spend our time basically saying, 'How do we, how do we win this?'

versus, like, bickering about strategy," right?

Alessio24:23

Uh, no, I just had to pull up some beautiful data center charts.

Jake Cooper24:26

Yeah.

Alessio24:27

Uh, I feel like you've done others. I couldn't, I just couldn't find them. But anyway.

Jake Cooper24:30

Well, these are good. I mean, like- ... they all kinda look the same. Like, you just, the servers in a rack, right?

Alessio24:33

Look at our box.

Jake Cooper24:34

Yeah, exactly. This is our box, right?

Alessio24:36

Such a gorgeous box.

Jake Cooper24:36

Yeah, it's like, "Do you wanna see more racks?" It's like, "Oh, yeah." It's like, you know.

Alessio24:39

What's the, what's the Jake Cooper signature edition-

Jake Cooper24:41

Yeah, yeah

Alessio24:41

... uh, box?

Jake Cooper24:42

We have, we, we actually have, uh, we have plans, uh, internally. Uh, yeah. So it'll be fun. Uh, we got a few different promos that we're gonna do, uh, and, like, stunts for, for the year. Uh, so those will be fun.

Alessio24:51

Yeah.

Swyx24:52

You had a tweet about data centers in space just before we wrapped this section.

Jake Cooper24:55

Yes. Yes. Now-

Alessio24:56

Why, why no data centers in space, man?

Swyx24:58

Why no data centers in space?

Jake Cooper24:58

Uh, okay.

Alessio24:59

Why you, why you hate, why you hate so much?

Jake Cooper25:00

Okay, so, so it's not no data centers in space because actually I think, like, my hot take is, like, I think this is solvable. I've just never seen anybody solve it, right? Uh, because you need to, like-

Alessio25:09

No, no, no. What... You, you just s- you said, you said, "How you, how you gonna dissipate that much heat in a vacuum?" You're making a physics claim.

Jake Cooper25:16

Yeah, yeah.

Alessio25:16

But you're not-

Jake Cooper25:16

Well, because, uh, because I haven't seen anybody, like, prove how you're gonna go and dissipate that much heat in a vacuum, right? Like, it, it doesn't mean that it's not possible.

Alessio25:23

I know.

Jake Cooper25:23

It just means that, like, nobody's kind of put it up yet.

Alessio25:26

Astrophage.

Jake Cooper25:26

Pardon?

Alessio25:26

Astrophage.

Jake Cooper25:27

I don't know what that is.

Alessio25:28

The Mar- the Martian thing. Okay. You're very logical.

Jake Cooper25:29

Okay. Yeah, yeah. That's fair. But yeah. I don't know. I mean, it could work in general, right? But I think a lot of people... And I, I think incidentally this is probably what you have to sorta do, um, is like they're putting almost the cart before the horse.

It's like, "Oh yeah, we're gonna put data centers in space." It's like, okay, but how? It's like, well, we have some period of time to basically figure it out, right? It's like, it's like, you know, in The Martian where they're like, "Oh, how are we gonna, like, intercept the meteor?"

Alessio25:49

Yeah, yeah. Astrophage. Yeah.

Jake Cooper25:50

Oh, okay. Right. Um, it's like, how are we gonna do that? It's like, well, we'll figure it out. We have so however long to go in and figure that out, you know? So.

Alessio25:57

Yeah, yeah. Yeah, making, making a bet on, like, human invention is weird because you just have to blind trust it is, can be solved. But-

Jake Cooper26:03

100%, right?

Alessio26:04

I feel like physics i- and there's, like, some first principles bounds that you can put on, like, maybe not.

Jake Cooper26:10

Yeah, I know, right? And, and it's-

Alessio26:11

Like, maybe you're asking to travel time here or, like-

Jake Cooper26:14

Right

Alessio26:14

... break, break some fundamental thermodynamic law.

Jake Cooper26:16

Yeah.

Alessio26:17

Um-

Jake Cooper26:17

And I don't know how VCs do this, like, incidentally too, because it's like how do you know what's, like, basically not possible and, like, is a grift versus, like, uh, is possible but, like, sounds completely insane, right? And you're like, "Oh, cool.

Like, you know, we're gonna put data centers in space." It's like, okay. Coin flip as to whether that's, like, one or the other. You just don't know, I guess. And I guess you'll know in, like, 10 years.

Alessio26:40

Yeah.

Jake Cooper26:40

Cool.

Alessio26:41

Yeah.

Jake Cooper26:41

That's one cycle.

Alessio26:42

Okay. Okay.

Jake Cooper26:43

Yeah. So.

Alessio26:43

Moving back to agents.

Jake Cooper26:44

Yep.

Alessio26:45

I think the branching that you do, the, the fast spin-up and orchestration, it's kind of like the pre-work-

Agent Needs26:45

Jake Cooper26:51

Yep

Alessio26:51

... that, like, happened to be exactly what agents want.

Jake Cooper26:54

Yeah.

Alessio26:54

What do agents want differently than humans?

Jake Cooper26:57

What do agent want differently than humans? I think they want the ability to, uh, version thing. So it's, it's not, like, actually that different. Like, there's just almost, like, slight deviations in terms of how it kinda materializes, right?

So agents want a way to be able to go in and, and test changes incrementally, right? Like, we just, we have feature flags as, like, engineers or whatever, right? Like, is there any reason why they can't just use feature flags, right?

I don't think so. Like, I think there's ways that you can go, just go in and do that, right? They want version control. Is there ways we can use Git or not Git? I think that one's, like, realistically completely up in the air, right?

And I do think that something ultimately outside Git will emerge, uh, in terms of how we're gonna go and version a lot of these things over time. They need observability. You need to be able to go in and essentially query what happened, at what point in time, which steps failed, uh, traces, logs, uh, metrics, all of those other things.

They need, like, network compute and storage. They need the ability to write files, save files, iterate on files, snapshots file system, all of those other things, right? And so I think a lot of the stuff that we roughly needed is, like, very, very kind of in line with a lot of the stuff that, that agents also need, right?

Um, and so, like, the branching and forking stuff, like, it's not different. Like, we're just moving a thousand times quicker than, than we used to, and so some of these things, like, look like you really need, like, something massively, massively different, but it's just you need something massively better than what currently existed, right?

You need orchestration. You need something massively better than Cube, right? You need, like, networking. You need something probably better than Envoy, right? Like, it... And it just goes all the way down the stack essentially, in terms of, well, if the, the workload profile doesn't change so much as it gets, like, massively, massively compressed because you need to do thousands of these things, what assumptions change, right?

Like, etcd is gonna melt, right? Like, you know, you need to replace it with something, right? And then I think you can go all the way down the stack and basically say, "Okay, well, that part has to change, and that part has to change, and that part has to change."

And the interesting thing about the kinda, like, super exponential curve is that you have to build your systems in such a way where you can rip out those parts at any point in time because a new bottleneck might emerge because, you know, you start getting really, really good at, like, parallel agents, right?

Um, and then that's, that's kind of where the new bottleneck is, right? And, and that breaks a different part of your system, right? Um, so I think it's very much, like, similar kind of stuff that, that kind of like, like, humans have needed.

You just need it at a 1000x scale, right?

Swyx29:16

Yeah.

Jake Cooper29:16

So, like, how do you co- how do you do code review in the age of the agents, right? I guess this is more of a question.

Swyx29:20

You throw more agents at it.

Alessio29:21

You don't.

Jake Cooper29:22

Like, yeah, right? Um, but then, like, who, who reviews things for, like, CVEs and, like, all of those other things? You can-

Swyx29:28

More agents.

Jake Cooper29:28

Yeah, more agents. Right. Okay. And, and then that's how we, we hit the inference wall at some point, right? Um, and you can continually throw agents and agents and agents at that problem, right? But, like, you know, I think there's, I think there's a limit to, like, the amount of, of agents you can kinda throw at a problem.

Swyx29:45

You started, though... You already had a CLI before it was cool, I guess. How has the shape-

Jake Cooper29:49

CLIs have always been cool, by the way. Um, but yeah.

Swyx29:52

How has the shape of, like, what you're exposing changed, if at all?

Jake Cooper29:56

Yeah.

Alessio29:56

Yeah.

Jake Cooper29:56

So, so I think, um, the CLI changes because the way that we think about this is like how do you give Claude or Codex or Chat or, like, whatever, like, any of these models almost like- A handhold.

Alessio30:08

Mm-hmm.

Jake Cooper30:09

And like a CLI is a single command when you think about it, right? It's like, okay, well, you're gonna do a deploy or whatever, right? Um, you're gonna get logs, you know, whatever, right? Like, things that were prohibitively annoying to humans are not actually prohibitively annoying to agents.

They're really, really nice, right? And so if, if I wanted to hand you a CLI and I said, "Hey, guess what? This CLI has forty arguments and-

Alessio30:31

Right

Jake Cooper30:31

... six hundred flags," you'd be like, "Wow, that's crazy. Like, I'm never gonna use all, all those things in general," right? But if you hand it to an agent and you say, "Hey, there's forty arguments, this... and six hundred flags," it'll be like, "Oh yeah, this is excellent."

You know-

Alessio30:41

Mm-hmm

Jake Cooper30:41

... like I, I have so many handles that I can go in and kind of like work on with this, right? And so I think incidentally if you're going to go in and try and expose things for agents over, over that mechanism, you wanna just basically have as many handles a-as possible where they can get information, query additional dynamic information, and then see how it can close that loop, like as quickly as possible.

Most of the kind of like problems right now are actually just how do you close the loop as, as quickly as possible. Where does the agent get stuck and, and how can you go and kind of remove that?

That's why incidentally like telemetry is very, very important because if you can tell where the agent gets stuck from the CLI and you say, "Hey, listen, like twelve percent of people are actually getting, uh, deviated from the happy path because of this thing, and now I go and add this arg and that drives it down to two percent," you've massively increased the like rate of the, the loop closing for, for a lot of people in general, right?

So that's kind of the way that we think about not just the CLI, but every point in, in the dashboard, right? Like it is a user journey from I hear about Railway-

Alessio31:38

Mm-hmm

Jake Cooper31:38

... I go and get s-something deployed, I get my first green build, whatever, aha moment, I see an endpoint, I see some logs, I see whatever, and then I go in and, and iterate, right? And then I go in and an iterate loop is indefinite and infinite until the end of time, right?

It's basic like user wants to deploy a new thing, user wants to deploy, deploy a new Postgres instance, user wants to change their, their, their code, user wants to iterate all, all over time, right? And so if you just focus on a lot of those iteration loops and, and figuring out what's blocking that loop from closing as quickly as possible, like one of the things we talk about internally is you never ever, ever wanna be waiting on compute anymore.

You always wanna be waiting-

Alessio32:14

Yeah

Jake Cooper32:14

... on intelligence, right? And if you're waiting on compute, there's a bottleneck that needs to be destroyed there because at some point that bottleneck will be so, so, so large that some other workflow will kind of emerge to go in and, and, and change a lot of that stuff.

And I, I think incidentally like, you know, we've built a really, really awesome product where you can push code, and then you build the code and all those other things, right? But that like push, pull, whatever kind of like loop, I just fundamentally believe it's gonna go away, right?

Like it's... We're gonna get to a point where you make a small change in production, that change is versioned across your entire kind of infrastructure. You're working alongside, you know, copy and write versions of your, your database, all of your infrastructure, and then you merge it in and instantaneously it's, it's like live, right?

'Cause that's like the, the holy grail of, of loops, right? Um, but that like push, pull, rebuild thing, right, is a point of friction that we are like removing entirely from, from our loops.

Alessio33:03

Yeah. It's incredibly fast. So if anyone hasn't tried it, like-

Jake Cooper33:06

Yes.

Alessio33:06

Yeah. Uh, that, that fast feedback is great. You know, my hot take is that, you know, Railway was kind of famous for its canvas, which sort of visualizes your infrastructure and lets you manipulate it visually. But that was for humans.

Jake Cooper33:19

Yeah.

Alessio33:19

And actually now for the next phase in growth, like Railway CLI is more important than canvas, which is what you were famous for.

Jake Cooper33:26

Yeah. So I think the, the canvas is funny because like it's actually just a mechanism to, to show you changes over time. But I think you're totally right in the sense that like we have previously used it a lot as an input, and its goal moving forward is actually a lot more like an output.

Um, what I mean by that is you would go to the canvas, and you'd make some changes and all these other things, whatever, right? Um, and you'd see them and, you know, your agents or, or your, your infrastructure would evolve over time, right?

Now, you just have a bunch of agents that like they have access to CLI, and they can go in and make those changes in general, right? And so the canvas actually instead of becoming this like input thing where you're like, "Oh, cool, like, how do I go in and make this happen?"

It's actually just more of an output thing. It, it basically says what information-

Alessio34:04

It's a dashboard. Yeah.

Jake Cooper34:05

Yeah. What information does the human need at this point in time to make, uh, suitable decisions about, about, um, like control requests of do I approve this, do I not approve this, right? Like that's realistically all the canvas, uh, becomes at that point in general.

Alessio34:18

Yeah.

Jake Cooper34:18

Right? And also a way, and I think this is important, and I think this is, um, lost on a lot of people who are like building some of these like canvas experiences, it has to be almost like a, an anchor for your context.

It has to be like a port in the storm. It has to be like you have to think basically about it as like layers in like a file system almost to like get to the next spot, right? And so you have all your infrastructure and like this is why the canvas starts as like it's just a project, right?

And then you dr- you have a drill down chart, right? Like it's like I'm breaking down into these services or this like section that just is like a function or code or anything else like that because you wanna actually be able to represent the entire thing, not just in your head, but in this, in this canvas so that other people can also get that representation, so that they can think on the same wavelength as you, so that they can move as, as quickly, right?

I think a lot of orgs, especially as they scale, they get in trouble because all that context lives in somebody's head basically. And it's like, "Oh, how does this microservice work?" It's like, "I have no idea. Go, go ask this specific person," right?

And then you have entire categories and classes of products that are built around like how do you do context discovery at, at all these things, and I think a lot of that stuff gets just gets melted in terms of if you can have a really, really solid hierarchy and you can infinitely nest services, infinitely less nest code, infinite less context, in-infinitely nest all these things all the way down, um, that's h- what allows you to kind of build these kind of like structures up over time, you know?

Alessio35:38

Yeah.

Jake Cooper35:38

And I think it's, it's also what's gonna allow us to like build, I've written a bit about this like these like hyperstructures, like things that are way, way bigger and like, you know, you look at the Golden Gate Bridge, and you're like, "How, how did we build that?"

Like, you know, we've-- there's that whole meme of like, "Oh, so how do we build this?" Like, "We lost the technology" or whatever.

Alessio35:53

We don't know how.

Jake Cooper35:53

We don't know how anymore, right? It's like, it's like, well, yeah, I mean, to some extent, yes, because a lot of the coordination that we, that built those things like has evolved, right? Um, and like has changed, and there are new things that we've lost almost like some of the art of like building that structure as we've just like- ...

jammed everything into Slack, right? And we're just like, "Everything happens through Slack," and it's just-

Alessio36:12

But you do everything in Discord, so

Jake Cooper36:13

Yeah, yeah. Well, it's, it's the same point. It, it doesn't, it doesn't really matter. It's just like message passing and interrupts, message passing and interrupts, message passing and interrupts, right? Like-

Alessio36:21

So you're arguing that there should be something better, more structured?

Jake Cooper36:24

Than Slack?

Alessio36:26

Yeah.

Jake Cooper36:26

Yeah.

Alessio36:26

Okay.

Jake Cooper36:27

Oh, for sure. I think Slack, I think, and incidentally, I think Discord's awful, too. Um-

Alessio36:30

This is, this is the equivalent of my mom test, right? Like, what have you done that has your solution to this?

Jake Cooper36:35

So internally, we've built a, a tool called Central Station that allows us to go in and aggregate all the context from all of our users. Um, so every piece of feedback, every piece of customer support, every single thing like that, uh, gets aggregated into, uh, what we call, like, clusters.

If you have a, an incident brewing or, like, anything else like that, now we can go and determine how many users are affected, all of those other things, et cetera, and then we can actually break off a discussion based on that.

And I think a, a lot of that is actually a lot more, more helpful and, and more correct in terms of instead of, like, having just these, like, long-running channels where you're just like, "Which channel should I put this thing in," right?

Like, if you can-

Alessio37:11

Oh, yeah

Jake Cooper37:11

... dynamically aggregate that information and dynamically route it to the right person based on the context, right? We know, we know internally, like, these four people are pretty close on networking, right? And so if we see, like, okay, we've got a networking thing, you can roughly, like, drill it down to, like, those four people, right?

And if you're sitting, like, oh, okay, cool, it's actually with this part, you can just go and, like, look at the, the commits, right? And this is, like, no longer a manual process internally. Like, this is the whole point of why we built...

If you go to, like, uh, Station or help.railway.com, there's a whole reason we built this thing, right? Is because we wanted to figure out how we were gonna go in and scale with, like, a massive, massive, massive amount of leverage to go and aggregate all this feedback, you know?

Alessio37:47

This is built in-house?

Jake Cooper37:48

Yep.

Alessio37:49

Okay. So and then I, I remember, uh, helping out on this one with Angelo-

Jake Cooper37:53

Yeah

Alessio37:54

... uh, in, in 2023. Yeah, you scale a lot with, uh, with a very small team.

Jake Cooper37:58

Yeah. Yeah, yeah. So we're like 10 times bigger now.

Central Station38:00

Alessio38:00

Oh my God, you have your full developer count here?

Jake Cooper38:03

Yeah.

Alessio38:03

Okay. All right. Very well.

Jake Cooper38:05

Yeah. Oh, if you go to, if you go to Rail-

Alessio38:05

I can just, like, cron this and then just have your life.

Jake Cooper38:07

Well, you don't, you don't even have to cron it. We expose this as, like, a, a pub subble with a... So go to railway.com/stats.

Alessio38:13

Oh, there you go. Yeah, that's your board. Yeah.

Jake Cooper38:14

And so it's like all real-time metrics for all of this stuff. There's a way to get this as, like, a JSON too somewhere, um-

Alessio38:19

Yeah, maybe later

Jake Cooper38:20

... if you care or anything else like that.

Alessio38:21

We'll look it up.

Jake Cooper38:22

Yeah. But yeah, yeah, we're, we're big on, like, trying to build everything in, in public, talk about a lot of the stuff that we're working on. You know, like, we've had some issues or whatever in the past, and we're like, "Hey, cool," like, "here's how we're fixing these things."

Like, we've, you know, we've got both compliments as well as some flak for our incidence reports and, like, always trying to, like, make them better over time just to, like, talk with people, right? So...

Alessio38:41

Yeah. Yeah. Um, any, uh, obviously you had a big one recently. Uh, I liked that it was only-

Jake Cooper38:46

Yeah, that was sucks

Alessio38:46

... scoped to 3,000. Uh, you used, presumably used Central Station. Like, any talk, uh, talking through, like, what happened and I guess how do you, how do you address it, you know-

Jake Cooper38:57

Yeah, totally

Alessio38:57

... internally as a team?

Jake Cooper38:58

Yeah. So internally as a team, this one, like, really, really sucked. You know, it was, it was, like, to do with an upstream provider that, uh, didn't... They didn't do the behavior that the s- they said they were documenting, which is unfortunate given they, like, wrote the RFC on how the behavior should work.

But we rolled those things out, and then Central Station kind of caught that, uh, initially, where we had a couple users being like, "Oh," like, "caches aren't invalidating for some of this stuff," right?

Alessio39:24

Right.

Jake Cooper39:24

Um, and so turned it off immediately, et cetera, right? Um, but when you go and kind of roll out to those, like, that, like, large user base of, like, 3 million people, right, you know, like, you have a lot of different disparate behaviors that, that can kind of come up, right?

And so try as we will, we tested those things in, you know, staging. Uh, we have tests for them, uh, like all of this other stuff. Um, you know, and unfortunately, we, like, hit a kind of an, an edge case there, right?

And we've incidentally s- like, gone and hardened a lot of those systems, and now we can, like, make a lot of that stuff better. But, um, yeah, it was, it was a, it was a tough one unfortunately.

Alessio39:59

Yeah. I always wonder how, like, the private disclosures are supposed to work. If people find an, an, an issue, are they supposed to, like, contact you first? Is there... Like, when you run a platform, these things are going to happen.

Jake Cooper40:12

Yeah.

Alessio40:12

And what channels should people pursue to, like, you know, quietly resolve it before it becomes, like, a, a much bigger incident?

Jake Cooper40:19

Yeah. So I, I think there's, like... Uh, there's responsible disclosure. Um, we kind of err on the side of, like, we'd rather over-disclose and know that you know that something is wrong versus almost, like, having your provider gaslight you.

Um, and so yeah. Um, you know, we've, we've, we've kind of ov- we've erred on the side of, like, sharing those things kind of more publicly, even if they go and, uh, impact a small s- like, subset of, of those users, right?

Um, and that's kind of just a decision that we've made internally. It's, it's under, like, we have four values. One of them is honor. Um, and so, like, what's the honorable thing to go in and do? It's like, well, you notify, you notify people, you know, to the widest degree at, at which they may have been, uh, you know, affected or if there was an issue or whatever, and then we kind of confront that head-on and be like, "Why did that happen?

What can we do better in the future?" All of those things kind of like that, you know? So...

Alessio41:05

Yeah. Not the whole user base.

Jake Cooper41:07

No.

Alessio41:07

And that's because of, like, incremental rollouts and other things.

Jake Cooper41:11

Yeah, progressive rollouts-

Alessio41:12

Yeah

Jake Cooper41:12

... um, and, and stuff like that, right? Um, so yeah.

Alessio41:14

Interesting. Yeah, yeah. I feel like that should just be the norm at all large platforms, right? Like-

Jake Cooper41:19

Oh, it totally should.

Alessio41:19

Yeah.

Jake Cooper41:19

And, and a variety of-

Alessio41:20

Wish you did this at Google.

Jake Cooper41:20

Yeah, in a variety of companies it totally is, right? Like, what, uh, there's that whole quote of, like, Meta runs, like, 10,000 versions of different versions of, of Meta in general. And, like, to our earlier point about agents, right, like, they need the same thing.

They need to build the shadow traffic. They need to build all these other... I think we've built so much ceremony around, like, production is sacred, all of the, all of these other things, that, like, we need to get to a point where it's just trivially easy to test different behaviors, right, in a safe environment.

'Cause then you can make those mistakes in, in an environment that's, that's, like, safe in, in general, right? So...

Alessio41:50

You mentioned somebody brought it up. Do you see a world in which these things get automatically caught, not necessarily by your agent, but, like, your customer agent? You know what I mean? That, like, the, the cache invalidation thing seems like a pretty easy thing to, to check if you know to look for it.

Jake Cooper42:04

It's hard because then you almost need... Well, for, for us to, like, determine it, like, we need almost, um- We'd have to hook in with, like, your observability infrastructure in general, right? This is, like, why we almost have the template loop on the platform, is to be able to kind of roll those things out progressively, where you say, "Hey, listen, you know, I can roll this out to, like, Johnny vibe coder initially," right?

Or I can push a shard, and you can almost, like, uh, consume that at your own leisure and be like, "Oh, okay, I'm gonna update to this specific version," right? Or have this kind of, like, roll out over a period of weeks where you're pushing a new version, and then it goes to, you know, 0.1% of people, 1% of people early do- like, whatever, and then rolls out all the way there, right?

That's the kinda, like, non-deterministic version control that, that we've kinda, like, talked about earlier. So yeah, 100%, right? And I do believe that, like, that's where most things should go, go towards because I think ultimately most companies end up building that stage rollout system in, in-house, right?

Um, and it's just the same thing built again and again and again at every single one of these, these different companies. So there's a massive opportunity to consolidate a lot of, like, developers' debt.

Swyx43:05

You, you should have a free tier, like the model providers give you free tokens if you let them use the data. Like, we'll give you free compute if you're like the number one shard that goes out, and you let us plug into your observability.

Jake Cooper43:15

Yeah. Like incidentally we, we do that, right? And that's why the, the-

Swyx43:18

Two packs

Jake Cooper43:18

... you know, we, we talked about... Yeah, we talked about, you know, the, the impact of that on, like, 3,000 people or whatever. We start with the, the kinda lower impact people.

Swyx43:25

Mm-hmm.

Jake Cooper43:25

Like the, you know, larger companies, et cetera, on the platform, right? Like, they're the last ultimately that should receive those kind of rollouts-

Swyx43:31

Mm

Jake Cooper43:31

... so that they have a version of the platform that's like deeply, deeply stable, right?

SRE Agents43:36

Swyx43:36

I have three services, so I'm sure I get, I get the first rollout. You can nuke, nuke, you can nuke my thing at any time, man. Uh, I guess my other question is like there's all these, like, SRE agent companies.

There's like the observability people also wanna have agents that fix your upstream problems. How do you kinda... You have your own agent in the canvas now-

Jake Cooper43:56

Yeah

Swyx43:56

... that you can try with. How do you kinda see that play out?

Jake Cooper43:59

It's almost like the stacking entropy thing in general, right? Like, I think if you don't have the primitives to make iterating in production safe, it becomes very, very difficult, right? And so if you're an observability provider, and you're like, "Oh, here's this fix to this, this error," right?

Like, assume like 80% of those, they're probably actually good. Like, they're, they're gonna make, they're gonna make sense, et cetera, right? But then the last, like, 20% of that long tail of like-

Swyx44:23

Mm

Jake Cooper44:23

... kind of complex kind of issues in general, ultimately rolling those changes out, if you just kind of let somebody say like, "Oh, cool, this looks good," and just like stamp it, there's an opportunity for you to have an issue or an incident or anything else like that.

And I think that's why it's really, really important to have those kinda like forked environments in general, and people have staging, et cetera, but it always end ups, ends up like deviating from prod, right? And so you need the primitives and the workflows and the experience like, like built in our, in our mind on like a, as a first party-

Swyx44:53

Right

Jake Cooper44:53

... thing on the platform, so that you can fork any point at any service at any point in time so that you can almost like... You know, I think I consider the canvas almost as like a little like p- sheet of transparency paper-

Swyx45:03

Mm-hmm

Jake Cooper45:03

... and the agent is kind of like this little guy that you push up, and it's like it should be able to like- ... pop up in the canvas, and it should be like, "Oh, cool, like, well, I need to copy that service.

I need to copy that service, so I can test these two things," right?

Swyx45:12

Mm-hmm.

Jake Cooper45:12

That's my hypothesis as, like, an agent or whatever. Um, okay, cool, I can go in and do that. Looks good for all the, this stuff. Ideally, I get a read-only copy, uh, of, of production. Anything that's PII, et cetera, is kinda like marked, um, as, as like a transform when we, when we automatically clone that database or go for a copy on write version of it or, or read from it, and it just makes those changes.

It says, "Does, does this actually work?" Right? Like, as close to production as, as possible, right? Because ultimately, that's how close you have to be, or you just have a massive amount of drift where, oh, I've changed this thing and, and then it just kinda gets out of sort, right?

The, the system gets a lot more unstable. And, and I think that that's like what you see with a lot of these kind of almost massive systems that these companies built on top of, like, Docker for local and then like Cube for production and, like, this specific thing for whatever, right?

It's like all of that complexity ends up getting to a point where it slows down the developers, yes, but it just, it just gets to a point where it's so unstable at scale that it becomes hard for people to go and iterate and make-

Swyx46:11

Right

Jake Cooper46:11

... those changes, right? And so we wanna compress a lot of that stuff way down and just say like as, like, as close to prod as you could possibly be, that's where we wanna be, right?

Swyx46:21

I was texting, uh, Erica, and sh- uh, for, for questions, and she, she says actually you were originally not a believer in AISRE.

Jake Cooper46:28

Oh, yeah, yeah. Uh, I mean, I've, I've kind of-

Swyx46:30

Have you come around on it?

Jake Cooper46:31

Yeah. Well, I flipped. I'm actually still not a believer on the AISRE because I believe that you need, you need the primitives to make those things safe. And if you just unleash an AISRE on your production infrastructure, and you don't have, like, safe primitives for, like, copying volumes, making sure that this is fine, it's gonna nuke your production database.

Like, it's, it's not a matter of if, it's a matter of when it's going to nuke that database, right? I'm a big believer in, in making those kinda like loops safe in general. I think I was a pretty deep, like almost...

I don't wanna say AI skeptic until like 2023 and then 2024 I've kind of like was like, "Okay, like, maybe I can make this thing roughly do it," et cetera. 2025 I was like, "Okay, now I can, like, hold this," et cetera.

And then, like, over the whole Christmas break, I think you just saw like, I guess winter break, uh, but you s- just massive, like everybody come back, came back and be like, "Oh, my God, it's, it's almost impossible to hold this thing."

Swyx47:22

It was serious you on the cloud docs?

Jake Cooper47:23

Yeah.

Swyx47:24

Uh, cloud bot.

Jake Cooper47:24

Yeah.

Swyx47:25

Well, open cloud.

Jake Cooper47:26

But it's gotten to a point where it's, it's almost like it's harder to hold it wrong than it is to hold it right, you know? And it's like, you know, the, there's that scene, uh, in, like, Avengers or whatever, where Vision's like, "It's terribly well-balanced," you know?

Like, when he picks up Thor's hammer or whatever.

Swyx47:39

Okay.

Jake Cooper47:40

Like, like, damn, like, this, this thing just kind of like self-balances and, like, works quite well from that perspective. So yeah, I'm, I'm a deep believer at this point in terms of that will be the dominant species, right?

Again, you know, assembly, C, C++, JavaScript, words, right?

Swyx47:56

Yeah. It feels like a big jump.

Jake Cooper47:57

Yeah, it feels like a big jump. And, and it is too, right? Um, and I, I think, like, there's, it's not like you abandon, uh, like CPU-based discrete logic in general and just move straight to fuzzy logic. You, you need both, right?

So your skills should call code or applications or, like, whatever, some sort of like static structure, and you can use the skills to kind of distill, um- What the almost like procedure should be, or like how the code should act, right?

Alessio48:22

Yeah.

Jake Cooper48:23

I'm kind of coming to this thesis, which is you need three points essentially, which is you need a s- a clear spec of what, what defines the system, you need the code, and then you need the tests, right?

And I think when you say this thesis out loud, it's like, well, if you've been in engineering for any amount of time, you're like, well, no sh- like, yeah, of course you... Like that's, that's a RFC, like a request for comment.

That's tests, and that's your code, right? But they all matter a lot, and having them all be actually together so that they can reinforce each other and say, "Well, the spec and the tests match, but the code doesn't.

Let me reconcile that."

Alessio48:54

Yeah, definitely. Yeah.

Jake Cooper48:54

"Oh, okay. Now, the tests and the spec match. Let me go and reconcile this other thing," right? And then you can kind of move through that period of, of basically saying, "Well, this is fuzzy, and these two are, are either discreet in the th- the case of tests, or slightly fuzzy, slightly discreet in the t- in case of code," right?

And that's kind of your iteration loop. I think that's also incidentally why you're seeing a lot of people be like, "Software factories, and I wanna write this doc, and like have it go and reconcile all this other stuff," which I think is a bit of architectural astronomy if you like don't actually go in and implement it.

But I do think generally that's kind of that loop is kind of where most things are gonna ultimately end up.

Alessio49:28

Yeah. Uh, for listeners, we've been talking about this on the pod for three years, the, the holy trinity of specs and tests.

Jake Cooper49:33

Oh, okay, cool.

Alessio49:33

Uh, Itamar Friedman from, uh, Kodo is, is the, is the reference where people wanna look it up.

Jake Cooper49:37

Nice.

Alessio49:38

One thing I, one thing I do wanna mention on just on the OpenClou thing is, uh, also like the i- idea that you can self-modify-

Jake Cooper49:44

Yeah

Alessio49:45

... uh, which is kind of interesting. I don't know how exactly Railway would support it, but I do have my OpenClou, uh, and I just tell it that it has the Railway CLI, it can do whatever. Um, and in theory, you can just, whatever capabilities or new infra you need, you can just call the Railway CLI, provision it, and add it to itself.

And so th- the agent h- can, like, modify its own infra-

Jake Cooper50:05

Yeah

Alessio50:05

... which I think is-

Jake Cooper50:06

Yeah, it's, it's not... We have a, we have a loop that I've kind of, uh, set up, which is you put the Rail- Railway CLI on top of, uh, something that runs on top of Railway.

Alessio50:15

Mm-hmm.

Jake Cooper50:15

Right? And so you're essentially authenticated as whatever the current box is, uh, in, in general, and you can make any sort of changes to it. And then you just call Railway deploy, and it deploys itself.

Alessio50:25

Yeah.

Jake Cooper50:25

Right? Like, it's just like, "Oh, cool, I need to go and spin up this instance of this environment." "I already exist in this environment. Excellent. I've got access to a Postgres instance now," right? Like, and this is kind of where we wanna go with a lot of the, like, agentic almost, like, self-replicating, like, infrastructure, is like that's your loop.

Like, you iterate in production. That's your loop, right? And you're gonna just continue to, to make some sort of change, and either it will work, and you're gonna wanna go in and merge it and say, "Cool, that's great," like put it into our upstream, or it will not work, and you can just kind of throw it away, et cetera, right?

How do you go and make those throwaway copies, like, as trivial as possible to spin up, run super cheap, et cetera? I think the, the era of like, I have an AWS instance and I'm gonna, you know, get four vCPU and 16 gigs of RAM, it's gonna get, like, completely destroyed, right?

Because it's like if you do that for, for agents or anything else like that, you now need 1,000 of those m- machines, right? Like, it's so prohibitively cost expensive versus like, you know, we've spent a ton of time trying to figure out how do we go in and make these deploys, whatever you wanna call them, uh, you know, Cloudflare's got the, like, isolates.

Everybody's like cloud sandbox, like whatever. That, like, that atomic unit of deploy on- like, only pay for what you use, spin up instantaneously, close loop as quickly as possible.

Alessio51:35

Yeah.

Jake Cooper51:35

Right? Because if the, if the system can self-replicate the system, and it can do so safely and say, "This is my environment, I'm making these changes," et cetera, it can come back with, "Hey, does this look good? Like, this is a new state of infrastructure given this prompt.

I think I've solved this problem," right?

Alessio51:52

Yeah.

Jake Cooper51:52

And then you can go back to the agent and say, "Actually, like, looks a little bit different," goes and does the loop again, and you're like, "Cool. Excellent. Apply."

Alessio51:59

Yeah. Um, I think, I think that's retroactively obvious-

Jake Cooper52:02

Retroactively obvious

Alessio52:03

... kind of, kind of, kind of like the, the most, uh, uh, useful kind. I don't know. Any other comments on like just, like, agent deployment on, uh, Railway?

Jake Cooper52:11

No. I mean, it's, it's getting better every day and, you know, I'm on, I'm on, uh, X or Twitter or whatever you wanna call it, and like, you can, you can always, you know, yell at me about the experience-

Alessio52:20

Oh

Jake Cooper52:20

... not working, uh, as well as it, it should because there's plenty of things that should work way, way better.

Alessio52:24

I was gonna say, I think at this, at, we're at this stage and juncture when people want like the massively or embarrassingly parallel s- compute, they usually talk serverless. And I feel like there's a n- new serverless that has emerged compared to like the last, the previous five years of serverless.

Jake Cooper52:41

Yep.

Serverless Shift52:41

Alessio52:42

You're kind of like in that, in that new bucket.

Jake Cooper52:44

Yep.

Alessio52:44

Uh, kind of peop- I, I don't know if you have like comparisons or philosophical differences that you wanna call out.

Jake Cooper52:51

No, I think it's like, it's, as you kinda mentioned, it's somewhere in between, right?

Alessio52:55

Yeah.

Jake Cooper52:55

It's like the ability to run stateful, long-running, like you wanna call them workflows, you wanna call them executions, you wanna call them whatever.

Alessio53:02

Yeah.

Jake Cooper53:02

Um-

Alessio53:03

Which like Vercel has Fluid Compute-

Jake Cooper53:06

Yeah

Alessio53:06

... and then Cloudflare has some container thing.

Jake Cooper53:08

Yep.

Alessio53:08

Uh, Google has always had, um, the App Runner and-

Jake Cooper53:11

App Runner and um-

Alessio53:13

The, the new one

Jake Cooper53:14

Yeah.

Alessio53:14

I forget the-

Jake Cooper53:14

There's like a bunch of them.

Alessio53:15

Yeah.

Jake Cooper53:15

Um, yeah. Yeah, I think like that's kind of where everything roughly it, it... And this is why we've been working on it for the last like six years, is like we just believe, like you do need access to a computer.

You'd like a l- a box that speaks Linux, right? Um, so that you can deploy the things that you wanna go in and deploy on it, right? Like other things are going to, I mean, they're gonna change the, uh, almost like surface area of what you can kind of go in and build and, and for us, we're always like, no, like users need a computer, and they need, they need to be able to deploy anything that they truly want, right?

And that's why we focused on long time, uh, for a long time on tho- those primitives, right? Of like network, compute, and storage, right? 'Cause if we can give you those things, and we can expose them, them to you and allow you to run these things indefinitely, right?

Um, that's of course like w- where we believe that, that it's gonna go in general, right? And so I think you're seeing right now where again, the, the whole like Twitter has no nuance versus everybody's just like, "Serverless," right?

Serverless is like, no, it's like it's always, it, it's always somewhere in the middle. You know? Like it's always some sort of convergence of, well, I wanna run it for, for a long time, but also I don't wanna like provision this resource statically or pay for just things that I'm not using or anything else like that.

And that's always been our thesis from like day one. It's like pay only for what, what you use, run it indefinitely. It is just like full, full Linux, basically.

Alessio54:32

Yeah. I think that's why I like, I like them for still naming a fluid. It's like, it's, well, it's fluid. It's flexible.

Jake Cooper54:38

Yeah.

Alessio54:39

Um, another milestone, and then I wanted to ask one more technical question, uh, which is the Heroku official deprecation or what, what did they, uh... B- basically, you know, you're, you're one of-

Jake Cooper54:49

Yeah

Alessio54:49

... the presumptive new Herokus. New Heroku has been a category for, like, as long as I've been in developer tooling.

Jake Cooper54:54

Yeah, right.

Alessio54:55

It's finally happening.

Jake Cooper54:56

Yeah.

Alessio54:56

What was that like when, you know... Is there any behind the scenes of like, "Well, this is the moment"?

Jake Cooper55:02

Yeah. I mean, you just have, you have so many people just like, you're just like, "Like, you were running stuff on here? Like, you as this company?" You're like, "It's crazy that like you," whatever, like, name that you would know is running this thing, and then you're coming to us to be like, "Yeah, we kind of like want to like move a lot of this stuff off," or whatever.

And like, "Oh, okay, cool." Um, but yeah, it's kind of just nuts. Like, I think-

Alessio55:20

Any insights behind the scenes on what's, what is, why, why did Salesforce let Heroku kind of just stagnate?

Jake Cooper55:26

Well, I mean, I can only, I can only like m- ... guess, I guess, right? Like, I mean, I think it's just hard when, like, it's not your business. Like-

Alessio55:34

Yeah

Jake Cooper55:34

... the business of Salesforce is to build a really, really good CRM.

Alessio55:38

Yeah.

Jake Cooper55:38

You know? Right? And like that's their focus, right? They should be really, really focused on building a really, really great CRM. And then you acquire this business as a compute business that's kind of an offshoot of your, your business in general, right?

And then I think, like, you know, a lot of the early Meta folks have talked a lot about, like, focus, right? And like, o- I think Boz has a whole, like, writeup that he's done basically where he, he talks about, uh, in the early days of Meta, we had no money and like we were forced to get focused, right?

And then we basically turned on the money... This is all like, you know me, um, verbatim reph-

Alessio56:05

Rephrasing. Oh

Jake Cooper56:05

... yeah, rephrasing or whatever. We turned on the money tree, and then we had no reason to like not, like, have focus, 'cause we just had infinite money where we could go and split all of our focus, right?

But that ends up diluting your product. It ends up, like, making these things where you kind of have these offshoots, where you're just like, "Is that the focus of the business?" Right? And it ultimately ends up not being if it's not the core of your business, right?

And so to me, it's, it's like kind of no wonder that like it languished in general, right? Um, because it just wasn't the core focus of the business. And I think that a lot of companies get in trouble, um, with this when they kind of like split out their focus in general, because it, it means that you're almost like fighting a like multi-fronted war trying to like compete with all these things and, and not just like compete with them externally, but compete with them internally for like alignment and like where are we going?

What are we doing? What is our purpose here, right? Like, if you're, you know, if you're really, really, you know, like, c- Salesforce built and you're like, "Hey, listen. I love Salesforce and I really wanna like work on all those things," like, you know, and you're, you're mission driven, which is like the, the as- aspiration for a, a company in general of like why do people work on things, right?

Like, it's like they wanna work on something interesting, right? Like Heroku is off to the side.

Alessio57:11

Yeah.

Jake Cooper57:11

It's like it's not the core of, of the business, right? And so to get those resourcing, uh, you know, like budget or focus or alignment or whatever internally, it's, it's just pushed away, right? So it was, it was literally just a matter of time for, for it to happen, uh, in, in our mind, right?

Alessio57:27

Yeah. I mean, kudos for them to like actually call it out instead of just letting it be unknown or-

Jake Cooper57:32

Yeah

Alessio57:32

... maybe, yeah.

Jake Cooper57:32

Well, their whole, their whole release was a little bit odd because they like, you know, they, they kinda called it out, but they didn't.

Alessio57:37

They, they did the our, our incredible journey.

Jake Cooper57:39

Yeah.

Alessio57:39

Right?

Jake Cooper57:40

Yeah. Yeah.

Alessio57:40

They didn't say they were like shutting it down, but they're like, "Yeah."

Jake Cooper57:43

Yeah, yeah. So yeah. And, and then, you know, behind the scenes I think they issued some, some stuff to people being like, "Hey, yeah, you should like close these accounts down." "Like we are going to go in and deprecate this and like remove it every time."

So yeah. I mean, it's just like... A- and it's crazy because like some of my first deployment experiences were like on Heroku. Like, like-

Alessio58:00

I learned like-

Jake Cooper58:01

Exactly.

Alessio58:01

Yeah, I learned to code on Heroku, man.

Jake Cooper58:02

It's, it's like a foundational thing where it's like-

Alessio58:04

I had this fucking alias in my bash for like Heroku deployments.

Jake Cooper58:07

Yeah, right? Like, you, you start with like dragging stuff into an FTP server and then like you move on to like trying to get a deploy working. You're like, "How do I go in and make this happen?" And then it's like, Heroku, right?

Alessio58:16

Did you know about Heroku Packs and the, the other stuff?

Jake Cooper58:17

Yeah, exactly, right? Like, and you, and you learn about all this and it was the on-ramp for us, right? Um, you know, but the wheel turns regardless, right? Like there's, there's new stuff that's emerging and like-

Alessio58:26

Yeah

Jake Cooper58:26

... we're very, very happy to like almost like continue to like, you know, carry the torch on for a lot of that stuff. But we-

Alessio58:31

Yeah

Jake Cooper58:31

... we don't wanna be the new Heroku. We wanna be the way in which people are building and deploying software, and ultimately the way that people monetize software over time, right? Um-

Alessio58:39

Yeah.

Jake Cooper58:40

So-

Alessio58:40

I mean, still it's a big crown to be a new Heroku. Like there's like-

Jake Cooper58:42

Yeah

Alessio58:42

... 50 companies that fought for it.

Jake Cooper58:43

Oh, yeah. Ev- everybody's kind of like, you know- ... holding some portion of this-

Alessio58:46

Yeah

Jake Cooper58:46

... being like, "Ah," you know? Um, but yeah. I think, you know, for us, we're, we're just happy to go in and support people, companies, et cetera. Um, the platform works a bit differently, so it's like-

Alessio58:55

Yeah

Jake Cooper58:55

... you know, it's, it's obviously kind of almost this, the similar kind of like game loop of-

Alessio59:00

CIC life cycle

Jake Cooper59:00

... these like... Yeah, exactly.

Alessio59:01

Yeah.

Jake Cooper59:01

Right? Um, but we've been quite dogmatic in terms of where we believe these things are gonna go in terms of, uh, primitives, you know, the, the agents, uh, kind of fan-off, all of those other things, right? And so some things will fit, and then some things will...

You know, you have to change a few other workflows, et cetera. Like we don't have, uh... And what's that feature that people really love? Pipelines? From Heroku?

Alessio59:21

Uh, Heroku?

Jake Cooper59:21

Yeah.

Alessio59:22

Okay.

Jake Cooper59:22

Right? Um, like there was, we have some approximation of it with the environment system in general, right? But yeah. So it's, it's been super exciting. We've got a ton of people that we're able to go and support, so...

And it's growing a lot. Um, so.

Alessio59:32

Yeah. Any other technical... I, I have one more-

Jake Cooper59:35

Maybe

Alessio59:35

... about, uh, Temporal.

Jake Cooper59:36

Yeah.

Temporal Deep Dive59:36

Alessio59:36

Okay. So Temporal. I have sold my shares. Uh, you are a power user. You're, you're one of our earliest customers I think.

Jake Cooper59:44

Yeah.

Alessio59:44

I met you through Temporal or something. You're a big Temporal business, like your business is built on Temporal. You have, uh, complaints. I think this is the neutral, most neutral, most informed our, um, conversation that anyone will ever hear about Temporal without someone working at the company.

Jake Cooper1:00:00

Mm. Yeah, that's fair.

Alessio1:00:01

Because it's the two of us.

Jake Cooper1:00:02

Yeah, yeah.

Alessio1:00:03

So-

Jake Cooper1:00:03

No, I think that's fair.

Alessio1:00:05

What's your-

Jake Cooper1:00:05

Yeah, so- ... I, I have used Temporal for almost like 10 years now, right? Because like Cadence-

Alessio1:00:10

Yeah, Uber

Jake Cooper1:00:11

... Uber, all of those other things like that.

Alessio1:00:12

And like at, uh... Just, just give people a scale of what, what Cadence is at Uber. Like-

Jake Cooper1:00:16

Yeah

Alessio1:00:16

... what people don't know.

Jake Cooper1:00:17

Yeah, so, so Cadence was the precursor to Temporal, and it powers like all of the trip actions, the rides, the like, you know, when you, you like rent a jump bike or a scooter or like any- anything else like that, or a car.

It's like you're running these workflows for a period of time, and you're basically saying, "This ride will run for an indefinite period until it like finishes," right? And you can go and attach information, whether it's like, oh, you paused it in this zone and so, you know, you need to add this dollar charge to like the, the bill or anything else like that.

And then when you end the trip, like your workflow's done, right? That whole experience, uh, behind the scenes, I don't know about today in general, but it was like powered by, by Cadence-

Alessio1:00:52

Yeah

Jake Cooper1:00:52

... at that point in time. And so it's a really, really like-

Alessio1:00:54

Yeah, and I used to say like it's like imagine if you could program the entire user journey top-down as one function.

Jake Cooper1:00:59

Yeah. Yeah. And it's, it's such a- ... it's such a powerful idea, and it's so, so important. It's also incidentally so important for the next, uh, phase of the agentic journey, where, like-

Alessio1:01:10

Hmm

Jake Cooper1:01:10

... you want an agent to do a specific task, and then you want it to, like, be complete or incomplete on that task, and then move on to the next thing.

Alessio1:01:17

Yeah.

Jake Cooper1:01:17

Right? Like, you need a way to be able to go in and manage these workflows. You need a way to be able to go and manage these workflows dynamically. And I think for me, Temporal was always, like, really, really, really great in theory, and it was really, really great when you got it working the way that you want it to in production.

It's just it required you to, like, model that entire journey in your head, and if you didn't have the entire journey in your head, you could put yourself into a spot where you would cause, like, issues where, like, replaying the state of the entire workflow, like, causes, like, a non-determinism issue.

Alessio1:01:46

Yeah, because it works on, like, deterministic workflow history.

Jake Cooper1:01:48

Yeah, exactly. Right? Um, and so it's very, very easy. It's like, uh, the, the way that I kind of like would describe it is like, well, it's a jet engine, right? Like, if you know how to, like, go in and operate it, if you know how to go in and run it, all of those other things, right?

But you can't hand it to people who are trying to build things that end up being complicated, but don't have that whole kind of like state in their head, right? Um, so if you have a large... Like, we run our whole deployment pipeline on, on top of it, right?

And so that's, like, a reasonably complicated workflow, right? Like, there's, uh, pre-commit hooks, there's, like, signaling, there's queuing, there's, like, all of this other stuff in, in general, right? And we kinda ran into the same thing at Uber, where, like, as you tried to express this large workflow, as you mentioned, like, going all the way down, it got more and more complicated, and it got more and more states in the state machine that you had to, like, map the state machine back to, like, the, the workflow.

Alessio1:02:35

It's a lot of ifs, right?

Jake Cooper1:02:36

Yeah, exactly, right?

Alessio1:02:36

If this, if that.

Jake Cooper1:02:37

Yeah. And so, so at Uber we built a system for, you know, doing the, the state machine and, like, testing the state machine and all that other stuff, and we've started to, like, go and build some of those things here, 'cause, like, it's, it's grown, uh, you know, quite heavily, right?

But it's like, it's such a, like, you know, I don't wanna say love-hate relationship, 'cause that's, like, too broad in, in general. Like, it's f- it was, when it works really, really well, it works, like, super, super well, right?

But then you run into a spot where you just, like, somebody who hasn't interacted with the system or doesn't have the full context of the system goes and puts something in the system that invalidates some of the state or causes a non-determin-issue, it's non-determinism issue or, um, spins off a ton of activities or anything else like that.

And then you have to kinda keep track of, like, almost underlying SRE knobs of, like, "Oh, we have, uh, you know, the amount of activity slots in, in this thing," right? And it's like, well, these should just scale with, like, memory, vCPU, all of those other things in general, right?

So it ends up becoming a bit of a bear to, to kinda scale out in general.

Alessio1:03:31

Hmm, yeah.

Jake Cooper1:03:31

You know?

Alessio1:03:31

So you need, like, a very capable CIS admin-

Jake Cooper1:03:34

Yeah

Alessio1:03:34

... running things behind the scenes for you.

Jake Cooper1:03:35

Yeah, yeah.

Alessio1:03:36

If you were to move off, what, what do you do?

Jake Cooper1:03:39

I think we would build our own workflow, uh-

Alessio1:03:41

Engine? Yeah

Jake Cooper1:03:41

... uh, we have a few internally, um, that, that we've kind of, like, worked on. So yeah, because it's like... Yeah.

Alessio1:03:47

This, this is one of those things where, like, you know, uh, this is one of those classes of things where, like, you, you typically wouldn't vibe code it, but I'm wondering if you can-

Jake Cooper1:03:54

Well, I don't think you should vibe code it still.

Alessio1:03:55

Yeah.

Jake Cooper1:03:55

Like, you still wanna run, like, uh, devson tests and stuff like that, like, to make sure that, like, you-

Alessio1:04:00

Yeah. I, I mean, then, you know, like, it's not like Temporal had to invent that from scratch either, right?

Jake Cooper1:04:05

No.

Alessio1:04:05

Like, so, like, there's libraries for those things-

Jake Cooper1:04:08

Yeah

Alessio1:04:08

... that you can, that you can run. And, like, on, on top of that, it's just a state machine and, you know, that you, that you have to really map out. But, uh, ultimately, you define those abstractions that you want, and you run it through a state machine, and that's it.

Jake Cooper1:04:19

Yep.

Alessio1:04:19

Like

Jake Cooper1:04:20

Yeah. It's, it's very, very doable. Um, so yeah, I think the workflow stuff is very, very interesting. Like, there's a few really cool companies that I think, like Restate's doing some neat stuff here.

Alessio1:04:29

Yes.

Jake Cooper1:04:30

Um...

Alessio1:04:30

So you're very tied into JavaScript. You're, like, a JavaScript maxi.

Jake Cooper1:04:33

Internally, we have JavaScript, we have, uh, or we have TypeScript, we have Rust, and we have Go. Those are the three languages, right?

Alessio1:04:39

Yeah.

Jake Cooper1:04:39

Like, we don't add any more stuff. Actually, that's not true. We have a little bit of C 'cause we write BPF code and, and, like-

Alessio1:04:44

Ooh

Jake Cooper1:04:44

... and, and it's hooks and stuff like that. So, um, but those are, those are the, the kind of, um, languages-

Alessio1:04:48

Is this for this, like, the side container things, sidecar stuff?

Jake Cooper1:04:52

No. Well, so this is for, uh, the networking stack-

Alessio1:04:55

Yeah

Jake Cooper1:04:55

... uh, as well as the, uh, volumes and, and stuff like that. Um, so yeah. But it's like... Yeah, we, we used the, like, the TypeScript stuff a lot-

Alessio1:05:04

Okay

Jake Cooper1:05:04

... because it, it's like what powers the dashboard, but we're gonna move a lot of the kind of workflow stuff off of the kind of dashboard stack into actually the infrastructure stack we're... just recently.

Alessio1:05:14

Yeah. Yeah, don't power things on front end, guys. Like Uh, oh, even though it's free compute.

Jake Cooper1:05:20

Yep.

Alessio1:05:20

Yeah, yeah. Cool. Any other technical infrastructure cool stuff? Railpacks? I don't know if that's still, um-

Railpack1:05:27

Jake Cooper1:05:27

Yeah. Uh, yeah. I mean, we built an engine for determining, uh, dependencies based on your source code, which is super cool. It's called Railpack. We built the first version called Nixpacks, which is on top of Nix. Uh, and then yeah, we moved.

Alessio1:05:37

People have been trying to get me to adopt Nix and NixOS for, like, four years.

Jake Cooper1:05:41

Yeah.

Alessio1:05:42

Is it gonna ever gonna be a thing?

Jake Cooper1:05:43

I don't, I don't know. Like, we're super excited about it in general, but it's like it has a bunch of different, uh, kind of, uh, pain points in general, 'cause if you just think of it, it's like, it's a s- it's a stack of versioned source code at...

or it's a stack of versioned binary at specific slices in time, right? And so if you want version X and version Y, you end up bloating a lot of your kind of, uh, like, package, like, space, right? Which blows up the size of your images and-

Alessio1:06:09

Yeah

Jake Cooper1:06:09

... and makes it really, really difficult, um, for, like, really real world workloads. I think if you-

Alessio1:06:13

But you, you know, you, you content address it, and you cache it. It's, you know, there's, there's a lot of optimizations that in theory you should be able to do.

Jake Cooper1:06:20

In theory, yes. Right? Um, and what, what happens ultimately is, like, you have a large enough user base, and you have a disparate enough set of, uh, machines that you kind of run into the, the problem that, uh, there's a paper that Meta released, uh, XFAAS.

They're, like, internal, uh, kind of like, um, serverless, uh, system. It becomes an... it ends up being, being very, very difficult to go in and, and do that at scale, unless you break out specific runtimes.

Alessio1:06:45

Hmm.

Jake Cooper1:06:45

Basically. Um, which we, uh, did not wanna go in and, and do, right? Because we wanted to truly allow you to deploy anything, right? Which was our initial kind of thing with Nix. Uh, but we've moved towards some interesting stuff that I think we'll, we'll be able to talk about a little bit later, um, that we've, we've built for- Doing conscious addressable file systems to be able to, like, lazy load, um, anything from any point, um, and then just, uh, page that into memory.

Alessio1:07:08

Amazing. Okay.

Jake Cooper1:07:09

Yeah, it's gonna be fun. There's, uh, the, the whole future is very, very bright. It's, it's crazy. It's, it's gonna be nuts.

Death PRs1:07:14

Alessio1:07:14

Uh, okay. Founder journey stuff?

Swyx1:07:16

Yeah. And your, uh, cloud usage. You tweeted you're gonna spend 300K this month?

Jake Cooper1:07:21

Yeah, I think we got-

Alessio1:07:22

What?

Swyx1:07:22

Is that all-

Jake Cooper1:07:22

I think we got 200

Swyx1:07:23

... coding agents or-

Jake Cooper1:07:24

Yeah

Swyx1:07:24

... across the company? Yeah.

Jake Cooper1:07:26

Yeah.

Swyx1:07:26

You only have 35 people, so-

Jake Cooper1:07:27

Yeah, I know

Swyx1:07:27

... I'm sure they're not all spending 10K a month. What's kinda like the distribution?

Jake Cooper1:07:31

I think I'm at about 25, uh-

Swyx1:07:33

Nice

Jake Cooper1:07:33

... in, in general. Um, and then we have some, you know, power users kind of all, all the way down. I don't know, like, w- we came back and, or from, from the winter break, and I was basically like, "If you're writing code by hand, you are doing this wrong," right?

Like, their code- the tools are good enough at this point, like, that you can move extremely, extremely quickly. And like, yes, there are issues and pain points and all of these other things, um, but you should be reviewing the code that you are writing instead of trying to go in and write it by hand.

Like, all of those architectural patterns, all of those other things, like-

Swyx1:08:01

Yeah

Jake Cooper1:08:02

... you're not... Just don't, like, you're not gonna throw them in the garbage or whatever. Actually, they matter more now than, than, than any other time. But you just shouldn't spend your time generating code that you would write.

Like, if you know how to go in and write it, just, like, ask the agent to go in and write it and then reconcile it until it looks like you would've written it yourself, right? And I think incidentally, like, people misconstrue, uh, my propensity to, like, push people towards agents for, like, "Hey, we're growing really, really fast, and we've had some kinda, like, bumps in reli-"

Swyx1:08:29

Yeah

Jake Cooper1:08:29

... like, they're not necessarily related, uh, in terms of that. Um, but I think people should really, really understand, like, the tools are good enough for you to be able to move extremely, extremely quickly to build things way, way larger than you could've-

Swyx1:08:40

Mm-hmm

Jake Cooper1:08:41

... possibly built before, right? Um, and so to our point about, uh, way earlier about, like, how do you cool data centers in space? It's like, well, I don't know, actually. Right? Um, but you're at a point now with, with software, you can actually be like, "Well, how would I build block storage from scratch?

How would I go in and do these things? I have ideas because I've got history and I've read all these papers in general, right? Let me go in and work them out in general-

Swyx1:09:01

Mm-hmm

Jake Cooper1:09:01

... and let me build, like, massive test benches with, like, thousands of tests," right? Because they're free to, they're free to author right now, right? To go in and make sure that, like, this system can now, can, can be built, right?

And I think that if you're not using the, the kind of AI systems to almost, like, speed run your roadmap to, like, go in and, and figure out where you need to go in and be to reconcile your existing system onto the future, then you're kind of missing a large point of, of what is currently happening right now.

Swyx1:09:28

Yeah.

Jake Cooper1:09:28

Right? Because you can just template out anything and validate it on the side for free, right?

Swyx1:09:32

What's the path to spend 3 million a month? Like, is it bound by, like, ideas and things that the customers can absorb?

Jake Cooper1:09:39

I think for, for most companies, it's actually bound by deployment at this point in time, and I think that's why we've seen a lot of, like, a massive boon in terms of, like, users trying-

Swyx1:09:46

Mm-hmm

Jake Cooper1:09:46

... like, not just users, like companies, like, you know, Fortune 50s, like, you know, below, et cetera. Like, going and being like, "How do we get our developers to, like, go in and move quicker," right? Um, I think you're probably gonna hit, um, your CFO before you hit, uh, any of these limits in general, 'cause they're gonna look at this and be like, "There's an eye-watering amount of, like, money being spent on these tokens."

Like, I think, uh, I don't know which... I think it was the Uber C- C-

Swyx1:10:09

Yeah, yeah, yeah.

Jake Cooper1:10:10

... it was like they blew their token budget for the entire year or whatever, right? Um, and so inference has- costs have to come down, but they're also, you know, they're, we're inference constrained at this point in time, right?

And so you're gonna almost get this, like, price discovery of, like, what makes sense for an org to go in and adopt?

Swyx1:10:26

Yep.

Jake Cooper1:10:26

And I think what you're gonna end up with is actually you're gonna almost, like, end up with the, like, F1 driver concept, um, which is if you have a, if you have somebody who's, like, really, really adept at these things, uh, it makes sense to go and put them into, like, a $3 million car or whatever, right?

Swyx1:10:40

Mm-hmm.

Jake Cooper1:10:40

But if you're not, then, like, it probably doesn't actually make sense for you to go in and do that. And we're gonna take a few of these people and say, "You can drive the F1 car. We need to go in this general direction, figure out w- if this works, and, like, almost go ahead and prototype it," right?

Swyx1:10:53

Yeah.

Jake Cooper1:10:53

Um, and so we've done a few of those things. We're like, we've vastly accelerated our roadmap in terms of, oh, we thought we were gonna be able to go in and ship this thing in the next, like, few years, but actually we can probably ship it in the next, like, few months now, right?

Because we're saying, "Oh, validated it out, it works. Don't have to even, like, build it incrementally. We can now skip steps to, like, go and, and just move towards where our vision is for a lot of this stuff."

And I think that that's kind of where we end up with a lot of it, you know?

Swyx1:11:18

Yeah. I think a lot of people are realizing the roadmap doesn't always have a business impact, and so it's like, "Oh, it's too expensive to run these tokens." But, like, if your roadmap was actually built to make more money by the time you built the whole thing, you would have some sort of token pricing for it, the same way you do with sales.

Jake Cooper1:11:34

Yep.

Swyx1:11:34

Like, you would spend a billion dollars in sales if you knew you would get $2 billion of revenue out of it.

Jake Cooper1:11:39

Exactly, right? Um, and I, I think the, the, the really naive way to go in and measure this is almost, like, your percentage of tokens that end up in production.

Swyx1:11:47

Right.

Jake Cooper1:11:47

And so if you can measure that you are getting this level of, of impact because those tokens are ending up in production, that's awesome, right? Um, but I think the, the kind of burden of proof is now gonna kind of, like, arise, and you see it internally too on, on our stuff.

Like, we have a growing number of pull requests that, like, are s- are, like, haven't yet been merged-

Swyx1:12:03

Solving

Jake Cooper1:12:03

... right?

Swyx1:12:03

Yeah, yeah.

Jake Cooper1:12:04

And, and you're just like, "Okay, how do you get this into production," right? Um, and so it's really about, like, how quickly you can go and kind of build and deploy that software, right? Um, which is exciting 'cause our whole thing is-

Swyx1:12:13

You deploy software

Jake Cooper1:12:14

... we build and deploy software, you know? Right. So yeah.

Alessio1:12:17

Yeah. The SDLC is changing, and it's something that both of us are, like, super interested in exploring as well. Um, my, one of my thesis, or it's not my thesis, it's the pull request is dying.

Jake Cooper1:12:28

Yeah.

Alessio1:12:28

Right? It's, it's gonna be the prompt request.

Jake Cooper1:12:30

Yep.

Alessio1:12:30

And then beyond that, code review is also kinda dying because do you really need to, uh, if you have all the other systems in place. What else is changing about the SDLC?

Jake Cooper1:12:40

What else is changing? Well, I think the-

Alessio1:12:41

AISRV

Jake Cooper1:12:42

... the AISRV, the tools to, to ma- So the AISRV is, like, one of those things where it's like, you know, it's a pie in the sky aspirational.

Alessio1:12:49

Yeah.

Jake Cooper1:12:49

What, what does it take to get an AISRV? Like what tools do you need-

Alessio1:12:52

And by the way, you should expose your tooling to your customers at some point. Right?

Jake Cooper1:12:55

Yeah. Uh, well, which tooling?

Alessio1:12:56

The central command center

Jake Cooper1:12:59

Oh, Central Station?

Alessio1:13:00

Yeah, yeah

Jake Cooper1:13:00

So yeah, yeah. The, like we, so we have it for template maintainers, right?

Alessio1:13:03

Right.

Jake Cooper1:13:03

So template maintainers can, like, deploy and maintain templates, and they get feedback on a lot of that stuff, right? And so we're 100%, like, going to go in and explore those things, like incrementally. Um-

Alessio1:13:11

Yeah, but like, you know, clustering around incidents. Like, everyone has a version of that, but like-

Jake Cooper1:13:15

Totally

Alessio1:13:15

... I don't think anyone's solved it.

Jake Cooper1:13:16

Yeah, yeah, right. Um, and I don't wanna say we've solved it internally, but it's gotten, it's gotten so good that, like, now we can see, uh, those incidents forming, like, pretty quickly.

Alessio1:13:25

Yeah.

Jake Cooper1:13:25

Right?

Alessio1:13:25

Real time-

Jake Cooper1:13:26

Yeah, yeah

Alessio1:13:26

... and like AI clusters.

Jake Cooper1:13:27

Yeah.

Alessio1:13:27

Yeah.

Jake Cooper1:13:27

Um, so the, at some point, those will, those will be things that either somebody else goes and builds or we go in and build, but we've always built stuff that, like, was purposeful for us. And if it made sense and, and there was a way to go in and make it useful for users or monetize it or, or make sure that that loop becomes like a profit center instead of a cost center, like, we wanna go in and-

Alessio1:13:45

Yeah

Jake Cooper1:13:46

... and do that at some point, right? Um, so but yeah.

Alessio1:13:48

Do, do you-

Jake Cooper1:13:49

Port- Portus definitely dying

Alessio1:13:49

... do you do first party, uh, feature flagging and, um, incremental rollout type stuff as well?

Jake Cooper1:13:54

So we have a feature flagging engine that we built-

Alessio1:13:56

Yeah

Jake Cooper1:13:56

... internally that at some point we will, we will roll that out.

Alessio1:13:58

Because I don't see it as a user.

Jake Cooper1:14:00

Yeah, yeah.

Alessio1:14:00

Yeah.

Jake Cooper1:14:00

Yeah. Um, so-

Alessio1:14:01

So like that, you know?

Jake Cooper1:14:02

That would be... That's good, right? And, and-

Alessio1:14:03

How come you didn't give us what you have?

Jake Cooper1:14:05

Well, because, because we have to, we have to beta test it. Like, we actually care a lot, a lot, a lot about the quality of the, the things. There's plenty of stuff that, like, we've used internally, and then we've got it to a point where, like, it di- it doesn't make its way entirely through the journey because it fails.

Alessio1:14:19

Yeah.

Jake Cooper1:14:19

Right? It's, it's like this, this holds for one service, but it doesn't hold for multiple services, right? So we'd have to go and build these things for multiple services to go and make this work, right? And we know for a fact that if we release this thing, we'd have to go and rebuild this thing again, and again, and again.

And some things are worth doing to go in and do that. But a lot of them are basically like that also, that kinda just informs our roadmap of, okay, well, like, for us to go and make that actually a bit easier, we can do the few of these things first, and then we get to that experience, right?

Alessio1:14:44

Yeah.

Jake Cooper1:14:44

Um, we don't wanna dilute the experience by basically saying like, "Oh, yeah, this works, but only for this service," right? Um, unless it's like a very, very core initiative, which is like, you know, over the next, like, few months we're gonna roll out a few things that are like, okay, it works for a single service, and then it works for multiple services, and then it works multiple services across the environment, right?

Um, but you have to be very, very deliberate about those things, otherwise you end up with a bunch of broken disparate experiences, which ultimately end up creating a ton of support load. 'Cause people are like, "How do I use this feature?

How do I go in and do this other stuff," right? Um, so-

Alessio1:15:12

Yeah

Jake Cooper1:15:12

... it's kinda the, the thing earlier about like, you expand your company i- in general to, to get those, like, features, and then you almost compact it, smooth out those things, so the experience is like really, really stellar.

Like we were talking in the hallway earlier, where you're like, "Oh my God, it's gotten so much better," and I'm like, "Oh, man," like, just internally we're like, "Damn, this part really sucks." Um-

Alessio1:15:29

Yeah

Jake Cooper1:15:29

... you know, we gotta make this significantly, significantly better.

Alessio1:15:31

No, I can, I can attest, uh, you know, over the last three years that I've, uh, watched you build Railway. Uh, but yeah, no, I would call to, to listeners, if you're not aware, like, the importance of feature flagging, it's a very part, big part of Uber culture.

Jake Cooper1:15:44

Yep.

Alessio1:15:44

So much so that they have too many feature flags, and then they have another thing to remove feature flags.

Jake Cooper1:15:49

Yep. 100%.

Alessio1:15:50

What was it called? I, I-

Jake Cooper1:15:50

Uh-

Alessio1:15:51

There's a, there's a paper about this

Jake Cooper1:15:52

... there's Flagger and, and there's, there's been another one. Uh-

Alessio1:15:54

There's a thing that, like, looks for feature flags.

Jake Cooper1:15:55

Death PRs has Gatekeeper. Uh, yeah, so they're, they're really important.

Alessio1:15:59

And agents are gonna need this. That- that's like the fundamental thing behind, you know, like, just incremental rollouts.

Jake Cooper1:16:05

Yep.

Alessio1:16:05

OpenAI acquired Statsig.

Jake Cooper1:16:07

Yep.

Alessio1:16:07

Uh, and this basically GPC5 is just routing and flagging, you know, through, through different models and like-

Jake Cooper1:16:13

Yep

Alessio1:16:13

... that's like-

Jake Cooper1:16:14

And it's, it's super important, right? Because if- ... if you, if you assume the, the software development life cycle is 100% gonna go in and change, but it's gonna change because we're trying to do things a thousand times faster and a thousand times more concurrent than-

Alessio1:16:26

Yeah

Jake Cooper1:16:26

... we're, we were currently doing, right? And so-

Alessio1:16:28

This is routing

Jake Cooper1:16:28

... yeah, right. Um, and so what ends up becoming important at scale, you know, before I even, you know, started Railway, I actually built a feature flagging product.

Alessio1:16:36

Oh.

Jake Cooper1:16:36

And I tried to, I tried to go in and sell it to people.

Alessio1:16:38

Okay.

Jake Cooper1:16:39

Right? 'Cause I was like, "Oh, it's like a, you know, like a, it's an easier version of like LaunchDarkly or whatever," right?

Alessio1:16:42

Yeah, yeah.

Jake Cooper1:16:42

And then I ran into this situation which is like, anybody who's small enough to adopt your technology doesn't care about feature flags, right? And then anybody who's large enough to try and actually need feature flags needs so much scale that you have to, like, build out all the existing infrastructure to end up scrapping that.

But it's, what is old is, is new again because now companies are trying to move really, really quickly, but you can't just YOLO this, like, vibe-coded thing straight into production. You need to basically say, "Hey, here's my blast radius, here's my impact, here's my, like, whatever.

I wanna shadow it for these users," right? Feature flags, right? Like, you're gonna need those tools that ultimately those larger companies ended up having to go in and build to maintain their structures. Everything's just gonna get compressed by like 1000x so that everybody can go and do that, and everybody can build those structures really, really quickly, right?

Alessio1:17:27

Yeah, yeah.

Jake Cooper1:17:27

And that's like exactly where we're at right now is like you're compressing the software development life cy- life cycle, and then we're gonna expand it and, and add way more new things to it, you know?

Alessio1:17:35

Yeah. Uh, and then the other term that comes to mind with when this kind of discussion happens for me, for newer developers who haven't heard this term, cattle not pets.

Jake Cooper1:17:43

Yeah.

Alessio1:17:43

Right? Because like your prod, it- people treat it like a pet.

Jake Cooper1:17:47

Yep.

Alessio1:17:47

Like it has a name. I, I, baby, you know, I have to keep it alive.

Jake Cooper1:17:50

Yep.

Alessio1:17:50

But when it's cattle, you can just mass farm and you can like roll out and you can like, you know, uh, portion out parts of them and kill them or whatever.

Jake Cooper1:17:58

Yeah. Yeah, exactly. Um, I actually-

Alessio1:17:59

Yeah

Jake Cooper1:17:59

... I actually think that maybe that's the, the hot take, but I think that that's actually gonna change, and I think-

Alessio1:18:03

Yes, please

Jake Cooper1:18:04

... you can move towards having pets so long as, uh, you have a, and this is gonna be a jump, so long as you have a cloning machine for your pets.

Alessio1:18:12

Uh, yeah, yeah.

Jake Cooper1:18:13

If you can snapshot every single thing at every frame, then like it actually doesn't matter if, you know, that thing gets obliterated because you have some sort of like snapshot of it, right? All of the things that we have built right now are to essentially block out, uh, any sort of, um, changes or alterations or whatever from that like hermetically sealed DevOps like line-

Alessio1:18:34

Mm

Jake Cooper1:18:34

... or whatever. It's like, okay, well, you have to write a Docker file because I only need these specific instance, like only this specific cut of the file system, et cetera, right? What if you just had the whole file system?

What if you just snapshot it? What if you lazily load the entirety of the file system, right? Then you could get around this problem entirely. You don't need the ceremony of those, you know, having a Docker file or like having an Ansible script or like having all of these other things.

You can just iterate on that loop and then like snapshot it. It's like, is this the right loop? Is this the right thing at this point in time? Okay, cool. Like, now I'm gonna go and merge it in production.

Like, go merge the file system.

Alessio1:19:06

Yeah. Yeah

Jake Cooper1:19:06

Right. It's gonna be really fun.

Alessio1:19:07

Yeah. This is like a whole other can of worms, but, like, I think the number of things that are state full in a VM, uh, I think if you just like kinda cataloged them and just like developed dedicated solutions for solving each of them, you can actually kind of cut this down, problem down a lot, and it's surprising that people weren't really trying until now .

Jake Cooper1:19:25

Yeah. Well, uh, so it's, it's surprising... I mean, it- it's always been surprising to me because these are the things that we would work on, 'cause I'm- they're just like, I'm like, it's so obvious.

Alessio1:19:31

Yeah, first principles, you need them.

Jake Cooper1:19:32

Yeah, right.

Alessio1:19:33

Everyone, like, in theory, needs them. And then like the big clouds don't do them, so you're like, "It's impossible." There's something. I don't know.

Jake Cooper1:19:38

Yeah, exactly. Right? You're like, oh, well, they've... You know, Meta has all the people who write like, you know, uh, EBPF code and, and they're like doing something with them. Um, you know, but like you need that kinda stuff to solve these problems, right?

Alessio1:19:49

Yeah.

Jake Cooper1:19:49

And, and like talked about earlier, like, it's like whatever is required, however deep we have to go in and like get to like solve those problems, right? Like, all the way down to like the kernel's TCP/IP stack, right?

Like, we're gonna go and figure that out. Is there something that we need to go in and modify to like go in and make that work for the mental model that we have for the universe moving forward? Like, yeah, 100% we're gonna go in and do it, and we'll just keep going that way down.

Alessio1:20:13

Man, it sounds fun. Let's see.

Jake Cooper1:20:14

It's super fun. It's like s- it's so much fun. Like, I have to literally peel myself away from the fun, interesting problems that we have to make sure that we can scale the company in, in a way that like works.

And there's so many different fun, interesting problems, whether it is like how do you get the information from the customer to, um, support to the person who built the thing internally, right? Um, or it's like how do you get, do safe iteration, or how do you get context like from the dashboard to users, or like how do you drill down all the way to the infrastructure layer?

How do you manage orchestration as like a real-time operating system versus a feedback control system, right? Like, it's just so fun, you know?

Alessio1:20:50

Yeah. I mean, speaking of, maybe you talk about the founder side. Um, you're famously, like, you know, the YC con- uh, the, the SF consensus is you go to YC, you get a co-founder.

Solo Founder1:20:50

Jake Cooper1:20:58

Yep.

Alessio1:20:58

You get to do all these things. Uh, you have done none of that.

Jake Cooper1:21:01

No. I've just... Yeah. I've, I've like done a lot of different things in general.

Alessio1:21:06

In the, in the elevator you were like, actually, co-founder, it kinda makes sense if like one person is the tech person, the other p- is the biz dev person.

Jake Cooper1:21:13

Yep.

Alessio1:21:14

And, but you have to contain all those multitudes yourself.

Jake Cooper1:21:17

Yeah.

Alessio1:21:18

How do you do it?

Jake Cooper1:21:18

Oh, okay. I was gonna ask, is there a question in there or what? Yeah. Um, how do you-

Alessio1:21:21

The question is, what the hell?

Jake Cooper1:21:22

How do you, how do you... Um, yeah.

Alessio1:21:23

The question is, how, how are you alive right now?

Jake Cooper1:21:26

Yeah. I... Well, I mean, yeah. I mean, just try to get eight hours of sleep. Uh- ... you know, like.

Alessio1:21:31

Is there like a balance that you ideally, like 50/50, 30/30/30? Like, what, what's the mental model that you use-

Jake Cooper1:21:37

There's no-

Alessio1:21:37

... as a solo founder?

Jake Cooper1:21:37

There's no balance. There's... Like, you just, you just have to think about all these things and, and be obsessed with all of these things. Like, whether it is being obsessed with like how do people think about your product from a go-to market perspective, or being obsessed from a perspective of like, well, like, if I can make this change at the like kernel level, then I can make it so that the user's SSH connection never drops, right?

Like, because that's what I want. Like, I want a universe in which I can go and like snapshot all these things, and it looks exactly like you would just kinda iterate on, on a VM, right? And I think you just have to be obsessed with all those things, like, at every layer of, of the stack.

And I think that that's what makes it easier for me. I think some people, like, they're obsessed with different portions of the, the kind of like journey cust- like, the company, like, whatever, right? And I think that that's when you can get really, really good almost like cohesion by like segmenting out these things, right?

And so, you know, in the elevator I was talking about, like, you know, you have a technical kind of like person, et cetera, and then you have the customer kind of like person in, in general, right? Um, and I think, like, if you can segment those lines out really, really well and you can be very, very clear about what your areas of ownership are for yourself or your, you know, company or any, like, just where you're gonna operate, you're gonna have a good time, right?

If you can't be clear about those things, right... And this is why I was saying like two is the worst number of co-founders, is because you have no tiebreak, right? You, you basically are like, "Well, I disagree on this thing, and I disagree on this thing," right?

It's like, well, how do you resolve that, right? If you-

Alessio1:22:59

Well, usually someone's CEO, right?

Jake Cooper1:23:00

Right. Exactly, right?

Alessio1:23:01

Then you're like, "Okay, you have the tiebreaker."

Jake Cooper1:23:03

Yeah, totally. I mean, listen, it's hard, it's hard every single way you cut it, right?

Alessio1:23:06

Yeah.

Jake Cooper1:23:06

It's hard, it's hard if you get help. It's hard if you do it yourself. It's, it's just, it's just hard to like run things, roughly speaking, right? But it's so rewarding. It's so fun, you know?

Alessio1:23:16

What have you found useful? Like a coach, um, any ad- advice that has been really helpful, um-

Jake Cooper1:23:21

I like to write a lot. Um, I, I got in trouble, I get in trouble a lot for my Twitter. Um, I guess there's a pattern. Um-

Alessio1:23:27

Who do you get in trouble with?

Jake Cooper1:23:28

The people on, on Twitter.

Alessio1:23:30

Oh, okay.

Jake Cooper1:23:32

Um, you know, um, you know, I was talking about it and I was like, "Hey, if you, you know, if you're working weekends, you're kind of messing up your planning roughly," right?

Alessio1:23:37

Mm.

Jake Cooper1:23:37

Um, and I've gone kind of back and forth on that, right? Because I think actually right now we're kind of at an extenuating time in general where it actually makes sense to like work more, right? Because the, the goals are pretty clear in, in my mind, right?

Um, and so if you have the vision and you know where you're going, you should work a little bit harder to distill that vision and go and do those things. But if you don't have the like, "We're, we're like, I think we should be going this direction, I'm not 100% certain, I wanna get a little bit of clarity," I think what you need to do is you need to like disconnect and you need to take your weekends like very, very seriously.

You need to write about where are you, what do you wanna do, where do you wanna go, what problems are you trying to go and solve, and like think about a lot of these things, right? So, um, you know, like, writing is important.

Uh, sitting down, like, I, I don't like the word like meditation or whatever, but like whatever gets you into the, the state of like your mental clarity, like that's the thing that's like really, really important when you're trying to go on these journeys of, of saying, "Well, we're here, and we really need to be, be here in, in general," or like, "We're here, and I think we need to be roughly in this kind of like space for this to like work," right?

So those are the things and then, you know, disconnect, hang out with the people that you love, and then like work super, super hard when, when you're like, you know. Like I try and work like sun up to sun down, Monday to Friday, all out in general, and then I try and disconnect on Saturday, and then I come back to work on, uh, Sunday afternoon, right?

And then I do my writing, plan for the week, all those other things, and it, it works really, really well for me. Um, but another hot take is like most advice is to be digested and to be thrown out the window, and if it's helpful, it'll come back, right?

If it's helpful, you'll have kinda like learned it over time through experience or anything else like that. But yeah, you mentioned like the, the kinda standard, you know, YC advice, all of those other things. We have a lot like- We've made failure as a society very, very expensive, and it makes it difficult for people to kinda trod off the paths, right?

So...

Alessio1:25:23

Yeah. Makes sense.

Swyx1:25:25

Any other SOP books you wanna get on? Like, anything that you have not tweeted and gotten in trouble with that you wanna preview to the world?

Jake Cooper1:25:33

Hmm. No, I think the, the agent stuff is like, it's just like, it's crazy. It's, it's gonna be, it's gonna be the dominant way in which people are doing pretty much everything, right?

Swyx1:25:42

Yeah.

Jake Cooper1:25:43

Um, it, provided we can, of course, get the amount of inference required for, for that to go and, and happen. Uh, but over the next, like, 10 years, right, you just, you see a fundamental shift in terms of how people are thinking about even just authoring the logic that's in their head, right?

So...

Swyx1:25:56

Yeah.

Alessio1:25:57

My, m- m- you know, maybe one, one way of phrasing this is if all birds can become a GPU provider, so can Railway.

Jake Cooper1:26:04

Y- yeah. So I don't... I think there's a lot of harm in us actually not becoming a GPU provider.

Alessio1:26:08

Yeah.

Jake Cooper1:26:08

I think, I think you're, you're defined almost more by the things that you don't do than the things that you do because it's e- it's really, really easy for you to just say yes to a bunch of different things, right?

Alessio1:26:17

Yeah.

Jake Cooper1:26:17

And I think, like, it's gonna be very, very interesting to watch, uh, you know, I, I think Anthropic is, like, an amazing company and, like, super, super stellar, and they're moving into a variety of different zones, right? They're moving into, like, the Figma kinda, like, stuff that they're, they're after.

Like, they're moving-

Alessio1:26:29

Today-

Jake Cooper1:26:29

Yeah, right? Uh-

Alessio1:26:30

... as of recording-

Jake Cooper1:26:31

They're, they're, they've got Cloud-

Alessio1:26:32

... and so Mike Figueredo-

Jake Cooper1:26:33

... they've got all this other-

Alessio1:26:33

... was on Figma's board, and then they removed him, like, Monday, and then they launched this today.

Jake Cooper1:26:37

Yeah. Yeah.

Alessio1:26:37

It's like, ugh.

Jake Cooper1:26:39

Yeah, so I mean, things, things move very, very fast, uh, right now. Um, so, but yeah. It's just gonna be the way in which people are, are, are-

Alessio1:26:45

Okay, so, so your answer is focus, no GPUs for now-

Jake Cooper1:26:48

Yeah, focus. Focus-

Alessio1:26:48

... but never say never.

Jake Cooper1:26:49

Yeah, yeah. Right? Um-

Alessio1:26:50

Yeah

Jake Cooper1:26:51

... like, I can tell you for a fact that we will not be doing GPUs now, but we 100% will be doing GPUs at some point in the future. And that's-

Swyx1:26:58

Oh, okay

Jake Cooper1:26:58

... and that's not, like, me leaking our roadmap-

Swyx1:27:00

Yeah

Jake Cooper1:27:00

... because we don't have plans to go and do GPUs. It's just a function of at some point you need flops, right? Like, at some point you want... Like, if you're fully vertically integrated and you wanna make it really, really trivial for people to go and iterate and build and deploy things, you need access to this core piece of fundamental logic, right?

Um, so yeah.

Alessio1:27:18

Yeah. And then, like, at s- at some point, uh, presumably the, your own data center traffic is, like, a minority of your workload right now, but is there, like, a majority or, you know, you just kinda completely turn it off of, of the clouds?

Jake Cooper1:27:30

Oh, it's at s- at some point we got to 100% data center. Like, we-

Alessio1:27:33

You own data centers

Jake Cooper1:27:34

... our own data centers.

Alessio1:27:35

Yeah.

Jake Cooper1:27:35

Yeah, yeah, yeah. Um, it's, and it's, right now, it's the vast majority of the stuff that exists on, on our, our bare metals data centers, right?

Alessio1:27:40

Okay.

Jake Cooper1:27:41

So it-

Alessio1:27:41

So you're already there, like, vast majority.

Jake Cooper1:27:43

Yeah, yeah. Right?

Alessio1:27:44

Okay.

Jake Cooper1:27:44

Um, like, the data-

Alessio1:27:45

I, I, I didn't, I didn't know the, the extent-

Jake Cooper1:27:46

Yeah

Alessio1:27:46

... of the transition.

Jake Cooper1:27:46

Yeah, totally. It was completed at some point, and then we grew so fast, um, that we had to basically, like, go and, and scale back on, on the-

Swyx1:27:54

Take us back.

Alessio1:27:55

Yeah.

Jake Cooper1:27:55

Yeah, basically.

Alessio1:27:55

Sorry, Google Cloud.

Jake Cooper1:27:56

Yeah, it was funny. Like, we got- It was funny. We got to, like, on, on the Datadog dashboard, it's like it got to 100%, and then it, like, divoted back down into the, like, 90s or whatever 'cause we were like, "Yeah, yeah."

Future Cloud1:28:00

Alessio1:28:04

And you're adding capacity.

Jake Cooper1:28:05

Yeah.

Alessio1:28:05

Yeah. It's, it's interesting. You're, you're literally building a new cloud, a- a- and that's independent, and people assume that that could never happen post, you know, the AWS.

Jake Cooper1:28:14

Yeah, and, and it's- ... and it's hard, right? Like, you know, we, we, we're gonna, you know, figure out a bunch of different things, uh, to, to make sure that, like, the platform is deeply, deeply reliable.

Alessio1:28:23

Yeah.

Jake Cooper1:28:23

But you have to break ground on a, on a lot of new things when you basically decide you're gonna build a cloud from scratch but not copy the hyperscalers, right?

Alessio1:28:31

Yeah.

Jake Cooper1:28:31

Like, we've been very, very deliberate to, like, invent our own infrastructure from scratch based on reading a ton of papers in general, but, like, almost, like, promising to ourselves that we wouldn't copy somebody else's homework, right? Um, because we were saying, "Hey, listen, you know, if we co- if we, if we copy somebody else, we lose."

Like, we just, you're just gonna become them over time, right? And so you have to have a core thesis about, like, why does this business need to go and exist at this point in time? And for us, it's always been about the activation energy to get something to go and deploy it in production, uh, at any of the hyperscalers as, as of right now is far too high, right?

And we believe that it should be instantaneous. We believe that there should be no friction in between what your thought is and reality that kinda comes out that you can share with your friends, right? Um, and so that's, that's what we're kind of, like, building toward a- again, at every layer of the stack.

Like, if we gotta go down to energy, we'll go down to energy at some point, right? Like, it, it, it just, it matters a lot for us from, from the experience of, of giving people access to this tooling because it's, it's gated behind...

Like, it's not even just gated for regular kinda, like, these citizen developers that are now vibe coding. It's like you have multiple layers. You have the citizen developer, you have the front end developer, you have the back end developer, you have a DevOps person, you have, like, all of these layers, right?

And they all need to go in and disappear so people can just, like, ship like that.

Alessio1:29:40

Amazing. All right. That's the future-

Swyx1:29:42

Awesome. Thanks for coming

Alessio1:29:43

... of cloud. Yeah.

Jake Cooper1:29:43

Thank you. Thank you for having me. It's been wonderful.