# Browserbase: Browser Infrastructure For Your AI Agents

Latent Space · 2025-02-28

<https://addtry.com/6a58431b-e247-4f52-8c48-4082afe8d112>

Paul Klein IV, CEO of Browserbase, argues that headless browser infrastructure is a critical primitive for AI agents to automate the web. Browserbase solves the hard problem of running thousands of browsers in the cloud—too big for Lambda, requiring Firecracker microVMs—and exposes APIs via its open-source framework Stagehand (act, extract, observe). It handles proxies, CAPTCHA solving, and offers a live view iframe for human-in-the-loop control. Primary use cases include browser automation and web scraping as a heavy-hitter renderer when simpler methods fail. Paul explains his choice to go solo founder and how Browserbase's in-person, 10–5 culture enables fast execution. He predicts the future of software will involve software acting on other software via browsers.

## Questions this episode answers

### What is Browserbase and why is it necessary for AI agents?

Browserbase is a headless browser infrastructure platform for AI, according to CEO Paul Klein IV. It runs browsers in the cloud, abstracting away the complex scaling, state management, and security needed to automate the web. This is critical for AI agents that must interact with JavaScript-heavy sites, click buttons, and fill forms, without developers managing server fleets or container orchestration themselves.

[0:02](https://addtry.com/6a58431b-e247-4f52-8c48-4082afe8d112?t=2000)

### How does Browserbase solve CAPTCHAs and manage proxies for reliable web automation?

Paul explains that Browserbase integrates multiple CAPTCHA solvers and proxy providers, continuously monitoring them and routing around failures. It offers a proxy network for IP-based location control, sparing developers from dealing with unreliable, 'sketchy' services. Browserbase sees CAPTCHA solving as a temporary measure while they work toward 'good bot' authentication through partnerships with companies like Cloudflare.

[0:21](https://addtry.com/6a58431b-e247-4f52-8c48-4082afe8d112?t=21000)

### What is Stagehand and how does it use LLMs for web automation?

Stagehand, open-sourced by Browserbase, is a framework extending Playwright with three natural-language APIs: act (perform actions), extract (get structured data via Zod schemas), and observe (list possible actions). As Paul describes, it lets developers write one generic script that adapts to any website's structure, using LLMs to generate the specific Playwright code on the fly, making automation resilient to page changes.

[0:33](https://addtry.com/6a58431b-e247-4f52-8c48-4082afe8d112?t=33000)

## Key moments

- **[0:00] Intro**
  - [0:50] Browserbase is one year old and already raised a Series A with a team of 20, serving hundreds of AI companies.
  - [1:14] Paul Klein's first vacation was interrupted by the launch of OpenAI Operator and DeepSeek, forcing him back to building.
- **[2:00] What is Browserbase**
  - [2:00] Browserbase provides headless browser infrastructure for AI agents, born from Paul Klein's previous startup Stream Club.
  - [3:24] "I truly do only have one trick. And like I, I'm screwed if it's not for headless browsers, you know?"
  - [3:36] Browserbase was in the AI Grant batch but used zero dollars on AI spend, being purely an infrastructure company.
  - [4:11] Q: What AI-specific challenges does Browserbase solve compared to traditional headless browser services?
  - [5:57] LLMs enable writing one generic automation script that works across many websites, unlike pre-LLM static scripts.
  - [7:13] Paul Klein didn't think vision would be a big driver for UI automation, but computer use models are pushing things forward.
- **[11:02] Scaling Challenges**
  - [11:11] Browserbase's website was designed by Parisian agency herb.paris, which builds beautiful consumer-like dev tools interfaces.
  - [12:26] Running headless browsers at scale is hard because Chrome is too large for Lambda and requires Kubernetes orchestration.
  - [15:28] "When I see a complex distributed system, I see an opportunity to build a great infrastructure company."
  - [16:09] Q: How does Browserbase spin up thousands of browsers in milliseconds?
  - [17:45] Browserbase had to go lower level than AWS Fargate, despite being a large Fargate customer, because infrastructure companies need deeper control.
  - [19:12] Browserbase offers proxy networks to ensure browser traffic appears from specific locations, like the US, for reliability.
- **[19:56] Proxies & CAPTCHAs**
  - [22:33] Paul Klein predicts CAPTCHAs will eventually differentiate between good and bad bots, and Browserbase aims to be an 'arbiter of good bots'.
  - [24:11] Paul Klein predicts agent authentication, like an OAuth flow for AI agents, will replace CAPTCHAs for identifying good bots.
- **[24:59] Agent Identity**
  - [26:25] Paul Klein applied to 500 internships, got rejected by all except Twilio and Amazon, and became a tech lead in two years.
- **[28:09] Live View**
  - [28:22] Browserbase's live view iframe feature uses Chrome DevTools Protocol to stream browser screens and enable remote control.
  - [29:24] Browserbase's live view streams browser screens as PNG images, allowing human-in-the-loop intervention for tasks like 2FA.
- **[33:44] Stagehand**
  - [33:44] Stagehand is an open-source AI web browsing framework that exposes act, extract, and observe APIs, using LLMs to generate Playwright code.
  - [36:23] Stagehand is MIT licensed and brings-your-own-LLM; Browserbase only monetizes if you run it on their browser infrastructure.
- **[38:37] Open Operator**
  - [38:54] Q: What does Paul Klein think of OpenAI's Operator?
  - [40:28] Paul Klein doubts OpenAI will release Operator as an API due to brand risks associated with CAPTCHA solving.
  - [41:24] Open Operator is a reference project demonstrating how to build an AI agent using Browserbase and Stagehand, not a product push.
- **[43:18] Use Cases**
  - [43:41] Browserbase's main use case is browser automation, not web scraping; they are too expensive for bulk scraping.
  - [44:41] Beni uses Browserbase to automate food stamp rebate submissions by filling government forms, a surprising but impactful use case.
- **[46:05] Market & Future**
  - [46:05] Paul Klein believes browser automation is complex enough to remain a separate primitive, even as conflicting tools emerge.
  - [48:05] "Do we really need to run an entire operating system just to control a browser? I don't think that's necessary."
  - [50:30] Paul Klein envisions a future where software uses software: clicking a button in one app triggers actions across multiple websites using AI agents.
  - [52:29] Paul Klein predicts Browserbase will become a billion-dollar company within five years.
- **[52:46] Solo Founder**
  - [53:11] Paul Klein argues solo founders can move faster by avoiding co-founder alignment overhead, especially in dev tools.
  - [55:53] Browserbase has a five-day in-office culture with 10 AM–6 PM hours, avoiding the excesses of 9-9-6 schedules.
  - [56:50] Browserbase hires ex-YC CTOs and future founders, finding that they thrive at a company with product-market fit.
- **[59:37] Dream AI**
  - [59:51] Paul Klein would mine local city hall meeting recordings to predict new commercial developments and invest in real estate.

## Speakers

- **Swyx** (host)
- **Alessio** (guest)
- **Paul Klein IV** (guest)

## Topics

Agent Infrastructure, Cloud

## Mentioned

Amazon (company), Beni (company), BrowserBase (company), Mux (company), OpenAI (company), Stream Club (company), Twilio (company), Chrome (product), Chromium (product), Docker (product), Fargate (product), Firecracker (product), Kubernetes (product), Open Operator (product), Operator (product), Pig.dev (product), Playwright (product), Puppeteer (product), Selenium (product), Stagehand (product)

## Transcript

### Intro

**Alessio** [0:05]
Hey, everyone. Welcome to the Latest Space Podcast. This is Alessio, partner and CTO at Decibel Partners, and I'm joined by my co-host Swyx, founder of Small AI.

**Swyx** [0:13]
Hey, and today we are very blessed to have our friend, uh, Paul Klein IV.

**Alessio** [0:19]
The Fourth.

**Swyx** [0:19]
The Fourth.

**Alessio** [0:20]
Uh, CEO of Browserbase. Welcome.

**Paul Klein IV** [0:22]
Thanks, guys. Yeah, I'm happy to be here. I've been lucky to know both of you for, like, couple years now, I think. So-

**Swyx** [0:27]
Yeah

**Paul Klein IV** [0:27]
... it's just like we're hanging out, you know-

**Swyx** [0:29]
Just hanging out with mics in front of us

**Paul Klein IV** [0:30]
... with three ginormous microphones in front of our face. This is totally normal hangout.

**Swyx** [0:35]
Yeah. Uh, we've actually mentioned you on the podcast, I think, more often than any other Solaris tenant, just because, like, you're one of the, you know, best performing, I think, LLM tool companies, uh, that have started up in, you know, in the last couple years.

**Paul Klein IV** [0:50]
Yeah, I mean, it's been a whirlwind of a year. Like, Browserbase is actually pretty close to our first birthday, so we are one years old. And going from, you know, starting a company as a solo founder to, you know, having a team of 20 people, you know, a Series A, but also being able to support hundreds of AI companies that are building AI applications that go out and automate the web, it's just been, like, really cool.

It's been happening a little too fast. I think, like, collectively as an AI industry, let's just take a week off together. I took my first vacation actually two weeks ago, and Operator came out on the first day, and then a week later, later DeepSeek came out, and I'm, like, on vacation trying to chill.

I'm like, "We gotta build with this stuff," right? So it's, uh, it's been a breakneck year, but I'm super happy to be here and, like, talk more about all the stuff we're seeing, and I'd love to hear kind of what you guys are excited about too and share with it, you know.

**Swyx** [1:40]
Where to start? So, uh, p- people, you've done a bunch of podcasts. I, I think I'd strongly recommend Jack Bridger's Scaling Dev Tools, as well as Turner Novak's The Peel and, and, you know, I'm, I'm sure, I'm sure there's others.

So you covered your Twilio story in the past. Talked about Stream Club, and you got ac-acquired to Mux, and then you left to start Browserbase. So maybe we just start with what, what is Browserbase?

### What is Browserbase

**Paul Klein IV** [2:00]
Yeah. So Browserbase is the web browser for your AI. We're building headless browser infrastructure, which are browsers that run in a server environment that's accessible to developers via APIs and SDKs. It's really hard to run a web browser in the cloud.

You guys are probably running Chrome on your computers, and that's using a lot of resources, right?

**Swyx** [2:21]
Yeah.

**Paul Klein IV** [2:21]
So if you wanna run a web browser or thousands of web browsers, you can't just spin up a bunch of Lambdas. You actually need to use a secure containerized environment. You have to scale it up and down. It's a stateful system, and that infrastructure is, like, super painful, and I know that firsthand 'cause at my last company, Stream Club, I was CTO, and I was building our own internal headless browser infrastructure.

That's actually why we sold the company is because Mux really wanted to buy our headless browser infrastructure that we'd built. And it's just a super hard problem, and I actually told my co-founders I would never start another company unless it was a browser infrastructure company.

And turns out that's really necessary in the age of AI when AI can actually go out and interact with websites, click on buttons, fill in forms. You need AI to do all of that work in an actual browser running somewhere on a server, and Browserbase powers that.

**Swyx** [3:09]
While you're talking about it, uh, it occurred to me, not that you're, like, gonna be acquired or anything, but it oc- occurred to me that it would be really funny if you became, like, the Nikita Beer of browser- ...

headless browser companies. Y- like, you just have one trick, and you just, you make browser companies that, that get acquired.

**Paul Klein IV** [3:24]
I, I truly do only have one trick. And like I, I'm screwed if it's not for headless browsers, you know? Like I, um, I'm not a Go programmer, you know. I, I'm in AI Grant, you know, Browsers was in AI Grant.

**Swyx** [3:36]
Yes.

**Paul Klein IV** [3:36]
But we were the only company in that AI Grant batch that used $0 on AI spend. You know, we're purely an infrastructure company.

**Swyx** [3:43]
Ah. Mm.

**Paul Klein IV** [3:43]
So as much as people wanna ask me about reinforcement learning, I might not be the best guy to talk about that, but if you wanna ask about headless browser infrastructure at scale, I can talk your ear off. So that's, um, really my area of expertise, and, um, it's a pretty niche thing.

Like, nobody has done what we're doing at scale before, so-

**Swyx** [3:59]
Yeah

**Paul Klein IV** [3:59]
... we're happy to be the experts.

**Swyx** [4:00]
You do have a AI thing, Stagehand, which we'll talk about, but, uh-

**Paul Klein IV** [4:04]
Yeah

**Swyx** [4:04]
... yeah, we can talk about the, the sort of core of Browserbase first, and then maybe Stagehand.

**Paul Klein IV** [4:08]
Yeah, Stagehand's our-

**Swyx** [4:09]
'Cause that's kind of chronologically

**Paul Klein IV** [4:09]
... web browsing framework. Yeah.

**Swyx** [4:11]
Yeah.

**Alessio** [4:11]
Yeah. And maybe how you got to Browserbase and what you, what problems you saw. So one of the first things I worked on as a software engineer was integration testing. Sauce Labs was kinda like the main thing at the time, and then we had Selenium, we had Playwright, we had all these different browser thing, but it's always been super hard to do.

So obviously you worked on this before. When you started Browserbase, what were the AI specific challenges that you saw versus there's kinda like all the usual running browser scale in the cloud, which has been a problem for years.

What are, like, the AI unique things that you saw that, like, traditional purchase just didn't cover?

**Paul Klein IV** [4:47]
Yeah. First and foremost, I think back to, like, the first thing I did as a developer, like, as a kid when I was writing code, I wanted to write code that does, did stuff for me, you know? I wanted to write code to automate my life, and I'd do that probably by using Curl or Beautiful Soup to fetch data from a website and parse that data.

And we all know that now, like, you know, taking HTML and plugging that into an LLM, you can extract insights, you can summarize. So it was very clear that now, like, dynamic web scraping became very possible with the rise of large language models, or a lot easier, and that was, like, a clear reason why there's gonna be more usage of headless browsers, which are necessary because a lot of modern websites don't expose all of their page content via a simple HTTP request, you know?

They actually do require you to run JavaScript on the page to hydrate this. Like, Airbnb is a great example. You go to airbnb.com, a lot of that content on the page isn't there until after they kind of run the initial hydration, so you can't just scrape it with a curl.

You need to have some JavaScript run, and a browser is that JavaScript engine that's gonna actually run all those requests on the page. So web data retrieval was definitely, like, one driver of starting Browserbase and the rise of being able to summarize that with an LLM.

Also, like, I was familiar with, like, if I wanted to automate a website, I could write one script, and that would work for one website. It was very static and deterministic. But the web is non-deterministic. The web is always changing, and until we had LLMs, there was no way to write scripts that could, you know, you could write once that would run on any website, that would change with the structure of the website, click the login button Could be mean something different on many different websites, and LLMs allow us to generate code on the fly to actually control that.

So I think that rise of writing the generic automation scripts that can work on many different websites, to me made it clear that browsers are gonna be a lot more useful because now you can automate a lot more things without writing.

You know, if you wanted to write a script to book a demo website, a demo call on 100 websites, previously you had to write 100 scripts.

**Swyx** [6:47]
Mm-hmm.

**Paul Klein IV** [6:47]
Now you write one script that uses LLMs to generate that script for each website in real time. That's why we build our web browsing framework, Stagehand, which does a lot of that work for you. But those two things, web data collection and then, like, enhanced automation of many different websites, it just felt like big drivers for more browser infrastructure that would be required to power these kind of features.

**Swyx** [7:06]
Yeah. And was multimodality also a big thing? Now you can use the br- you can use the LLM to look even though the text in the dome might not be as friendly.

**Paul Klein IV** [7:13]
Yeah. Maybe my, my hot take is, like, I was always kinda like, I didn't think vision would be as big of a driver for UI automation. I, I felt like s- you know, HTML is structured text and large language models are good with structured text, but it's clear that these computer use models are often vision-driven and, um, they've been really pushing things forward.

So definitely being multimodal, like rendering the page is required to take a screenshot to give that to a computer use model to take actions on a website. I mean, it's just another win for browser, but I'll be honest, that wasn't what I was thinking early on.

I, I didn't even think that we'd get here so fast with, with multimodal and vision models.

**Swyx** [7:51]
This is one of those things where, uh, I forgot to mention in my intro that I'm an investor in Browserbase.

**Paul Klein IV** [7:56]
Yes.

**Swyx** [7:56]
And I remember that when you, when you pitched, uh, to me, like, a lot of the stuff that we have today, we... like, wasn't on the original conversation. But I, I did have my, my original thesis was, um, something that we've talked about on the podcast before, which is take the GPT store, the custom GPT store, all the, every single checkbox and plugin is effectively a startup, and this was the browser- browser one.

I think the main hesitation, I think I actually took a while to get back to you. The main hesitation was that there were others, like you are not the first headless browser startup. It's not even your first headless browser startup.

There's always the question of, like, will you be the category winner in a place where there's a bunch of incumbents, to be honest, that are bigger than you? They're just not targeted at the AI space. They don't have the backing of Nat Friedman.

You know, there's a bunch of like you, you, you're, you, you're here in S- in Silicon Valley. They, they're not. I don't know if that's, that was it, but, like, that was a-

**Paul Klein IV** [8:50]
Yeah

**Swyx** [8:50]
... interesting barrier.

**Paul Klein IV** [8:51]
I mean, like, I think I tried all the other ones and I was, like, really disappointed. Like, my background is from working at great developer tools companies and nothing had like the Vercel-like experience. Um, like our biggest competitor actually is partly owned by private equity and they just jacked up their prices quite a bit and the dashboard hasn't changed in five years and I actually used them at my last company and, and tried them and I was like, "Oh man," like there really just needs to be something that's like the experience of these great infrastructure companies like Stripe, like Clerk, like Vercel, that I use and love, but oriented towards this kind of like more specific category, which is browser infrastructure, which is really technically complex.

Like a lot of stuff can go wrong on the internet when you're running a browser. The internet is very vast. There's a lot of different configurations. Like there's still websites that only work with Internet Explorer out there. How do you handle that when you're running your own browser infrastructure?

These are the problems that we have to think about and solve at Browserbase and it's, it's certainly a labor of love, but I built this for me first and foremost. I know it's super cheesy and everyone says that for like their startups, but it really truly was for me if you look at like the talks I've done even before Browserbase, and I'm just like really excited to try and build a category-defining infrastructure company and it's, it's rare to have a new category of infrastructure exist.

We're here in the Chroma offices and like, you know, vector databases, uh, is a new category of infrastructure.

**Swyx** [10:11]
Is it?

**Paul Klein IV** [10:12]
Is it? I mean, we can- I, I... We're in their office, so you know, we can, we can debate that one later.

**Swyx** [10:16]
Yes, that is one of the-

**Paul Klein IV** [10:17]
But I-

**Swyx** [10:17]
... the industry debates

**Paul Klein IV** [10:18]
... I guess we go back to the LLMOS talk, uh, that Karpathy gave way long ago, and like the browser box was very clearly there and it seemed like the people who are building in the space also agreed that browsers are a core primitive of infrastructure for the LLMOS that's gonna exist in the future.

**Swyx** [10:34]
Yeah.

**Paul Klein IV** [10:34]
And, um, nobody was building something there that I wanted to use, so I had to go build it myself.

**Swyx** [10:39]
Yeah, I mean, uh, exactly that talk that, that it-- honestly, uh, that diagram, every box is a startup.

**Paul Klein IV** [10:44]
Mm-hmm.

**Swyx** [10:44]
And, uh, there's the code box and then there's the, the browser box. I think at some point they will start clashing.

**Paul Klein IV** [10:50]
Mm-hmm.

**Swyx** [10:50]
Uh, there, there's always the question of do you, are you a point solution or are you the sort of all-in-one? And, uh, I think the point solutions tend to win quickly, but then the all, all-in-ones have, have a very tight cohesive experience.

**Paul Klein IV** [11:01]
Yeah.

**Swyx** [11:02]
Let's talk about just the hard problems of Browserbase. Um, you have on your website, which is beautiful.

### Scaling Challenges

**Paul Klein IV** [11:08]
Thank you.

**Swyx** [11:09]
Was there a agency that you used for that?

**Paul Klein IV** [11:11]
Yeah.

**Swyx** [11:11]
When I, when I shut them up.

**Paul Klein IV** [11:12]
Uh, herb.paris. They're amazing.

**Swyx** [11:13]
Herb.paris.

**Paul Klein IV** [11:13]
Yeah. It's H-E-R-B-E. I, I highly recommend for developer tools founders to work with consumer agencies

**Swyx** [11:19]
Yeah

**Paul Klein IV** [11:19]
... 'cause they end up building beautiful things and the-

**Swyx** [11:21]
And the videos-

**Paul Klein IV** [11:22]
... the Parisians know how to build beautiful interfaces, so I gotta give prep- prop- props to them.

**Swyx** [11:26]
And, and chat apps apparently are, they are very fast.

**Paul Klein IV** [11:28]
Oh, yeah.

**Swyx** [11:29]
Uh, the Mistral chat.

**Paul Klein IV** [11:30]
Yeah.

**Swyx** [11:30]
Yeah, Mistral.

**Paul Klein IV** [11:30]
Le Mistral. Yeah. Le Chat.

**Swyx** [11:32]
Le Chat. And then it all, your videos as well w- was professionally shot, right?

**Paul Klein IV** [11:35]
Yeah

**Swyx** [11:36]
... the Series A video?

**Paul Klein IV** [11:36]
Yeah, yeah, yeah.

**Swyx** [11:36]
Yeah.

**Paul Klein IV** [11:37]
Nick coded the videos. He's amazing.

**Swyx** [11:38]
Not the initial video that you shot at the new office.

**Paul Klein IV** [11:41]
First one was Austin, another, another video for us, but yeah, I mean, like I think when you think about how you talk about your company, you have to think about the way you present yourself. It's, you know, as a developer you think you evaluate a company based on like the API reliability and the P95, but a lot of developers say, "Is the website good?

Is the message clear? Do I like trust this founder I'm building my whole feature on?" So we've tried to nail that as well as like the reliability of the, of the infrastructure. You're right, it's very hard and there's a lot of kinda foot guns that you run into when running headless browsers at scale.

**Swyx** [12:11]
Right. So let's pick one. Uh, you have eight features here, seamless integration, scalability, fast or speed, secure, observable, stealth, that's interesting, extensible and developer first. What comes to your mind as like- The top two, three hardest ones.

**Paul Klein IV** [12:26]
Yeah. I think just running headless browsers at scale is, like, the hardest one.

**Swyx** [12:31]
So scalable.

**Paul Klein IV** [12:31]
And, and maybe... Can I nerd out for a second?

**Swyx** [12:33]
Yeah, go ahead.

**Paul Klein IV** [12:33]
Is that okay? I heard this is a technical audience, so I'll talk to the, the other nerds. Whoa. They were listening.

**Swyx** [12:39]
Yeah, they're upset.

**Paul Klein IV** [12:40]
They're, they're ready. Um-

**Swyx** [12:41]
The AGI is... The AGI is angry.

**Paul Klein IV** [12:44]
Okay. So how do you run, um, a browser in the cloud? Let's start with that, right? So let's say you're using, um, a popular browser automation framework, like Puppeteer, Playwright, and Selenium. Maybe you've written a code, some code locally on your computer that opens up Google, it finds the search bar, and then types in, you know, search for latent space, and hits the search button.

That script works great locally. You can see the little browser open up. You wanna take that to production. You wanna run the script in a cloud environment, so when your laptop is closed, your browser is doing something, the browser is doing something.

Well, we use Amazon at Browserbase. You know, the first thing I'd reach for is probably, like, some sort of serverless, uh, infrastructure. I would probably try and deploy it on Lambda. But Chrome itself is too big to run on a Lambda.

It's over 250 megabytes, so you can't easily start it on a Lambda. So you maybe have to use something like Lambda layers to, to squeeze it in there, maybe use a different Chromium build that's lighter, and you, you get it on Lambda.

Great, it works. But it runs super slowly. This is because Lambdas are very, like, resource limited. They only run, like, with one vCPU. You can run one process at a time. Remember, Chromium is super beefy. It's barely running on my MacBook Air.

Um-

**Swyx** [13:49]
I'm still downloading it from, uh, from-

**Paul Klein IV** [13:51]
From... Yeah. From the test earlier, right? Like, it's-

**Swyx** [13:53]
I'm joking. Yeah, yeah.

**Paul Klein IV** [13:54]
It... But it's big, you know? Um, so, like, Lambda, it, it, it just won't work really well. Maybe it'll work, but you need something faster. Your user wants something faster. Okay. Well, let's put it on a beefier instance.

Let's get an EC2 server running. Let's throw Chromium on there. Great. Okay. I can... That works well with one user, but what if I wanna run, like, 10 Chromium instances, one for each of my users? Okay. Well, I might need two EC2 instances, maybe 10.

All of a sudden, you have multiple EC2 instances. This sounds like a problem for Kubernetes and Docker, right? Now, all of a sudden, you're using ECS or EKS, the Kubernetes or container solutions by Amazon. You're spinning up and down containers, and you're spending a whole engineer's time on kinda maintaining this stateful distributed system.

Those are some of the worst systems to run because when it's a stateful distributed system, it means that you are bound by the connections to that thing. You have to keep the browser open while someone is working with it, right?

That's just a painful architecture to run, and there's all this other little gotchas with Chromium. Like Chromium, which is the open source version of Chrome, by the way, you have to install all these fonts. You want emojis working in your browsers because your vision model is looking for the emoji, you need to make sure you have the emoji fonts.

You need to make sure you have all the right extensions configured. Like, oh, do you want ad blocking? How do you configure that? How do you actually record all these browser sessions? Like, it's a headless browser. You can't look at it, so you need to have some sort of observability.

Maybe you're recording videos and storing those somewhere. It all kind of adds up to be this just giant monster piece of your project when all you wanted to do was run a lot of browsers in production for this little script to go to google.com and search.

And when I see a complex distributed system, I see an opportunity to build a great infrastructure company, and we really abstract that away with Browserbase, where our customers can use these existing frameworks, Playwright, Puppeteer, Selenium, or our own Stagehand, and connect to our browsers in a serverless-like way and control them, and then just disconnect when they're done.

And they don't have to think about the complex distributed system behind all of that. They just get a browser running, you know, anywhere, anytime, uh, really easy to connect to.

**Swyx** [15:56]
I'm sure you have questions. I'll, I'll just... My standard question with anything. So y- you know, essentially you're a serverless browser company, and, uh, there's been other serverless things that I'm, I'm familiar with in the past, serverless GPUs, serverless, I don't know, website hosting.

You know, that, that's, that's where I come from with Netlify. One question is just, like, you know, you promise to spin up thousands of browsers in milliseconds. I feel like there's no real solution that does that yet, and I'm just kinda curious how.

A part... And the only re- the only solution I know, which is to kinda keep a kinda warm pool of servers around, which is expensive, but maybe not so expensive because it's just computer, just CPUs. So I, I'm just like, you know.

**Paul Klein IV** [16:36]
Yeah. You, you nailed it, right?

**Swyx** [16:38]
Okay.

**Paul Klein IV** [16:38]
Like, I mean, like how do you offer, like, a serverless-like experience with something that is clearly not serverless, right? And the answer is you need to be able to run, uh, many browsers on single nodes. We use Kubernetes at Browserbase, so we have, you know, many pods that are being scheduled.

We have to predictably schedule them up or down. Yes, thousands of browsers in milliseconds is the best case scenario. If, if you hit us with 10,000 requests, you may hit us-

**Swyx** [17:00]
You would hit the limit

**Paul Klein IV** [17:00]
... a slower cold start, right?

**Swyx** [17:01]
Yeah, yeah.

**Paul Klein IV** [17:01]
So, um, we've done a lot of work on predictive scaling and being able to kind of route stuff to different regions where, you know, we have multiple regions at Browserbase where we have different pools available. You can also pick the region you wanna go to based on, like, lower latency.

Round trip time latency is very important with these types of things. There's a lot of requests going over the wire. So for us, like, having a VM like Firecracker powering everything under the hood allows us to be super nimble and spin things up, up or down really quickly with strong multi-tenancy.

But in the end, this is, like, the complex infrastructural challenges that we have to kinda deal with at Browserbase. And we have a lot more stuff on our roadmap to allow customers to have more levers to pull to exchange, do you want really fast brow- browser startup times or do you want really low costs?

And if you're willing to be more flexible on that-

**Swyx** [17:42]
Yeah, slider

**Paul Klein IV** [17:42]
... we may be able to kind of, like, work better for your use cases.

**Swyx** [17:45]
You seem to use Firecracker. Shouldn't Fargate do that for you, or did you have to go lower level than that?

**Paul Klein IV** [17:50]
We had to go lower level than that.

**Swyx** [17:51]
I find this a lot with Fargate customers-

**Paul Klein IV** [17:53]
Yeah

**Swyx** [17:53]
... which is alarming for Fargate.

**Paul Klein IV** [17:55]
We used to be a giant Fargate customer. Actually, the first version of Browserbase-

**Swyx** [17:58]
Everyone says this

**Paul Klein IV** [17:59]
... was ECS and Fargate. And I, um... Unfortunately, it's, it's a great product. I think we were actually the largest Fargate customer in our region for a little while.

**Swyx** [18:07]
No. What?

**Paul Klein IV** [18:08]
Yeah. Seriously. And unfortunately, just, it's a great product, but I think if you're an infrastructure company, you actually have to have a deeper level of control over these primitives.

**Swyx** [18:16]
Mm.

**Paul Klein IV** [18:16]
I think it's the same thing is true with databases. You know, um, we've used other database providers, and I think, you know-

**Swyx** [18:22]
Yeah, serverless Postgres.

**Paul Klein IV** [18:23]
Yeah. When, when-

**Swyx** [18:23]
Shocker

**Paul Klein IV** [18:24]
... when you're an infrastructure company, you're on the hook if any provider has an outage. And I, I can't tell my customers, like, "Hey," like, "we went down because so and so went down." That's not acceptable. So- For us, we've really moved to bringing things internally.

It's kind of opposite of what we preach. We tell our customers, "Don't build this in-house." But then we're like, "We build a lot of stuff in-house." But I think it just really depends on what is in the critical path.

We try and have deep ownership of that.

**Alessio** [18:47]
On the distributed location side, how does that work for the web, where you might get served different content in different locations, but the customer is expecting, you know, if you're in the US, I'm expecting the US version, but if you're spinning up my browser in France, I'm gonna get the French version?

**Paul Klein IV** [19:03]
Yeah. Yeah. It's a g- it's a good question. Well, generally, like, on the localization, there is a thing called locale in the browser. You can set, like, what your locale is if you're, like, in the EN US browser or not.

But some things do IP, um, loca- IP-based routing, and in that case, you may wanna have a proxy. Like, let's say you're running something in the, in Europe, but you wanna make sure you're showing up from the US.

You may wanna use one of our proxy features. So you can turn on proxies to say, like, "Make sure these connections always come from the United States," which is necessary too 'cause when you're browsing the web, you're coming from, like, a, you know, data center IP, and that can make things a lot harder to browse web.

So we do have kinda like this proxy super network where we'll pick the right proxy for you based on where you're going so you can reliably automate the web. But if you get scheduled in Europe, that doesn't happen, especially if you try and schedule you as close to, you know, your origin that you're trying to go to.

But generally, you have control over the regions you can put your browsers in. So you can specify West One or East One or Europe. We only have one region in Europe right now, actually.

**Alessio** [19:56]
Yeah. What's harder, the browser or the proxy? I feel like to me, it feels like actually proxying reliably at scale is much harder than spinning up browsers at scale. I'm curious.

### Proxies & CAPTCHAs

**Paul Klein IV** [20:06]
It's all hard. It's layers of hard, right?

**Alessio** [20:08]
Well, yeah, yeah, yeah. Of course.

**Paul Klein IV** [20:08]
I think it's different, different levels of hard. I think the, the thing with the, you know, the proxy infrastructure is that we work with, you know, many different, uh, web pop proxy providers, and some are better than others.

Some have good days, some have bad days. And our customers who've built, you know, browser infrastructure on their own, they went-- they have to go and deal with sketchy actors. Like, first, they figure out their own browser infrastructure, and then they gotta go buy a proxy.

And it's like you can pay in Bitcoin, and it just kinda feels a little sus, right? It's like you're, you're buying drugs when you're trying to get a proxy online. We have, like, deep relationships with these counterparties. We're able to audit them and say, "Is this proxy being sourced ethically?"

Like, it's not running on someone's TV somewhere.

**Alessio** [20:46]
Is it free range? Like what-

**Paul Klein IV** [20:47]
Yeah. Free range organic proxies, right?

**Alessio** [20:49]
Right.

**Paul Klein IV** [20:49]
We do a level of diligence. We're SOC 2s, so we have to understand what is going on here. But then we're able to make sure that, like, we route around proxy providers not working. There's proxy providers who will just...

the proxy will stop working all of a sudden. And then if you don't have redundant proxying on your own browsers, that's hard down for you, or you may get some serious impacts there. With us, like, we intelligently know, "Hey, this proxy's not working.

Let's go to this one." And you can kinda build a network of multiple providers to really guarantee, you know, the best uptime for our customers.

**Alessio** [21:15]
Yeah. So you don't own any proxies-

**Paul Klein IV** [21:18]
We don't own any proxies

**Alessio** [21:18]
... you're providing.

**Paul Klein IV** [21:19]
The team has been saying, "Who wants to, like, take home a little proxy server?" But I-- not yet. We're, we're not there yet, you know?

**Swyx** [21:26]
It's a very mature market. I don't think you should build that yourself. Like, you, you should just be a super customer of them.

**Paul Klein IV** [21:31]
Yeah.

**Swyx** [21:31]
Scraping, I think, is, is, is the main use case for that. I guess we'll-- that leads us into CAPTCHAs.

**Paul Klein IV** [21:36]
Mm-hmm.

**Swyx** [21:37]
And also auth, but let's, let's talk about CAPTCHAs. Uh, you, you had a little spiel that you wanted to, to talk about CAPTCHA stuff.

**Paul Klein IV** [21:43]
Oh, yeah. I, I was just... I'm-- I think a lot of people ask, if they're thinking about proxies, they're thinking about CAPTCHAs too. And I think it's the same thing. You can go buy CAPTCHA solvers online, but it's the same buying experience.

It's some sketchy website. You have to integrate it. Like, it's, it's not fun to buy these things. And, and you can't really trust that the docs are bad. What Browserbase does is we integrate, um, a bunch of different CAPTCHA providers.

We do some stuff in-house, but generally, we just integrate with a bunch of known vendors and continually monitor and maintain these things and say, "Is this working or not?" Like, "Can we route around it or not?" Um, and really try-

**Swyx** [22:16]
These are CAPTCHA solvers.

**Paul Klein IV** [22:16]
CAPTCHA solvers. Yeah.

**Swyx** [22:17]
Not CAPTCHA providers, CAPTCHA solvers.

**Paul Klein IV** [22:18]
Yeah.

**Swyx** [22:18]
Sorry. CAPTCHA solvers.

**Paul Klein IV** [22:19]
And, and we really try and make sure all of that works for you, you know? Like, I think as a dev, if I'm buying infrastructure, I want it all to work all the time, and it's important for us to kinda provide that experience by making sure everything does work and monitoring it on our own.

**Swyx** [22:33]
Yeah.

**Paul Klein IV** [22:33]
And, and right now, the world of CAPTCHAs is, is tricky. I think AI agents, in particular, are very much ahead of the internet infrastructure. You know, CAPTCHAs are designed to block all types of bots, but there are now good bots and bad bots.

And I think in the future, CAPTCHAs will be able to identify who a good bot is, hopefully via some sort of, like, KYC. For us, like, we've been very lucky. We, we have very little to no known abuse of Browserbase because we're pretty s-- we, we really look into who we, like, work with.

And, you know, for certain types of CAPTCHA solving, we only allow them on certain types of plans because we wanna make sure that we can know what people are doing, what their use cases are. And that's really allowed us to try and be an arbiter of good bots, which is our long-term goal.

Like, I wanna build great relationships with people like Cloudflare so we can agree, "Hey, here are these acceptable bots. We'll identify them for you and make sure we flag when they come to your website. This is a good bot," you know?

**Swyx** [23:23]
I see.

**Alessio** [23:24]
And Cloudflare said they wanna do more of this, so they're gonna set by default if they think you're an AI bot, they're gonna reject. I'm curious if you think this is something that is gonna be at the browser level or, I mean, the DNS level with Cloudflare seems more where it should belong, but I'm curious how you think about it.

**Paul Klein IV** [23:41]
I think the web's gonna change, you know? I think that the internet as we have it right now is gonna change, and we all need to just accept that, that the cat is out of the bag. And instead of kind of like wishing the internet was like it was in the 2000s, where you can have free content line that wouldn't be scraped, it's just, it's not gonna happen.

And instead, we should think about, like, one, how can we change the models of, you know, information being published online so people can, th- you know, adequately commercialize it? But two, how do we rebuild applications that expect that AI agents are gonna log in on their behalf?

Those are the things that are gonna allow us to kind of like identify good and bad bots, and I think the team at Clerk has been doing a really good job with this on the authentication side. I actually think that auth is the biggest thing that will prevent agents from accessing stuff, not CAPTCHAs.

And I think there will be agent auth in the future. I don't know if it's gonna happen from an individual company, but actually authentication providers that have a, you know, hidden login as agent feature, which will then, you put in your email, you'll get a push notification to say like, "Hey, your browser-based agent wants to log into your Airbnb."

You can approve that, and then the agent can proceed. That really circumvents the need for CAPTCHAs or logging in as you and sharing your password. Um, I think agent auth is gonna be one way we identify good bots going forward.

And I think a lot of this CAPTCHA-solving stuff is really short-term problems as the internet kinda reorients itself around how it's gonna work with agents browsing the web just like people do.

**Swyx** [24:59]
Yeah. Uh, Stitch rec- recently was on Hacker News for talking about agent experience, AX, which, uh, is a thing that Netlify is also trying to clone and, and coin and talk about. And we've talked about this on our previous episodes before in a sense that I, I actually think that's, like, maybe the only part of the tech stack that needs to be kind of reinvented for agents.

### Agent Identity

**Swyx** [25:20]
Everything else can stay the same, CLIs, APIs, whatever. But auth, yeah, we, we need, we need agent auth. And it's, it's mostly, like, short-lived. Like, it should not... It, it should be a dis- distinct identity from the human, but paired.

I almost think, like, in the same way that every social network should have your main profile and then your alt accounts or your, or your Finsta, it's almost like, you know, every, every, uh, human token should be paired with a agent token, and the agent token can go and do stuff on behalf of the human token but not be presumed to be the human.

**Paul Klein IV** [25:48]
Yeah, it's like, it's, it's actually very similar to OAuth is what I'm thinking. And, and, you know, Reid from Stitch is an investor, Colin from Clerk, Octaventures, all investors in Browserbase because, like, I hope they solve this 'cause it'll make Browserbase's mission more possible so we don't have to overcome all these hurdles.

But I think it'll be an OAuth-like flow where an agent will ask to log in as you. You'll approve the scopes. Like, it can book a apartment on Airbnb, but it can't, like, message anybody. And then, you know, the agent will have some sort of, like, role-

**Swyx** [26:14]
Yeah

**Paul Klein IV** [26:14]
... based access control within an application.

**Swyx** [26:16]
Yeah.

**Paul Klein IV** [26:16]
I'm excited for that.

**Swyx** [26:17]
The tricky part is just there's one more layer of delegation here, which is like you're authing my user's user or something like that. I don't know if that's tricky or not. Does, does that make sense?

**Paul Klein IV** [26:25]
Yeah. You know, actually, at Twilio, I worked on the login-

**Swyx** [26:28]
Identity, yeah

**Paul Klein IV** [26:29]
... identity and access management teams, right? So, like, I built Twilio's login page SSO-

**Swyx** [26:32]
You were the intern on that team, and then you became the lead in two years?

**Paul Klein IV** [26:35]
Yeah. Yeah, I started as an intern in 2016, then I was the tech lead of that team in two years later.

**Swyx** [26:38]
How? That's not normal.

**Paul Klein IV** [26:40]
I didn't have a life.

**Swyx** [26:41]
He's not normal.

**Paul Klein IV** [26:42]
Yeah.

**Swyx** [26:42]
Look at this guy.

**Paul Klein IV** [26:42]
I didn't have a girlfriend.

**Swyx** [26:43]
Nothing about him is normal.

**Paul Klein IV** [26:44]
I just loved my job. I don't know. I applied to 500 internships for my first job.

**Swyx** [26:48]
Okay.

**Paul Klein IV** [26:48]
And I got rejected from every single one of them except for Twilio and then eventually Amazon. And, um, they took a shot on me, and, like, I was getting paid money to write code, which was my dream.

**Swyx** [26:58]
Yeah .

**Paul Klein IV** [26:58]
I'm, I'm very lucky that, like, this coding thing worked out 'cause I was gonna be doing it regardless. And yeah, I was able to kinda spend a lot of time on a team that was growing at a company that was growing.

So, and it informed a lot of this stuff here. I think these are problems that have been solved with, like, the S- SAML protocol, with SSO. I think it's a relationship with, like, WebAuthn, like these different types of authentication, like, schemes that you can use to authenticate people.

The tooling is all there. It just needs to be tweaked a little bit to work for agents. And I think the fact that there are companies that are already providing authentication as a service really sets it up well.

The thing that's hard is, like, reinventing the internet for agents. We don't wanna rebuild the internet. That's an impossible task. And I think people often say, like, "Well, we'll have this second layer of APIs built for agents." I'm like, we will for the top use cases, but instead if we can just tweak the internet as is, which is on the authentication side, I think we're gonna be the dumb ones going forward, unfortunately.

I think AI's gonna be able to do a lot of the tasks that we do online, which means that it will be able to go to websites, click buttons on our behalf, and log in on our behalf too.

So with this kind of, like, web agent future happening, I think with some small structural changes, like you said, it, it feels like it could all slot in really nicely with the existing internet.

**Swyx** [28:09]
There's one more thing, which is your live view iframe, which lets you take, take control.

### Live View

**Paul Klein IV** [28:13]
Yeah.

**Swyx** [28:14]
Obviously very key for Operator now. But, like, was... Is there anything interesting technically there or that the pe- like, well, people always want this.

**Paul Klein IV** [28:22]
It was really hard to build , you know? Like, so okay, headless browsers, you don't see them, right? They're running in a cloud somewhere. You can't, like, look at them. That's why I really make... It's a weird name.

I wish we came up with a better name for this thing, but you can't see them, right? But customers don't trust AI agents, right? At least the first pass. So what we do with our live view is that, you know, when you use Browserbase, you can actually embed a live view of the browser running in the cloud for your customer to see it working.

And that's what the first reason is, to build trust. Like, okay, so I have this script that's gonna go automate a website. I can embed it into my web application via an iframe, and my customer can watch that thing go.

And then we added two-way communication. So now not only can you watch the browser kinda being operated by AI, if you wanna pause and actually click around, type within this iframe that's controlling a browser, that's also possible. And this is all thanks to some of the lower level protocol, which is called the Chrome DevTools Protocol.

It has a API called startScreenCast, and you can also send mouse clicks and button clicks to a remote browser. And this is all embeddable within iframes. You have a browser within a browser, yo.

**Swyx** [29:24]
And then you simulate the screen, the click on the other side.

**Paul Klein IV** [29:28]
Exactly. And this is really nice often for, like, let's say, a CAPTCHA that can't be solved. You saw this with Operator. You know, Operator actually uses a different approach. They use VNC. So, you know, you're able to see, like, you're seeing the whole window here.

What we're doing is something a little lower level with the Chrome DevTools Protocol. It's just PNGs being streamed over the wire. But the same thing is true, right? Like, hey, I'm running a window. Pause. Can you do something in this window, human?

Okay, great. Resume. Like, sometimes 2FA tokens, like if you get that text message, you might need a person to type that in. Web agents need human-in-the-loop type workflows still. You still need a person to interact with the browser, and building a UI to proxy that is kinda hard.

You may as well just show them the whole browser and say, "Hey, can you finish this up for me?" And then let the AI proceed on afterwards.

**Swyx** [30:10]
Is there a future where I stream my current desktop to Browserbase?

**Paul Klein IV** [30:14]
I don't think so. I, I think we're very much cloud infrastructure.

**Swyx** [30:17]
Server.

**Paul Klein IV** [30:18]
Yeah, you know. But I think a lot of the stuff we're doing, we do wanna, like, build tools. Like, you know, we'll talk about the Stagehand, uh, you know, web agent framework in a second. But, like, there's a case where a lot of people are going desktop first for, you know, consumer use, and I think Claude is doing a lot of this, where I expect to see, you know...

You know, MCP is really oriented around the Claude desktop app for a reason, right? Like, I think a lot of these tools are gonna run on your computer because it makes-

**Swyx** [30:39]
I think it's breaking out.

**Paul Klein IV** [30:40]
Yeah.

**Swyx** [30:40]
People are putting it on a server.

**Paul Klein IV** [30:41]
Oh, really? Okay. Well, sweet. We'll see. We'll see that. But-

**Swyx** [30:43]
I was surprised though.

**Paul Klein IV** [30:45]
I, I, I, I think that, like- The browser company too, with Dia browser. It runs on your machine, you know. It's going to be-

**Swyx** [30:51]
What is it?

**Paul Klein IV** [30:52]
So Dia browser, uh, as far as I understand, I-

**Swyx** [30:54]
I used to use Arc and then-

**Paul Klein IV** [30:55]
Yeah

**Swyx** [30:55]
... it died.

**Paul Klein IV** [30:55]
I haven't used Arc, but I'm a big fan of the browser company. I think they're doing a lot of cool stuff in consumer. As far as I understand, it's a browser where you have a sidebar where you can, like, chat with it, and it can control the local browser on your machine.

So if you imagine, like, what a consumer web agent is, which it lives alongside your browser, I think Google Chrome has Project Marina, I think. I almost call it, uh, Project Marinara for some reason. I don't know why.

It's-

**Swyx** [31:17]
No, I think it's, uh-

**Paul Klein IV** [31:18]
Yeah

**Swyx** [31:18]
... someone really likes the Waterworld.

**Paul Klein IV** [31:20]
Oh, 'cause I see.

**Swyx** [31:21]
The, the classic Kevin Costner...

**Paul Klein IV** [31:22]
Yeah. Okay. Project Mariner is a similar thing to the Dia browser in my mind, as far as I understand it. You have a browser that has an AI interface that will take over your mouse and keyboard and control the browser for you.

Great for consumer use cases, but if you're building applications that rely on a browser and it's more part of a greater, like, AI app experience, you probably need something that's more like infrastructure, not a consumer app.

**Swyx** [31:44]
Just because I, I have explored a little bit in this area, do people want branching? So I have the state of whatever my browser's in, and then I want, like, 100 clones of this state. Do, do people do that or...?

**Paul Klein IV** [31:57]
People don't do it currently.

**Swyx** [31:58]
Yeah.

**Paul Klein IV** [31:59]
But it's definitely something worth thinking about. I think the idea of forking a browser is really cool. Technically kind of hard. We're starting to see this in code execution, where people are, like, forking some, like, code execution, like, processes or forking some tool calls or branching tool calls.

Haven't seen it at the browser level yet, but it makes sense. Like, if an AI agent is, like, using a website and it's not sure what path it wants to take to c-crawl this website to find the information it's looking for, it would make sense for it to explore both paths in parallel, and that'd be a very, like-

**Swyx** [32:27]
A road not taken.

**Paul Klein IV** [32:28]
Yeah. Yeah. Um, and, and hopefully find the right answer and then say, "Okay, this was actually-

**Swyx** [32:32]
Yeah

**Paul Klein IV** [32:32]
... the right one," and memorize that and go there in the future. On the roadmap, for sure.

**Swyx** [32:36]
Well-

**Paul Klein IV** [32:36]
Don't mock my roadmap please, you know?

**Swyx** [32:38]
How do you actually do that?

**Paul Klein IV** [32:39]
Yeah.

**Swyx** [32:39]
How, how do you fork? I feel like the browser is so stateful for so many things.

**Paul Klein IV** [32:43]
Serializes state, restores state, I don't know.

**Swyx** [32:45]
Well, but yeah.

**Paul Klein IV** [32:45]
So it's one of the reasons why we haven't done it yet. It's hard, you know? Like, to truly fork, it's actually quite difficult. The naive way is to open the same page in a new tab and then, like, hope-

**Swyx** [32:54]
Oh, yeah

**Paul Klein IV** [32:54]
... that it's at the same thing. But if you have a form halfway filled, you may have to, like, take the whole, you know, container, pause it, all the memory, duplicate it, restart it from there. It could be very slow.

So we haven't found a thing... Like, the easy thing to fork is just, like, copy the page object-

**Swyx** [33:08]
The whole, yeah

**Paul Klein IV** [33:08]
... you know. But I think there needs to be something a little bit more robust there, so.

**Swyx** [33:12]
Yeah. So, uh, Moreflabs has this infinite branch thing.

**Paul Klein IV** [33:15]
Moreflabs, yeah. Exactly.

**Swyx** [33:16]
They, they wrote a custom fork of Linux or something that let them save the system state, including it.

**Paul Klein IV** [33:22]
Moreflabs, hit me up, I'll be a customer.

**Swyx** [33:24]
Yeah, that's, that's gonna be... Uh, that's, I think that's the only way to do it.

**Paul Klein IV** [33:27]
Yeah.

**Swyx** [33:27]
Like, unless Chrome has some special API for you.

**Paul Klein IV** [33:29]
Yeah, that's probably s-something we'll reverse engineer one day. I don't know.

**Swyx** [33:32]
Yeah. Uh, let's talk about Stagehand, the AI web browsing framework. You have three core components: observe, extract, and act. Pretty clean landing page. What was the idea behind making a framework?

**Paul Klein IV** [33:44]
Yeah. So there's three frameworks that are very popular already exist, right? Puppeteer, Playwright, Selenium. Those are for building hard-coded scripts to control websites, and as soon as I started to play with LLMs plus browsing, I caught myself, you know, code genning Playwright code to control a website.

### Stagehand

**Paul Klein IV** [34:03]
I would, like, take the DOM, I'd pass it to an LLM. I'd say, "Can you generate the Playwright code to click the appropriate button here?" And it would do that. And I was like, "This sh- really should be part of the frameworks themselves."

And I became really obsessed with SDKs that take natural language as part of, like, the API input, and that's what Stagehand is. Stagehand exposes three APIs, and it's a super set of Playwright. So if you go to a page, you may wanna take an action, click on the button, fill in the form, et cetera.

That's what the act command is for. You may wanna extract some data. This one takes a natural language, like extract the winner of the Super Bowl from this page. You can give it a Zod schema so it returns a, a structured output.

And then maybe you, you're building an agent loop and you wanna kind of see what actions are possible on this page before taking one. You can do observe. So you can observe the actions on the page, and it will generate a list of actions.

You can guide it, like, "Give me actions on this page related to buying an item." And you can, like, buy it now, add to cart, view shipping options, and pass that to an LLM, an agent loop, to say, "What's the appropriate action given this high level goal?"

So Stagehand isn't a web agent. It's a framework for building web agents, and we think that agent loops are actually pretty close to the application layer because every application probably has different goals or different ways it wants to take steps.

I don't think I've seen a generic, and maybe you guys are the experts here. I haven't seen, like, a really good AI agent framework here. Everyone kind of has their own special sauce, right? I see a lot of developers building their own agent loops, and they're using tools, and I view Stagehand as the browser tool.

So we expose act, extract, observe. Your agent can call these tools, and from that you don't have to worry about generating Playwright code performantly. You don't have to worry about running it. You can kind of just integrate these three tool calls into your agent loop and reliably automate the web.

**Swyx** [35:49]
A special shout-out to Anirud, who I met at your dinner, uh, who I think listens to the pod. So-

**Paul Klein IV** [35:53]
Yeah

**Swyx** [35:54]
... hey, Anirud.

**Paul Klein IV** [35:54]
Ani's the man. He's a Stagehand guy, you know.

**Swyx** [35:57]
I mean, the interesting thing about each of these APIs is they're kind of each a startup. Like specifically extract, you know, Firecrawl is extr-extract. There's, like, Expand AI. There's a whole bunch of, like, extract companies. They just focus on extract.

I'm curious, like, I feel like you guys are gonna collide at some point. Like right now it's friendly. Everyone's, everyone's in a blue ocean. At some point it's gonna be valuable enough that there's some turf battle here. I don't think you have a dog in the fight.

I think you can, you can mock extract to use an, an external service if, if they're better at it than you. But it's just an observation that, like, I, I... In the same way that I see each option, each checkbox in the set of custom GPTs becoming a startup or each box in the Karpathy chart being a startup, like, th-this is also becoming a thing.

**Paul Klein IV** [36:41]
Yeah, I mean, like, so the way Stagehand works is, uh, it's MIT licensed, completely open source. You bring your own API key to your LLM of choice. You could choose your LLM. We don't make any money off of the extract or-

**Swyx** [36:54]
Yeah

**Paul Klein IV** [36:54]
... really... We, we only really make money if you choose to run it with our browser. You don't have to. You can actually use your own browser, a local browser. Um, you know, Stagehand is completely open source for that reason.

And yeah, like, I think if you're building really complex web scraping workflows, I don't know if Stagehand is the tool for you. I, I think it's really more if you're building an AI agent that needs a few general tools or if it's doing a lot of, like, web automation intensive work.

But if you're building a scraping company, Stagehand's not your thing. You probably want something that's gonna, like- Get HTML content, you know, convert that to markdown, query it. That's not what Stagehand does. Stagehand's more, more about reliability. I think we focus a lot on reliability and less so on cost optimization and speed at this point.

**Swyx** [37:34]
I actually hear, like, Stagehand's-- So the way, the way that Stagehand works, like, it's like, you know, page.act, click on the quick start, right? It's kind of the integration test for the code that you would have to write anyway, like the Puppeteer code that you have to write anyway.

And when the page structure changes, 'cause it always does, then this is still the test. This is still the test that I would have to write.

**Paul Klein IV** [37:53]
Yeah.

**Swyx** [37:53]
So it's kind of like a testing framework that doesn't need implementation detail.

**Paul Klein IV** [37:56]
Well, yeah. I mean, Puppeteer, Playwright, and Selenium were all designed as testing frameworks, right?

**Swyx** [38:00]
Yeah, yeah.

**Paul Klein IV** [38:00]
And now people are, like, hacking them together to automate the web. I would say, and, like, maybe this is, like, me being too specific, but, like, when I write tests, if the page structure changes without me knowing, I want that test to fail.

So I don't know-

**Swyx** [38:12]
Oh

**Paul Klein IV** [38:12]
... if, like, AI, like, regenerating that. Like, people are using Stagehand for testing, um, but it's more for, like, usability testing, not like, uh-

**Swyx** [38:20]
Oh, okay

**Paul Klein IV** [38:20]
... testing of, like, does the front end, like, has it changed or not? So.

**Swyx** [38:24]
Okay, okay.

**Paul Klein IV** [38:24]
But generally, where we've seen people, like, really, like, take off is, like, if they're using, you know, something. If they wanna build a feature in their application that's kinda like Operator or Deep Research, they're using Stagehand to kinda power that tool calling in their own agent loop.

**Swyx** [38:37]
Okay, cool. So let's go into Operator, the first big agent launch of the year from OpenAI. Seems like they have a whole bunch scheduled. Um, you were on break, and your phone blew up. What's your just general view of computer use agents, is what they're calling it, the overall category before we go into Open Operator, just the overall promise of Operator?

### Open Operator

**Swyx** [38:54]
I will observe that I tried it once. It was okay, and I never tried it again.

**Paul Klein IV** [38:59]
That tracks with my experience too. Like, I'm a huge fan of the OpenAI team. Like, I, I, I think that I do not view Operator as the company killer for Browserbase at all. I think it actually shows people what's possible.

I think, like, computer use models make a lot of sense, and I'm actually most excited about computer use models is, like, their ability to, like, really take screenshots and reasoning and output steps. I think that using mouse click and mouse coordinates, I've seen that prove to be less reliable than I would like, and I just wonder if that's the right form factor.

What we've done with our framework is anchor it to the DOM itself, anchor it to the actual item, so, like, if it's clicking on something, it's clicking on that thing, you know? Like, it's more accurate.

**Swyx** [39:39]
No matter where it is.

**Paul Klein IV** [39:40]
Yeah, exactly-

**Swyx** [39:40]
Yeah

**Paul Klein IV** [39:40]
... 'cause it really ties in nicely, and it can handle, like, the whole viewport in one go, whereas, like, Operator can only handle what it sees. But-

**Swyx** [39:47]
Can you hover? Is hovering a thing that you can do?

**Paul Klein IV** [39:50]
I don't know if we expose it as a tool directly, but I'm sure there's, like, an API for hovering.

**Swyx** [39:53]
Okay.

**Paul Klein IV** [39:53]
Like, move mouse to this position.

**Swyx** [39:55]
Yeah, yeah, yeah.

**Paul Klein IV** [39:55]
I think you can trigger hover, like, via, like, the JavaScript on a DOM itself. But no, I, I think, like, when we saw computer use, everyone's eyes li-lit up 'cause they realized, like, "Wow," like, "AI is gonna actually automate work for people."

And I think seeing that kinda happen from both of the labs, and I'm sure we're gonna see more labs launch computer use models, I'm excited to see all the stuff that people build with it. I think that I'd love to see computer use power, like, controlling a browser on Browserbase, and I think, like, Open Operator, which was, like, our open source version of OpenAI's Operator, was our first take on, like, how can we integrate these models into Browserbase?

And we handle the infrastructure and let the labs do the models. I don't have a sense that Operator will be released as an API. I don't know. Maybe it will. I'm curious to see how well that works, 'cause I think it's gonna be really hard for a company like OpenAI to do things like support CAPTCHA solving or, like, have proxies.

Like, I think it's hard for them structurally. Imagine this New York Times, h- New York Times headline, "OpenAI CAPTCHA solving." Like, that would be a pretty bad headline.

**Swyx** [40:54]
Mm-hmm.

**Paul Klein IV** [40:54]
This New York Times headline, "Browserbase solves CAPTCHAs." Like-

**Swyx** [40:57]
No one cares. Yeah

**Paul Klein IV** [40:59]
And, like, our... No one cares. Like, and, like, our investors are bored. Like, we're all okay with this, you know? We're, we're building this company knowing that the CAPTCHA solving is short-lived until we figure out how to authenticate good bots.

Um.

**Swyx** [41:10]
Mm.

**Paul Klein IV** [41:10]
I think it's really hard for a company like OpenAI, who has this brand that's so, so good, to balance with, like, the icky parts of web automation, which it can be kind of complex to solve.

**Swyx** [41:19]
I'm sure OpenAI knows who to call when, whenever they need you.

**Paul Klein IV** [41:22]
Yeah, right.

**Swyx** [41:22]
Yeah.

**Paul Klein IV** [41:23]
I'm sure they'll have a great partnership.

**Alessio** [41:24]
And is Open Operator just, like, a marketing thing for you? Like, how do you think about resource allocation? So you can spin this up very quickly, and now there's all this, like, open deep research.

**Paul Klein IV** [41:34]
Yeah.

**Alessio** [41:34]
There's open, open all these things-

**Paul Klein IV** [41:35]
We started it, you know

**Alessio** [41:36]
... that people are building. You're the original open.

**Paul Klein IV** [41:39]
We're the original open Open Operator, you know.

**Alessio** [41:42]
Is it just, "Hey, look, this is a demo, but, like, we'll help you build out an actual product for yourself?" Like, are you interested in going more of a product route? That's kinda the OpenAI way, right? They started as a model provider, and then...

**Paul Klein IV** [41:54]
Yeah. We, we're not interested in going the product route yet. I view Open Operator as a reference project, you know? Let's show people how to build these things using the infrastructure and models that are out there, and that's what it is.

It's like Open Operator is very simple. It's an agent loop. It says, like, take a high-level goal, break it down into steps, use tool calling to accomplish those steps. It takes screenshots and feed those screenshots into an LLM with the step to generate the right action.

It uses Stagehand under the hood to actually execute this action. It doesn't use a computer use model. And it, it, like, has a nice interface using the live view that we talked about, the iframe, to embed that into an application.

So I felt like people on launch day wanted to figure out how to build their own version of this, and we turned that around really quickly to show them, and I hope we do that with other things like deep research.

We don't have a deep research launch yet. I think, um, David from Meomni actually has an amazing open deep research that he launched. It has, like, 10K GitHub stars now, so he's crushing that. But I think if people wanna build these features natively into their application, they need good reference projects, and I think Open Operator is a good example of that.

**Swyx** [42:53]
I don't know. I, actually, I'm actually pretty pr- uh, bullish on API-driven Operator because that's the only way that you can sort of... Like, once it's reliable enough, obviously, and, and now we're nowhere near, but, like, give it five years.

It'll, it'll happen, you know? And then you can sort of spin this up, and browsers are, are working in the background, and you don't necessarily have to know. And it just is booking restaurants for you, whatever. I can definitely see that future happening.

### Use Cases

**Swyx** [43:18]
I had this on the, on the landing page here. This might be a slightly out of order, but, you know, you have, like, sort of three use cases for Browserbase. Open operator or just the operator sort of use case is kind of like the workflow automation use case, and it competes with UI path in the sort of RPA category.

Would you agree with that?

**Paul Klein IV** [43:34]
Yeah, I would agree with that.

**Swyx** [43:35]
And then there's agents we talked about already, and web scraping, which I th- I imagine would be your, the bulk of your, your workload right now, right?

**Paul Klein IV** [43:41]
No, not at all.

**Swyx** [43:42]
It's agents?

**Paul Klein IV** [43:42]
I'd say actually, like, the majority is browser automation. We're-

**Swyx** [43:45]
Okay

**Paul Klein IV** [43:46]
... we're kind of expensive for web scraping . Like, I think that-

**Swyx** [43:48]
I see

**Paul Klein IV** [43:48]
... if you're building a web scraping product, if you need to do occasional web scraping or you have to do web scraping that works every single time, you wanna use Browserbase. But if you're building web scraping workflows, what you should do is have a waterfall.

You should have-

**Swyx** [44:01]
Ah

**Paul Klein IV** [44:01]
... the first request is a curl to the website. See if you can get it without even using a browser. And then the second request may be, like, a scraping-specific API. There's, like, 1,000 scraping APIs out there that you can use to try and get data.

**Swyx** [44:12]
ScrapingBee?

**Paul Klein IV** [44:13]
Like ScrapingBee is a great example, right? Yeah. And then, like, if those two don't work, bring out the heavy hitter. Like, Browserbase will 100% work, right? It will load the page in a real browser, hydrate it.

**Swyx** [44:21]
I see, 'cause they-- a lot of people don't render the JS.

**Paul Klein IV** [44:24]
Yeah.

**Swyx** [44:24]
Okay, cool. I just wanted to get a, a rough sense of that.

**Paul Klein IV** [44:26]
Yeah, exactly. So I mean, the three big use cases, right? Like, you know, automation, we-da- web data collection, and then, you know, if you're building anything agentic that needs, like, a browser tool, you wanna use Browserbase.

**Swyx** [44:36]
Is there any, any use case that, like, you were super surprised by that people might not even think about?

**Paul Klein IV** [44:41]
Oh, yeah.

**Swyx** [44:41]
Or is it... Yeah, anything that you can share?

**Paul Klein IV** [44:43]
The long tail is crazy. Um-

**Swyx** [44:44]
Yeah

**Paul Klein IV** [44:44]
... one of the case studies on our website that I think is the most interesting is this company called Beni. So the way that it works is if you're on food stamps in the United States, you can actually get rebates if you buy certain things.

Like, maybe you buy some vegetables and you submit your receipt to the government. They'll give you a little rebate back, say, "Hey, thanks for buying vegetables. It's good for you." Um- ... that process of submitting that receipt is very painful.

And the way Beni works is you use their app to take a photo of your receipt, and then Beni will go submit that receipt for you and then deposit the money into your account. That's actually using no AI at all.

It's all, like, hard-coded scripts. They maintain the scripts. It's-- they've been doing a great job, and they built this amazing consumer app. But it's an example of, like, all these, like, tedious workflows that people have to do to kinda go about their day-to-day lives.

And I had never known about, like, food stamp rebates or the complex forms you have to do to fill them, but the world is powered by millions and millions of tedious forms, visas. You know, Immigra- Lighthouse is a customer, right?

You know, they do the O-1 visa. Millions and millions of forms are taking away humans' time, and I hope that Browserbase can help power software that automates away the web forms that we don't need anymore.

**Swyx** [45:50]
Yeah, I'm, I mean, I'm very supportive of that. I hate forms. I do think, like, government itself should embrace AI more to do more sort of human-friendly form filling.

**Paul Klein IV** [46:00]
Mm-hmm.

**Swyx** [46:01]
But, uh, I'm, I'm not optimistic. I'm not holding my breath.

**Paul Klein IV** [46:03]
Yeah. We'll see.

**Swyx** [46:05]
Okay. I think I'm about to zoom out. I, I have a little brief thing on computer use, and then we can talk about founder stuff, which is I tend to think of developer tooling markets in impossible triangles, where everyone starts in a niche, and then they start, start to branch out.

### Market & Future

**Swyx** [46:19]
So I already hinted a, a little bit of this, right? We mentioned Morph. We mentioned E2B. Mentioned Firecrawl, uh, and then there's Browserbase. So there's, like, all this stuff of, like, have serverless virtual computer that you give to an agents and let them do stuff with it.

And either... And there, and there's various ways of connecting it to the internet. Uh, you can just connect to a search API, like SERP API, whatever other, uh, Exa is another, another one. That's where you're searching. You can also have a JSON markdown extractor, which, which is Firecrawl.

Or you can have a virtual browser like Browserbase, or you can have a virtual machine like Morph. And then there's also maybe, like, a virtual sort of code environment, like Code Interpreter. So, like, there's just, like, a bunch of different ways to tackle the problem of give a computer to an agent.

And I'm just kinda s- wondering if you see, like, everyone sort of happily coexisting in their, in their respective niches. And as a developer, I just go and pick, like, a shopping bas-basket of one of each. Or do you think that you eventually, people will collide?

**Paul Klein IV** [47:19]
I think that currently it's not a zero-sum market. Like, I think we're talking about all of knowledge work that people do that can be automated online, all of these, like, trillions of hours that happen online where people are working, and I think that there's so much software to be built that, like, I tend not to think about how these companies will collide.

I just try to solve the problem as best as I can and make this specific piece of infrastructure, which I think is an important primitive, the best I possibly can. And yeah, I, I think there's, there's players that are actually gonna launch, like, over-the-top, you know, platforms, like agent platforms that have all these tools built in, right?

Like, who's building the Rippling for agent tools-

**Swyx** [48:01]
Exactly

**Paul Klein IV** [48:01]
... that has the search tool, the browser tool, the operating system tool, right?

**Swyx** [48:05]
There are some. There are some.

**Paul Klein IV** [48:06]
There's some, right? And I think in the end, what I have seen as my time as a developer, and I look at all the favorite tools that I have, is that, like, for tools and primitives with sufficient levels of complexity, you need to have a solution that's really bespoke to that primitive, you know?

And, and I am sufficiently convinced that the browser is complex enough to deserve a primitive. Um, obviously, I have to. I'm the founder of Browserbase, right? I'm talking my book. But, like, I think maybe I can give you one spicy take against, like, maybe just whole envi- OS running.

I think that when I look at computer use when it first came out, I saw that the majority of use cases for computer use were controlling a browser.

**Swyx** [48:46]
Mm.

**Paul Klein IV** [48:46]
And do we really need to run an entire operating system just to control a browser? I don't think that's necessary. You know, Browserbase can run browsers for way cheaper than you can if you're running a full-fledged OS with a GUI, you know, operating system.

And I think that's just an advantage of the browser. It is like, browsers, like, are little OSs, and you can run them very efficiently if you orchestrate it well. And I think that allows us to offer 90% of the, you know, functionality and the platform needed at 10% of the cost of running a full OS.

**Swyx** [49:17]
Yeah. I definitely see the logic in that. There's a Mark Andreessen quote, I don't know if you know of that, this one, where he basically observed that the browser is turning the operating system into a poorly debugged set of device drivers 'cause most of the apps are moved-

**Paul Klein IV** [49:29]
Yeah

**Swyx** [49:29]
... from, from the OS to the browser. So you can just run browsers.

**Paul Klein IV** [49:32]
There's a place for OSs too.

**Swyx** [49:33]
Yeah.

**Paul Klein IV** [49:33]
Like, I think that there are some applications that only run on Windows operating systems, and- Eric from Pig.dev in this la- upcoming YC batch or last YC batch, like, he's building all run tons of Windows operating systems for you to control with your agent.

And, like, there's some legacy EHR systems that only run on Internet Explorer and Windows, and Browserbase doesn't ex- uh, support-

**Swyx** [49:54]
Scott Pigs.

**Paul Klein IV** [49:55]
Yeah, I think that's it. Um, I think, like, there are use cases for specific-

**Swyx** [50:00]
Well-

**Paul Klein IV** [50:00]
... operating systems for specific legacy software, and, like, I'm excited to see what he does with that.

**Swyx** [50:04]
Uh, I just wanted to give a shout-out to the Pig.dev website. The pigs jump when you click on them.

**Paul Klein IV** [50:08]
Yeah.

**Swyx** [50:08]
Uh, it's great.

**Paul Klein IV** [50:09]
Eric, he's the former, uh, co-founder of Banana.dev too.

**Swyx** [50:11]
Ah.

**Paul Klein IV** [50:12]
This is a different Eric.

**Swyx** [50:13]
Oh, that Eric.

**Paul Klein IV** [50:13]
Yeah.

**Swyx** [50:13]
That Eric. Okay. Well, he abandoned Bananas for Pigs, so I don't know-

**Alessio** [50:17]
Well, I hope he doesn't start going around with pigs now- ... like he was going around with bananas.

**Swyx** [50:21]
Little toy pig. Yeah, I love that.

**Alessio** [50:23]
Yeah. What else are we missing? I think we covered a lot of, like, the Browserbase product history, but...

**Swyx** [50:28]
What do you wish people asked you?

**Alessio** [50:29]
Yeah.

**Paul Klein IV** [50:30]
I wish people would ask me more about, like, what will the future of software look like, 'cause I think that's really where I've spent a lot of time about why do Browserbase. Like, for me, starting a company is, like, a means of last resort.

Like, you shouldn't start a company unless you absolutely have to. And I remain convinced that the future of software is software that you're gonna click a button, and it's gonna do stuff on your behalf. Right now, software, you click a button, and it maybe, like, calls a backend API and, like, computes some numbers.

It, like, modifies some text, whatever. But the future of software is software using software. So I may log into my accounting website for my business, click a button, and it's gonna go load up my Gmail, search my emails, find the thing, upload the receipt, um, and then comment it for me, right?

And it may use that using APIs, maybe a browser. I don't know. I think it's a little bit of both. But that's completely different from how we've built software so far, and that future of software has different infrastructural requirements.

It's gonna require different UIs. It's gonna require different pieces of infrastructure. I think the browser infrastructure is one piece that fits into that, along with all the other categories you mentioned. So I think that it's gonna require developers to think differently about how they've built software for, you know, application level so far.

And, um, I'm excited to kinda explore more what that means. And I think we've seen from, like, you know, the s- customers that use Browserbase so far, some really innovative ways to, like, take software and really reimagine it for AI and build things that, like, have chat interfaces, build things that have human-in-the-loop flows, build things that are more asynchronous 'cause AI is slower.

And those are patterns that are still emerging, and I don't think we have all the best practices yet.

**Swyx** [52:04]
I don't have much to-

**Paul Klein IV** [52:04]
All right. Sweet

**Swyx** [52:05]
... to feedback on that. Like, that's true.

**Paul Klein IV** [52:06]
Paul's right.

**Swyx** [52:07]
Paul's right.

**Paul Klein IV** [52:08]
Quoted by-

**Alessio** [52:08]
You heard it here first

**Paul Klein IV** [52:09]
... by Swix. Yeah. Amazing. I'm framing that.

**Swyx** [52:11]
It is not specific enough to be wrong.

**Paul Klein IV** [52:13]
That means Paul is right to me still. I don't, I don't know if I'm hearing that wrong.

**Swyx** [52:18]
I always try to prompt people for falsifiable predictions because, like, you can predict that things will be better generically, but how? And, like, tho- those are the, the things where you, you, like, put a little skin in the game where-

**Paul Klein IV** [52:29]
Yeah, I mean, I, I can predict that Browserbase will be a billion-dollar company one day. So let's check back in five years and, um-

**Swyx** [52:37]
Yeah. Yeah, yeah

**Paul Klein IV** [52:37]
... you know, if I'm a PM at Coinbase, then something went wrong, so.

**Swyx** [52:40]
Oh, boy. Yeah, yeah. Yeah, we, we picked out a couple of your tweets about founder... Yeah, you, I think you're a pretty building in public kinda guy. Um-

### Solo Founder

**Paul Klein IV** [52:46]
Yeah, I try to be

**Swyx** [52:47]
... the, I think the main thing that I want to highlight as well is, uh, you, you emphasized this at the start of your intro, which is you're a solo founder. I think that there's a movement towards more solo founders in the, in the Valley more generally, but people who are hearing this for the first time have no idea.

They're like, "What do you mean? YC forces me to get a co-founder." Like, what, what is this? So you've... I've, I've heard you talk about this before, but maybe you wanna recap your, your spiel for, for folks that haven't heard about it.

**Paul Klein IV** [53:11]
Yeah. Yeah. I mean, I've had co-founders in, in my past company. I love my co-founders. They're at my wedding. I think if you wanna move extremely fast as a company, one of the hard parts about having co-founders is that there's, like, you have to do the, the co-founder alignment and then the company alignment.

And then there's people on the team that probably tell things to one co-founder 'cause they have a favorite and then, like, that co-founder has to represent their interests. But at Browserbase, I... is a benevolent dictatorship, you know? Like, if I wanna make a change, I work with the team, and we all decide together, and we move quickly.

We don't have an extra layer of buy-in within the co-founder, uh, layer. And frankly, like, I think, especially with dev tools companies, if you're able to talk about your product, uh, and talk with customers, and you can build product, you don't need to have a business guy or a business side.

You know, I'm a developer f- first and foremost. I was raised by two salespeople, so I guess that's why I can talk to customers or something. But at my core-

**Swyx** [54:04]
What kinda sales?

**Paul Klein IV** [54:04]
I love... Uh, they did, uh, semiconductor and pharmaceutical sales, my mom and dad.

**Swyx** [54:07]
Oh, very different.

**Paul Klein IV** [54:08]
Yeah, very different.

**Swyx** [54:09]
But also very enterprise. Good.

**Paul Klein IV** [54:11]
Yeah. Yeah, yeah, yeah. I mean, like, I'm, I'm, I'm... It rubbed off on me in some way. I was just trying to play WoW as a kid, and they made me play sports, so I don't know how it worked out the way it did.

But it does all come back to, like, as a solo founder, you need to be willing to, like, go out there and, you know, talk about your product, go talk to customers, go convince people to work for you, but then also have core principles of, like, how you wanna build this company and, like, what product you wanna build.

And, um, thankfully, if you can do all of that, you can be a solo founder. You just have to hire fast and put the right team around you, and I'm, I'm lucky to have the team that we do that's surrounding me and kinda lifting the whole company up.

**Alessio** [54:45]
So there's kinda, like, the decision-making, and then there's, like, the culture of a company. Obviously, as a solo founder, you have huge influence on everybody. Apple is maybe the usual example of, like, you know, you have the-

**Paul Klein IV** [54:56]
Well, not solo

**Alessio** [54:57]
... Jobs and Wozniak.

**Paul Klein IV** [54:58]
Okay.

**Alessio** [54:58]
No, no. Of, like, you can have two co-founders that are, like, each polarizing in their own-

**Swyx** [55:01]
There was a third co-founder, by the way.

**Alessio** [55:03]
Yeah.

**Paul Klein IV** [55:03]
Who's the third co-founder?

**Swyx** [55:04]
Like, I don't know. He sold his shares very early on. Nobody talks about him, but he, he's, like, he always, um, has a, has a bit of a regret.

**Alessio** [55:11]
But anyway, yeah, have you thought about building the culture? You know, obviously, startups are, like, super intense, but you also cannot just run yourself to the ground all the time. Any insight doing it solo?

**Paul Klein IV** [55:22]
Yeah. I mean, like, I talked about, like, how it's easier for me to make decisions being a solo founder. The real cheat code is, like, having a great team that you give a lot of agency and ownership to.

A lot of people make the little, tiny decisions that go into everything that makes Browserbase great, like the website, for example. I, I had some, like, some involvement with that, but, like, a lot of that was the team, right?

And, and the product. I, I think the team really has ownership of a lot of these day-to-day decisions that add up to make a cohesive product experience. Culturally, like, we're fully in person. Maybe that's one crazy take that we do.

But we're also, like, not ... two in person. Like, our first meeting's at 10:00 AM. People leave around 5:00 or 6:00. We work Monday to Friday in person, and those, like, that's the r- the expectation, right? I think people have gone too far with in person, where they're like seven days a week in the office, 9:00 AM to 9:00 PM.

That's too much.

**Swyx** [56:12]
Just an anecdote.

**Paul Klein IV** [56:13]
Yeah.

**Swyx** [56:13]
I just visited an office, I'll keep them anonymous for now, but to my face, we are 9-9-6.

**Paul Klein IV** [56:18]
Yeah.

**Swyx** [56:18]
For those who don't know, 9-9-6 is 9:00 AM to 9:00 PM, uh, six days a week.

**Paul Klein IV** [56:21]
I think we've taken it a little too far. And for some teams, I know another anonymous company that does something like 9-9-6, and they're, like, crushing it right now, right? So, like... And, like, it does get results, but, like, I think for our culture, we gather in person, we put pants on every day and go to the office so that we can all work together.

Um, or shorts- ... I guess, right? And then, like, we all know we're gonna work outside of o- out of the office. We're gonna work at home sometimes. We might come in on a weekend. The weekends are for fun work, and that's really where, where we get to let people work on stuff that's not on the roadmap.

And that empowers them to build something and bring it back to the team on Monday and say, "Look what I built. This is cool." Culturally, we're a lot of, like, former YC CTOs, um, and, like, ex-founders or future founders, and I've just found that those people tend to be just really great early hires for a company.

They, they get it. And I think for them, especially kind of the ex-YC people who maybe didn't find PMF, coming in and being at a company with PMF, it's such a refreshing thing for them because they can just come in and execute.

And there's just so many clear things we have to go build, and if you're a talented engineer, being able to go build and make an impact every single day is, like, super fulfilling.

**Swyx** [57:26]
My question on, on, on the other hand is you also talk a lot about recruiting, especially in the podcast that you talk about. How come there's no Browser-based recruiting agent?

**Paul Klein IV** [57:34]
That's a good question. I think it's 'cause I don't do that much outbound. I do message people, but a lot of it's now through referral. It's very, like, targeted. Like, if I see somebody working on something really cool, I just message them.

So-

**Swyx** [57:47]
Okay

**Paul Klein IV** [57:47]
... I don't want, like, somebody g- trawling the web and, like, messaging every Kubernetes Firecracker expert. I try and, like, look for them in my passive web browsing. And, um, when I find somebody, I, I just wanna, like, take the time personally to, like, say, "Hey, I love what you're doing.

I think it's really cool, and let's have a conversation."

**Swyx** [58:03]
Yeah, off of Hacker News and other-

**Paul Klein IV** [58:05]
Yeah

**Swyx** [58:05]
... other stuff.

**Paul Klein IV** [58:06]
I'd love to hire off of Hacker News.

**Swyx** [58:07]
Yeah. Let you plug at the end. My attempt at this failed, which is I really hate LinkedIn Sales Navigator. I think that it is, uh, i- it is just grifting on top of people doing data entry for LinkedIn, and I, I, I hope that Browserbase will someday help to kill LinkedIn Sales Navigator.

That, that's my-

**Paul Klein IV** [58:22]
I don't know if we will directly, but one of our customers definitely is trying to do that. So I, um, I think there's a couple that are on it. These AI SDR companies are crushing it, so.

**Swyx** [58:31]
And, yep, the, the 9-9-6 company was an AI SDR company.

**Paul Klein IV** [58:33]
There we go.

**Swyx** [58:34]
Uh, yeah, very classic.

**Alessio** [58:35]
This was great. Anything, yeah, that we missed? Uh, you got the run clubs too. What other things do you mix in? Like, building the company culture-

**Paul Klein IV** [58:42]
Mm-hmm

**Alessio** [58:42]
... and, like, the community culture. I know you bring people together-

**Paul Klein IV** [58:46]
Yeah

**Alessio** [58:46]
... often.

**Paul Klein IV** [58:47]
I, I think, like, we, like, we try and build in public and, like, like, you can see a lot of the Browsebase people on Twitter. Every Monday, we have a run club. People go running together. We don't run very fast, but it's, like, a good way to spend time together.

I just look back fondly on my time being in person at my first company, and we have people, like, with a mix, like, people who've, like, are just early in career, people who've been in, you know, the workforce for 20, 30 years.

So it's not just, like, a young people company. Like, it's a huge mix. But when you make people make a polarizing decision of, like, I will come to an office five days a week, people then end up making more sim- decisions that are aligned with the culture.

So it's almost like if you can make your culture binary, or you're in or out, it becomes easier to assimilate and, like, keep a cohesive culture, and I think that starts with being in office for us. But for other people, it could be, like, moving or, like, using Discord versus Slack or, like, other, like, binary decisions that people may have to make.

### Dream AI

**Swyx** [59:37]
One thing I like as asking founders is, you know, you're famously not an AI company, or, you know, y- you serve AI companies, but you're not y- yourself a LLM sort of consuming company. But if you were, though, what company would you start?

What's, like, what's, like, obviously a good idea?

**Paul Klein IV** [59:51]
Yeah, I, I had this tweet, like, forever ago, which is like, there's so much money to be made in taking, like, proprietary research and then turning that into, like, uh, an automation, which is obviously, like, a very, like, Browserbase inspired one.

Like, listening to all the city halls or town hall meetings in, like, little towns and then knowing when they're gonna, like, approve a new Walmart or something, and then, like, buying up real estate around the Walmart because that will go up when they install this thing.

So it's, like, really interesting to think about, like, how can you find new channels for data that will allow you to make, like, high alpha decisions, um, and, and benefit you financially. So I think there's, like, some interesting stuff there of, like, just a bunch of conversations that happen in real life that are recorded, that are online, that you can go find using, you know, a web browser, uh, of course.

And then, like, making some interesting, like, decisions off of that. So-

**Swyx** [1:00:40]
Yeah

**Paul Klein IV** [1:00:40]
... I don't know, like, I, I like-

**Swyx** [1:00:41]
That's a very-

**Paul Klein IV** [1:00:42]
... the Browser stuff. Like, I, it's on brand, right? Like, I have to... I'm consistent at least. Just-

**Alessio** [1:00:46]
Do not look at it on your phone-

**Paul Klein IV** [1:00:47]
Yeah

**Alessio** [1:00:47]
... through a native app. Only look at it through the browser.

**Swyx** [1:00:50]
Uh, my favorite part of one of his videos, they had these, uh, these guys holding this bee behind them while they were doing the demo, so it was, like, a really Easter egg type thing.

**Paul Klein IV** [1:00:57]
Yeah.

**Swyx** [1:00:57]
Was... That was stagehand, right?

**Paul Klein IV** [1:00:58]
Yeah, the stagehand video, it's not... They're not holding it. They're actually wearing these bee boxes on their heads.

**Swyx** [1:01:03]
Oh.

**Paul Klein IV** [1:01:03]
And we shot it, like, five times, and poor Sean and Samil are, like, bobbing their heads back and forth with these bee boxes on 'cause, hey, we can't afford special effects, man.

**Swyx** [1:01:12]
This is a views website.

**Paul Klein IV** [1:01:12]
This is a really serious thing, you know?

**Swyx** [1:01:14]
Good detail. Good eye for detail there. Uh, yeah, thank you so much. Um, congrats-

**Paul Klein IV** [1:01:17]
Yeah. Thanks, Paul

**Swyx** [1:01:17]
... on all your success.

**Paul Klein IV** [1:01:18]
Thanks for, thanks for having me, guys. It's been a really good time.

**Swyx** [1:01:20]
Yeah. I'm sure we'll have you back again.

**Paul Klein IV** [1:01:22]
Yeah, I'd love to come back.

---

This library is powered by PodHood (https://podhood.com), the podcast website platform.
