Intro0:00
All right, we are here for a very special edition of Lane Space with my buddy, Jared Palmer, uh, SVP at GitHub and VP at CoreAI at Microsoft.
That's correct. Dual title.
Yeah.
Twice the fun.
Is it weird to have two jobs?
No. I'm only on-- I'm only... The full disclaimer, I'm only on day 13, I think.
Yeah.
So early days, so, so far, so good.
So far, so good. Uh, we've been trying to get you on the podcast for two years, I think.
I think so, yeah.
You-
We finally went for it.
Yeah. You're a busy guy.
We like to do it in person, so...
Yeah, we have to do it in person. Yeah, exactly.
Yeah, it's way better.
Uh, I should also plug that you have-- if, if Jared Palmer fans should dig into your previous podcasts with Ken Rivera.
How about that?
It's from, from the interim. Okay. Uh, so, um, but-
But shout out to Ken.
Shout out to Ken. Uh, before that, you were building, um, I, I guess, like, v0 and AI SDK, and you were just sort of VP of AI at-
At Vercel, yeah.
Vercel.
Yes, all, uh, AI initiatives and vibes.
Agent HQ Vision1:00
Yeah. And I feel like basically you went from, like, you sort of building one coding agent to now being the home for all coding agents.
Mm-hmm.
Is that, is that, like, the, the general vibe of Agent HQ?
I, I think that's right, yeah. So backing up, um, I spent the last sort of two years or so building v0 at Vercel and AI SDK. Uh, and then this summer took time off and now joined GitHub, and today we launched Agent HQ, among other things here at Universe.
And, uh, yeah, it's gonna be the home, we hope, of not only agents, but also developers, and it seems like, uh, the gravity well of this new collaboration space that we're trying to build.
Yeah. What do you think, like, basically that GitHub can do that you couldn't do at v0?
GitHub is an enormous platform, right? So these are-
Hundred and eighty-three million, it's some amount of-
Yeah, it's hundred and eighty million developers.
Pretty-
Um, it's just the scale is immense, right? Um, and v0 was focused on not only one language, but one framework.
Right.
And a specific problem space with a built-in renderer and, you know, um, you know, for those who are n- not aware of v0, it's, it's like a bolt or lovable, but it's built by Vercel, and, uh, it's focused on building Next.js apps, specifically Next.js apps.
That constraint was rather liberating for the team at the time, and it lets us really, like, laser focus.
It does. That's your world, edit.
I hope... Thank you. I hope so. Um, and obviously, at GitHub, you know, we're the home of all languages and, you know, and, and frameworks and, and developers. And so the scope is broadened, and yeah, it's just, it's just a different, a different part of the map, if you will, right?
Yeah. How I've seen it-- So you, you've been basically covering the entire journey of coding agents-
Yeah
... for, you know, from, from the start. Like, what do you think-- Uh, what's your jour- personal journey through coding agents, right? Like-
Sure
... is- We started out with Copilot, obviously GitHub-
Sure
... started the Copilot trend.
Sure, sure.
Uh, when... Tell us about the origin story of v0-
Yeah
V0 Origins2:52
... and then how that develops and, and maybe where, where, like, what you wanna see next with-
So-
... history.
It's, it's funny you ask that. As I've told this story multiple times, I feel like I've unlocked different parts of it in my brain-
Yeah
... by going back. You know, like, so maybe we'll have to figure out how retrieval memory works.
Which is, by the way, interesting how memory works for agents.
Totally. That's why I brought it up.
Like, this is wild.
It's, like, kind of amazing as you-
Yeah
... sometimes you discover new paths, right?
Yeah.
Anyway, um, the story goes like this. So when, um, ChatGPT first came out, obviously it was incredible, right? Like, world-changing immediately, fastest growing product ever. I looked back at, like, the timeline and dates, and, um, we were very early, like, in, uh, when I was at Vercel, uh, jumping into AI stuff.
But the journey kinda went like this. So at the time, actually, there was no AI division. There was no AI group. I was actually the director of engineering for all of Vercel Frameworks, and I was helping Next.js svelte, uh, svelte, right.
Yeah, so Next.js svelte, uh, TurboRepo, TurboPack, Webpack, and all internal dev tools at Vercel. And I was helping the Next.js team dog food, um, and test the initial implementation of server actions. And instead of building a to-do app, Guillermo, the CEO of Vercel, was like, "Why don't you build, like, a playground?"
And I was like, "Okay, cool." So that led to the AI playground, which is now just part of AI SDK. We'll get there in a second.
Which, by the way, iconic for, like, side by side, but also the-
Right, so nat.dev
... main scene is nat.dev.
Is nat.dev. So, so G-G told me the nat day of... I got a, I got a DM. It was like-- I remember 'cause I was at a bachelor party, and Guillermo, totally online, sends me a note like, "nat.dev is launching on Monday.
You have to ship." And I had been working on it previously, and so it was like-
So it was the same idea.
It w- Yeah, so he got wind of it, I guess. Uh-
Oh, interesting
... so I, I definitely had to, like, jump into motion. And I didn't think we-- we didn't even ship chat first. Guillermo sends me this, this, uh, DM over the weekend. I'm at a bachelor party, and he's like, "nat.dev," this, the side by side.
He sends me the link, and I play with it, and I'm like, "Cha." Okay, so I spring into gear, ship, um, the AI playground. What was cool about the AI playground was it, it, it forced me to go through every single model provider's API docs and figure out their s- their quirks of their nuanced streaming, 'cause at the time it wasn't like everybody used OpenAI.
Yeah.
It was like all little quirks. Some of them kinda were compatible. So that was my first foray into it. And then launched AI Playground. That shot to the top of Hacker News. And I remember we didn't-- I didn't even implement chat, 'cause that chat wasn't actually, like, a fa- like, wasn't as important.
It was just, like, completion solutions. So eventually we, we factored chat, and that, that project, uh, out of that came AI SDK, because I had already looked at all the model providers and all the, um, the combinations. It was like, "Okay, here's that chunk of streaming code you need."
And then AI SDK found that sort of niche of, like, how do we focus on the part that we're gonna be good at, which is, like, that UI aspect of it, but then also not get in your way.
So that, we shipped AI Playground, then AI SDK launched. And then, um, you know, we're always about, um, demos and having great starter templates at Vercel, and I ha- a- and I remember writing Guillermo. I was like, "You know what'd be cool?
This guy, ShadCN. Oh, man, he seems, like, amazing, and his UI library is doing great. Why don't we team up and ship a ChatGPT clone open source?" And we did. We shipped this, this, this awesome template, which is now called Chat SDK, but it's great.
And- Uh, what that did though at Vercel was it, um, set us up for like rapid experimentation because we had this really good, like pretty full featured ChatGPT ready to rock with all the latest features.
Yeah.
So when it came to like rapid prototyping that summer, now we're summer '23, it was so great. It was, it was like liberating. So I remember at that point I had gained some momentum internally and pivoted almost entirely to AI.
Um, and I had Shu Ding, uh, who you've, you're friends with, um, and Max Lighter and, uh, Shad CN now were cooking. And I think at that point, code execution had just come out. I think that's my timeline here.
Yeah, they're called sandbox, uh, interpreter.
Inter- code interpreter, that's what they called it at the time. And I had a very-- As soon as I saw this, I had a very ambitious idea and proposal to present to Guillermo, which was like, what if there was some...
And mind you, tool calls don't exist. The context window is four thousand-
Wow
... tokens.
Yeah.
So like there's not much here. Um, what if we had this thing where like you could prompt and sometimes it would do code interpretation, and then maybe we could sort of, sometimes we would do code interpretation, but then other times it would choose to render like UI or then it would render sometimes like a document or-
Or some line in the chat.
Yeah. And like it would just have different sort of render modes.
Generative UI.
Yeah. Maybe you could pipe them together. So like the output of one could pipe into... So if we did code interpretation, we coerced it to always emit, uh, like tabular data.
Yeah.
Maybe we could pass that to another prompt that was like a UI and just like some idea there. It's kind of crazy. But if it sounds like these are just tool calls, that's exa- that's exactly what these are.
Exactly what it was. Yeah.
Um, so it became pretty obvious that like the... So I came to Guillermo and we, at the time, we just had a sec-secure-- we had a sort of security debate like should we code interpretation with the ability to fetch data, um, like giving it internet access was like kind of...
Now they're like, "Fine, whatever, whatever. Do whatever you want. Wild, wild west." But at the time, it was like a little scary. So we kind of said, "Okay, no to the code interpretation, but this UI thing, it's pretty neat."
And so that, this like prompt to UI, that was like the aha moment of v0, but the models were not very good, right? So or relative to the way they are now.
And it was, uh, for GPT 4 era?
This is just into the GPT 4 era, and now we're probably at a, uh, sixteen thousand token context window. So you can't really do chat. So we had to kind of invent this kinda new paradigm of like fake it with completion, but that forced us to do sort of the click the...
Well, we-- The, the initial v0, which launched I think in September '23, um, it would look more Midjourney. In fact, if you look like the original suite, it was like Midjourney for React because it was all very visual and you could click on different, uh, components and elements and re-prompt.
But it was again, we were, we were kind of hacking this because we didn't have chat and we didn't have tool calls. Um, and then fast-forward again, um, you know, that launches and then probably like nine months later, um, it took us like nine months to, uh, uh, get to like a million ARR, uh, this little team.
But then the models progressed and from, you know, GPT 4, GPT 4 32K, big, big, big boy. We never really got GPT 4 Turbo working. I don't know why. Just never happened. And then switched to oth-other frontier models and, and then started doing our own models and stuff like that.
Uh, but fast-forward another ten months or nine months or so, and then we, we rebased towards chat, and now the models finally could do chat. Um, and the artifact pattern had evolved, so it was time to rewrite. When we launched v0, the chat version or the new v0, whatever you call it, uh, it, it's like fourteen days, another million MRR, fourteen days, another million MRR.
Right.
It was like a rocket ship after that.
Yeah.
And, um, that just proceeded, uh, and we just kept cooking. And so that's been the journey. Uh, we just kept perfecting. And, and what was really, um, liberating for us was actually the focus on just one stack or one framework.
When everybody else was trying to do general purpose coding agent, we were like, "No, we're just gonna focus on Next.js front-end and ShadCN," and, and that was really-- that allowed the team to focus. So that's the, that's the story arc.
I mean, to be fair, like bolt lovable, like because Next.js is so dominant, basically everyone has to be good at Next.js.
Right.
But being focused on like even right down to the UI library and component stuff like that, that actually helps a lot.
We also started working with all the frontier model labs to help because it was in our Vercel's best interest to have them be great at Next.js.
Yeah.
And also because of the post-train models, and you can read about this on the Vercel blog, like we can, uh, um, m- the post-training, um, like harness that we created, we were started sharing and stuff with other model labs and stuff like that, and we had all our data in a very hygienic state-
Yeah
... to work with them, so.
Did you ever debate internally, a-and because, uh, you know, I've-- I-- my-- from my seat at Cognition, I can also see this, where you should pick the best qualities of every model and string them together in the v0, or you have the model selector and you let customers choose.
We went back and forth, and I think-
Model Strategy11:13
Yeah
... um, you know, we launched a-- We, we went back and forth on all this. I think at the end of the day, there's pros and cons.
Yeah.
One of the benefits of having, uh, your own branded model or synthetic or composite is that you can stitch these things together.
Yeah. There's like-
You can get higher levels of... And, and, and, and now it's a little different with, with this agentic flow, but even look at what you, uh, like what you guys launched recently with-
Exactly
... rep, right? So search is gonna be a different model than what generation, but you know, the genesis, but like because it's, it's a shorter. But like search and genesis are two different mo-- like, like entire subsystems, right?
So you can have search evals that are gonna be totally different. And so how do you... So where we ended up now, uh, where are we? Where, where we've probably ended up now is like, um, for a long time we didn't have model selector, and then we had our own models, which were composites, which we talked about.
And then like, um, and that would allow us to s-- you know, mix, match, and I think that's probably if what It's also nice because you get-- As a-- This is like the product hap-- You get to brand it-
Yes
... right? And you can decouple it from the launch of the Frontier Lab.
Yes.
How do you guys-- How does Cognition even bill for it? Like-
ACUs, uh, compute units
... Right. Because it becomes some synthetic unit, right?
Yeah.
So it gets a little wonky.
Yeah.
We can go on for pricing this, uh, this stuff. It's, it gets challenging. But the nice thing about having the, like, brand name model is that, like, you get to co-launch with the provider, and they'll hype you up.
So, uh, but your billing needs to then is capped at whatever retail is, right? Or, or, or some-
Right. Right, right, right
... or-
You can't really charge a, uh, too much of a premium. You can, but-
Right
... then people are like, "Why?"
But then people are like-
"So I can only pay for-"
... "It's the API key," right? And it's like, well, then how do we charge you for SuiteRep or something, right?
Yeah, yeah.
Going forward, so.
And, and so, like I, I think it's some part of it is the cynical, like you wanna m- create a sustainable business and independence from the, the model labs. But the other part is genuinely, you actually do get better performance-
Right
... if you like string together all these things.
Yes. And so it's, it's tough. I think, uh, uh, what we've-- Switching gears to GitHub, like we're all about model choice now-
GitHub Approach13:18
Right
... and like making sure that-- Uh, and what's cool is that like it's, we also have Copilot, which is our harness, and Copilot CLI, but we also have third-party harnesses like Claude Code and, and Codex and Cognition now, like in a- uh, Agent HQ.
So you kinda get the best of both worlds.
Yeah.
And like I think that's gonna be awesome and ultimately what people want.
Yeah, I think, uh, also the model layer is not the right abstraction to do the switcher anymore. It's-- Which is weird because that's where you started with the ISDK.
Yeah, yeah, yeah, yeah, exactly.
Um, but now it's like the model and the agent have to be strictly tied together, like very, very strongly balanced. You can't loosely bound it and just do a generic interface, because then you're just gonna have the lowest common denominator of all the models.
If you're over in agent world-
Yeah, yeah
... which may just be better than chat world, like in general, like-
Yeah, yeah
... just better.
A-agent world is, is a much better abstraction.
I-I'm calling it agent world, but I mean by like a loop with maybe compute runtime-
Yeah
... and like files. I-
So do you-- That's not your definition of agent? You're dropping your official definition here?
No, don't put me on the spot. But, um, maybe. May-- No. So my, my initial definition of agent for, for AI is-- 'cause like I actually fought, I was dying on this hill, 'cause AI SDK, everyone else is an agent framework, and I think maybe they, they actually went to like this...
I don't know what it says on the front page now, but like who knows? But like-
In like six words
... an agent is, uh, you know, an agent is, uh, orchestrating, uh, ch- you know, an API request with a queue and a for loop. Okay, but a coding agent now has meant so much more. There's like these, you know, coding, coding agent SDKs, and you've got sandboxing and file systems and tool calls and, and I do think that is a uniquely d-- I'll call that agent world, and when charging coding agents here.
And yeah, I think that that seems to be where things are going. And even I believe the, the Claude Excel agent is basically... I was talking, um, to, to Mike Reager backstage, like I think it's like related to Claude Code.
Uh, it could be. I actually don't know-
Yeah, definitely
... how it's implemented under the hood.
It is, yeah.
Yeah.
So it wouldn't surprise me if it was.
Yeah, yeah. They seem very all, all, all in on Skills, which is kind of like an interesting-
What do you think of Skills?
Uh, it's kind of DXT, which is like the, the sort of bundled f- uh, version of MCPs.
Okay. I've heard of that one.
Uh, it wasn't... Yeah, it wasn't... The, the reason you don't know about it is because it wasn't very popular.
Okay.
Uh, so Skills is kinda like the, the, the second shot that is very LLM-pilled. Like, it's like just read my markdown, and just read this, this directory of files and go nuts. Uh, a-as long as like it can understand that you have the capability to run code, to read files, you're good.
Uh, and, and actually, that, that is the universal interface, which is a file system.
Right.
Um-
Back to agent file system.
Yeah , which is, which is kinda cool. Uh, yeah, so I mean, like I, I think like what you're hitting at is, uh, you know, th-this philosophy of like our understanding of what coding agents, the, like, the minimum bar is over the last two years.
Yeah.
Right? Like you've, you've lived this journey, and like now you're basically kind of like the kingmaker or like the -
I don't know about that
... you, you run-
You have frames
... you run, you run Agent HQ and, and like, I mean, I, I, I imagine you have other projects too, but like Agent HQ is like the big one that we're talking about here. Um, like what are you seeing from like the different agents?
Like what do you want to, this to become?
Such a good question. Um, I think that Agent HQ and GitHub itself need to like co-evolve and, you know, one of the things that Microsoft has done really well is by putting things that are, um, alike closer together.
CoreAI Integration16:25
And so you think about the new Core, CoreAI organization-
Yeah
... you've got Visual Studio-
Okay
... Visual Studio Code, GitHub, and parts of Azure all in one. Um, and obviously the GitHub team and the VS Code team have been working closely together for a long time, but now we're really close together. And I think for me, one of the cooler things that, um, Agent HQ can sort of offer is this seamlessness, this fluidity with your workflow, right?
So if you saw in the demo today, we saw a, um, a demonstration of, uh, you use Agent HQ, you fire off a task, and it creates a PR, but you can also open that PR up in VS Code in one click.
And that's, that's awesome. And I think my... the, the vision for, for, for GitHub as it evolves is to look at those touch points where AI can be sprinkled in, you know, uh, salt Bae style, um, into the native workflow, whether you're assigning an issue or maybe some new stuff that I think we should like focus on.
Maybe it could be like, how do we resolve a merge conflict?
Oh my God.
Right? Like, how do we maybe pop open an action or like get in, right? And so it's like-
I think, I think solving merge conflicts is my definition of AGI.
Sh- totally. But like you get that error on an action, and you're like, "We've all been in that, we've all been in that sort of flow where like actions kind of don't work lo-locally versus that tool act." If you're trying to...
"I don't have that set up on my machine. I haven't done this in a wh-"
Right.
So you're like pushing up, and you're, you're this like, "Okay, what if we could just put like, oh, you know, comment or kick off a task to solve this for, for you or..." There are things there. I think what, what I'm trying to describe is this like this workflow where it's just like seamless and fluid, and you can stay in a flow state across whether you're acro- like across all devices, mobile, web, on github.com, or in your local editor.
And I think that's where, you know, my focus is gonna be in the next-
Yeah
... six, six months or so.
Yeah, yeah. Um, just a side tangent on this, uh- Uh, so one of the things that Microsoft also owns, I don't know if it's Microsoft or GitHub, is Dev Containers. And I think, like, a very important concept for sandboxing environments-
Dev Containers18:44
Yeah
... whatever you call it, uh, it is kind of a light version of what Docker containers are, kind of.
Right.
Uh, do you see that as a standard that we should invest in as, like, a, like a thing? Like, 'cause it's supported in VS Code. I don't think it's just that popular outside of VS Code.
Yeah. Uh, it's used internally at GitHub too for, like, de- development at GitHub.
Oh, yeah, yeah.
Which is cool. Um, yeah, I think they were so far ahead almost. Like, it was... But now there's sand- this, like sandboxes, there's so many of these days, right? So I think Cloudflare just launched theirs.
Yeah.
Uh, there's Daytona.
Vercel.
Vercel. Um, Mo- Modal-
Yeah
... uh, which I think Lovable uses.
Uh, I have no idea.
I don't know. Yeah, I mean, you, you probably have your own. I don't know, what do you, what do, what do you guys use?
Uh, just some Kubernetes pods.
Okay, you guys are rolling it yourself. Very cool. I think that's maybe the runtime, but, but there's, there's work and discussion about what that runtime should be even internally at Microsoft, and we've got a couple different competing things, so we'll, we'll figure it out in the next, you know, in the next cycle here.
Right.
But, uh, there's a great point, like, there, there's a lot of cool stuff that i- is in a dev container. You've already got VS Code loaded, you've got a file system-
Yeah
... you've got a sandbox, you've got the security protocol.
Yeah.
It's also, like, wired into GitHub Enterprise-
Yeah
... and, like, ready to be packaged. So there's lots of goodness there.
Yeah, I see, like, the, the number one pain points that Cognition has, but also Codex, also presumably the other guys, is repo setup.
Yeah.
Which is effectively what Dev Containers and a Docker file does for you, is like run this thing, then that thing, set this up, do that thing. Uh, why is it so hard? Like, why haven't we solved it?
I don't know. I think it's hard because you can't predict what's in the, what's in the repo, right? So it's like...
Yeah.
And you don't know when they've bundled FFmpeg.
You don't-
You just don't know.
Right. It, it's nice when, like, if it's just Next.js, you just run PMPM install.
Correct. Correct. You, you can, like, do special ki- like, there's obviously cons- through constraints you can make optimizations, and I think the general purpose container is just, like, challenging. That being said, though, I think there's probably some work to do on, you know, auto-detection and preempting and, and stuff like it to be done there, but it's just a bigger, it's a broader problem space, right?
Yeah.
Um, and, uh, yeah.
So, so fun fact. When I was at Netlify, I actually wanted to reach out to Vercel to do, like, a standardized open source auto-detection thing of frameworks.
Oh, yeah.
And I, I... Like, we, we never, we never really, like, got internal momentum on that. It was an idea.
Okay.
I was like, "Shouldn't this be open source?" You know?
Yeah.
Like, auto-detection is a, is a common utility-
Yes
... that everyone needs.
Yes. Yes. I remember that. I'm having a flashback. It's like half a year then. The-
Yeah. Everyone builds their, their-
Yes, and yeah
... probably we shouldn't all build it.
Right. No, it's, um... And then also, like, what are your defaults? They're not exactly the same, which would be better to, like, just having even the same, like, uh, preference stack-
Yeah
... of defaults-
Yeah
... is the right-
Yeah, yeah
... um, would be great, 'cause then we can move the whole ecosystem together-
Oh, yeah
... from like the PMPM run, right?
Um, okay. So are there other, um, movements or protocols or standards that you're interested in? Like MCP was a big winner this year.
Yeah.
There's other like, I don't know, AC, A2A, ACP, all this, all these, like-
That'd be interesting. I'm not as familiar with...
ACP, the payments one or the-
Is that-
... the red one.
Oh. No, the red one. The, uh-
That one. Okay.
And then the, the payment, that was it Stripe or Coinbase?
Stripe.
Stripe?
Yeah.
Yeah, that's very cool.
We've had them on the pod.
Okay, yeah. That's very cool. Um, but it'd be interesting to see if that takes off. Um-
I mean, it's Stripe.
So yeah, but it still needs to be adopted by, like, the clients, right?
Yeah.
And, and, and I think that's, that's fascinating. Uh, but MCP is huge, it seems. Like, it is the way that a lot of the... Especially when it comes to, like, digital transformation or some of our enterprise customers, it's where they're gonna able to be a- or where they are able to add context.
In addition to that, we also have custom agents that we announced today too, so, like, you can work with prompts and stuff within your, within your Agent HQ and, and customize these agents for different tasks and, and those can have MCPs and such, and I think that's gonna be really powerful from like a, a platform perspective.
Yeah.
It gets me excited. That's what, like, I think is, like, shipping now in, in the next... But we're always on the lookout for, like, the next thing, and, uh, I don't know. What, what's on your... What's, what's top of mind for you?
Uh, for, like, standards or-
Yeah, standards
... standards?
Or should it be dev container?
Uh, dev container is like-
Dev container. A container
... like, is the... Look, I think dev container just is a PR problem.
Right.
It's a great idea.
Right, right.
Just no one makes it interesting.
Yeah, yeah.
I think you can do it, basically, this one.
Okay. Add it to my list.
But before that, probably you have a bunch of other stuff that I, I do wanna get to. Uh, but just staying on the AI stuff, like, uh, I, I think we're, we're, we're actively exploring computer use as, as a thing, like, because it kind of got going a little bit.
People were very excited, and then they found out it was slow and bad and inaccurate.
It is computationally intensive-
Intensive
... from my understanding.
It's getting better.
Yeah.
Uh, especially with open vision models like DeepSeek OCR and, and OMO OCR. Uh, it, it, it... Like, just give it a few more turns of the-
Yeah, it's scaling. It seems very, it seems like it's, like you need that edge case, uh, and primary. It just seems like a modality worth, worth pursuing.
Yeah, I think a lot of people are, uh, on the CodeGen side, the code agent side, a lot of people are trying to think about, all right, we had this evolution from Copilot to, like, uh, you know, a more agentic, uh, sort of cl- like a cloud code situation is where I think that the status is.
Like, what's next, right? What's, what's the, what's the, the, the, the obvious next step?
Agent Quality24:06
Oh, making them good?
Making them good. Yeah. You don't like cloud?
No, no, I, I don't. It's more just like, you know, the, the, the devil's in the details. Like-
Yeah
... going from 90%... Going to, like, hill climbing, it, it gets steeper.
Yeah.
In my opinion. It gets steeper, and so going from 90% success to 95 to 98 to 99 to nines of success-
Yeah
... I mean, really hard.
Paying Mercor a lot of money for, uh, expert programmers of open source maintainers. So, like, you know-
Then you realize along the way maybe the users aren't that good at it.
Right.
Like, um, no, but I, I just think there's a lot of work to do to finish the swing Um, and there's a big difference between ninety-eight percent and ninety-nine percent correct, uh, and that's, like, noticeable. Uh, and this used to hit, you know, if, if you're working on an AI product, you, you, you probably don't realize how-- You've probably seen this, like most people are blind, like living in la-la land about how, um, poor quality, uh, their, their AI product likely is.
Unless they're really measuring, like, the number of error-free sessions, like how many errors are coming from the in-infra, infra providers. Like, you know, how, how many requests are dropped, how fast, how fast these things are. Um, and so that's something that we cared about at Vercel quite a bit, and, like, we'll care about-
What are the pro-- Do you have like a daily review of your dashboard or how, how... I don't know.
Daily? Daily would be slow.
Yeah. Oh, okay. I thought you were gonna say daily is too much.
Daily is s-- No, and even thinking, it's like, uh, every three hours, k- uh, roll up of, of key metrics and stats.
Yeah.
Um, and, like, one of them was like error-free sessions and other things like that, that was, like, really important because, you know, in the-- especially now with agents attribute, like multi-turn. Um, I have a tweet about this that was like in twenty twenty-four, which is that, like, agents will really only work when we get to not only, like, the in- more intelligent models, but better reliability of the infrastructure providers, right?
These aren't, these are not-- Inference is not like a database, like-
Yeah
... update uptime, right? So there's still differences between providers, there's still differences between performance and difference uptimes, and that's why you see things like OpenRouter being very successful-
Mm
... and, and different gateway products because, like, reliability, you need to switch, they go down all the time. Um, so long story short, yeah, we would, we would do like, you know, uh... It was almost like video game style.
Like, we'd have, like, all the data coming in all the time.
Yeah.
Uh, and that allowed-- We-- I used to joke it's my mood ring, like good day, bad day. Um, so it was very successful for us. Uh, I think other teams should adopt that, like, data-driven approach, um-
I think one thing that's surprising is the lack, the relative lack I still see on data analyst agents, where you can sort of chat it, uh, like add a Slackbot for, uh, the, the precise analytics that you wanna generate, um, because I think we're still in the BI era.
Yeah.
Isn't that weird?
It-- Yeah, I totally agree. Um, it's interesting that that space hasn't been, like, captured as much.
Yeah.
Like, I guess maybe now. Actually, I'm interested in this, like, shift to a way into, in, like, into knowledge work tasks with coding agents.
Non-Coding Agents26:57
Okay.
I wonder-
Using coding agents for non-coding tasks?
Correct.
Do you, do you do that personally?
I do, yeah, yeah.
Yeah, what, what do you do?
Well, like, I was-- This summer, I was doing, like I was helping-- I was trying to automate some of my dad's workflows and stuff like that. And just, like, some of his... He's got some Excel spreadsheets and, uh, for like accounting, like, like, uh, financial accounting or managerial accounting, I guess.
Um, and, uh, yeah, just like point cloud code of that stuff and see what happens. And like, I, it's, it, it ends up doing Python, um, and, and, and generating some scripts, and it kinda got off down other hairs, but it was like even he saw that it was better at it than the chat client that makes sense.
Super obvious. They, they-
It became, it became kinda obvious. Yeah, it, it felt better.
I wonder if he can try cloud for Excel and see if-
Yeah, yeah
... it makes sense.
You're probably right. The, um... And then, of course, you got the browser, the... I don't even want-- Not the browser. Browser-based agents, but not, not computer use, but browsers with agents. Agent browsers. Wait, wait, wait, wait.
Oh, right, right.
Agent browsers.
Logter, Perplexity-
Yeah, everybody's coming right at that guy. It's like, is that the better... If that's true, then maybe the, the general purpose injection point is there.
Yeah.
Uh, what do you-- Or have-- Which-- Have you tried any of the agent browsers?
All of them.
All of-- Which one-- What, what's your take?
I, I am very, uh... I'm currently maining Atlas, mostly because I just wanna give ChatGPT a fair go.
Okay.
Um, I... But I'm very stuck to the Arc and, like, the vertical tabs.
Oh, okay.
I think, like, any pro user, like I have multiple businesses.
How many tabs do you have?
I'm context switching, right? I, I have hundreds of tabs open. I made a, I made an open source tool called Chrome Dump, you can find it on my GitHub, uh, where it literally dumps all the tabs open, it summarizes them, and then I can close it by deleting them on Markdown.
That's pretty cool.
And it's, it syncs to .
So you just go on like a-
Oops
... you just go on like a, like a, like a, a bender.
Yeah.
And then you just dump it.
Yeah.
Right.
It should be as easy to close as Markdown, and Chrome is, isn't that good at the performance side of things yet.
And you were working on, like, some browser comparisons just like before.
I, I was. So I tried to build it in Tauri.
Okay. How'd that go?
And Tauri explicitly doesn't want you to build a browser, and I, I tried to fight it too much.
I see. Yeah, I see. Very cool.
So just to wrap things up then, and, and we're, we're around about time, there are other side projects, uh, tasks, and things that you, you've, uh, announced here. Uh, first of all, redesigned GitHub homepage, which a lot of people don't even know GitHub has a homepage.
I, I leg-- I'm legitimately one of them.
Homepage Redesign29:18
I had a tweet, which is, uh, re- Riz's tweet, like printed out, like-
What?
You've seen the tweet. Like, there's a tweet, um, from Riz, and it's like-
Love it
... no one uses-
No one, you know, GitHub homepage
... none-- All of this stuff is totally useless. I'll, I'll pull it up. Like-
Yeah, yeah, yeah
... I, I have-- I pulled it out today 'cause we were like when we launched. Let me, let me get it right 'cause I gotta do it right. Hold on.
Okay.
Was, um, "Incredible how pretty much the entire GitHub homepage is useless," and is one point three million views and nineteen thousand likes. And this was, uh, May two twenty-five. So heard, uh, the team, the team made improvements, and today they launched a new GitHub homepage.
Yeah.
Which I'm very proud of. Um, and they should be really proud of. It's got tasks at the top. It's got recent PRs. Some stuff is still there, like your recent repositories. I still-- I, I think there's more work to do, but it's, like, really overhauled, and they did an amazing job with it.
Yeah.
So they nailed it. Uh, but, you know, more work to do, never done, and, like, hopefully, we can keep iterating with, uh, the community and, and everyone and, uh, and keep going.
The last thing I wanna hit you on is Stacked Diffs.
Oh, yeah.
Stacked Diffs30:21
You asked everyone when you joined, "What should I work on?" Or something.
Yeah, what hap-
I don't know if this is your job specifically.
It wasn't.
But-
It wasn't
... why do people want Stacked Diffs so much? Let's-- I, I think you have some history there.
Yes.
Anyone who's interacted with anyone at Facebook knows-
Right
... about Fabricator.
They're religious about it, yeah.
Uh, so can you explain why, why it's been so-- like, what it is, why is it so hard?
Okay. So this, um- This concept of a pull request, which we're all familiar with, uh, you write some commits, you open a PR, and then you merge the PR, and you go about your day. So as you scale, like, larger organizations, um, and, like, you look at your history, and, and, and there are people who are, like, very r- I'll say, like, have, like, r- near religious beliefs about how to Git, how to do Git right.
Rebase vers- rebase versus merge.
There's, there's a crowd that wants to fast-forward the repository, so to preserve all the history, and then there's a crowd that wants to squash and, and, and merge into the main thing.
I, I'm team squash.
Okay. Anyway, Facebook... And I've, I've never worked at Facebook, but, um, in my previous, uh, startup before Vercel, Turborepo, I did a lot of research on build systems. And at Facebook, they have, not only do they have their custom build tools called Buck, they also have a custom file system, and they don't use Git.
They use, uh, Mercurial, which ... And then now it's sort of custom, and it's all wired together. Um, and at, at, at Facebook, they don't use pull requests. They have a different sort of philosophy. You can, um... It's sort of like the best way to anal- like, think about this is like imagine every PR just had one commit in it.
Yeah.
You could branch them, um, and res- and the critical thing is you can restack them. And then, um, if you restack or make a change later, like earlier in the stack than later in the stack, and these stacks are just diffs, right?
The commits are just diffs, and that's the term, stack diffs. You can then collapse them and merge the last one and merge them towards the end. It just gives you a little bit of a nicer workflow. And it's what people who, if you work on a monorepo or you work on a very, very large code base, it's a really, really nice way to work, especially if you've got a system that will re- that will d- automatically restack.
Um, and then if you think even more deeply about it, like, and get really deeper into the weeds, uh, you can decide which, um, diffs in the stack CI should run against, if you get fan- if you get fancy.
Okay.
May not be-
By, by like certain commit messages and or-
There's always... Yeah, you could decide, like, maybe this one doesn't need it or skip that one or, or whatever, and you end up getting these, like, sort of these groups, these, these stacks. And it's really nice from a code review perspective because, um, when you go to update or you can update a different part of the stack, it just makes it a little bit more fluid, and so it's what people want.
There are a couple tools out there in the market that do this kind of behavior. One's called Graphite. There are a couple others.
Too many tools called Graphite. You see there's another Graphite right here.
Yeah, yeah. Um, and it's just got, it's a great workflow. And so it's been the top pull request, or sorry, the top, um-
Feature request
... feature request, thank you, uh, at GitHub for years. From, from the community's perspective. I don't know at GitHub. But, and then, so as soon as I joined, the first, the first thing I did was go look this up.
And, uh, well, not the first thing. I asked how I should make it up better, and it was the top feature request, right? Uh, and then I went to go, like, okay, investigate, like any good product person would, and it, there's been multiple attempts at this in- internally, um, going back to like 2020.
Okay.
And there was one very, very, very, um, polished attempt too in 2022, and it just... I, I don't have all the context, so, but it was, it was, there was a pretty good implementation. All of the work was done on the client, and it reintroduced this new concept called stacks outside the pull request into GitHub, and it was a little too risky.
It was sort of deemed too risky, too big of a change. That's just what I was told. So anyway, we're, we, we had a couple meetings internally already, and, um, we're trying to weave it into planning and the roadmap, and so hopefully we'll be able to share more updates soon.
But, like, it's a top of the list known feature request.
As, again, heard.
Yeah.
And so
We're, and, like, we're, we're working on it. Uh, obviously, like, something the size of GitHub to move to, like, support this kind of new, this new feature is, like, not just like, um, a walk in the park because of the size of GitHub's and, and GitHub's Git implementation, but, uh, it's something that we're actively exploring.
Yeah. Well, I think, you know, just to wrap all that up, you know, I think it's really nice for someone who's so deeply e- en- engaged and, like, coming from, like, one of us literally-
Yeah, yeah
... that you now run things at, at GitHub, and we can just at you. You know, like
You can just at me.
Yeah, I think, like, uh, Ajay Karpathy was, uh, the other day was saying, like, every company needs one of these where you can just like, "Hey, like, this, this really should exist at GitHub. We love GitHub. We use GitHub, but, like, come on."
Outro35:09
And
Well, yeah. Feature requests welcome. Like, my DMs are always open. It's like-
Oh, careful. I don't know. 180 million developers.
I, I like- 100, whatever. I, I am of the philosophy that, like, all feedback is a gift. Like, it's all a signal.
Yeah.
Um, and the more signal we can collect, the better decisions we can make, and the, truly build this really, really useful website, uh, and company, like, together, and that's gonna be the future. And if we focus just on that, we're gonna be okay.
Yeah. We're gonna be okay. All right.
Yeah.
Well, thanks so much, Jared.
Yeah. Thanks.
This was a real pleasure-
Awesome
... catching up.
Yep. Likewise.
Congrats.





