LALatent SpaceDec 17, 2023· 1:33:26

The "Normsky" architecture for AI coding agents — with Beyang Liu + Steve Yegge of SourceGraph

Beyang Liu and Steve Yegge of Sourcegraph explain how their AI coding assistant Cody achieves a 30% completion acceptance rate by prioritizing high-quality context over agents. They introduce the 'Normsky' architecture—a blend of Norvig's data-driven models and Chomsky's formal systems—arguing that non-agentic, deterministic context retrieval (using parsers, trigram indexes, and the BFG code graph) outperforms multi-hop LLM-based approaches for code completion and chat. The duo details their data pre-processing moat, the 'bin packing' challenge of stuffing relevant code into context windows, and why they use open-source StarCoder for completions while leveraging Claude/GPT-4 for chat. They also discuss the death of DSLs, the limitations of LSP and LSIF, and how Cody's web version lets users query any public repo instantly.

  1. 0:00Intros
  2. 6:20Grok to Cody
  3. 13:18Cody vs Copilot
  4. 16:49Context is King
  5. 20:33Chomsky & Norvig
  6. 30:06Normsky
  7. 36:00Context Techniques
  8. 46:15Protocols & Graphs
  9. 1:02:00Tech Stack & Models
  10. 1:14:35For Managers
  11. 1:26:00Lightning Round

Powered by PodHood

Transcript

Intros0:00

Alessio0:01

Hey everyone, welcome to the Latent Space Podcast. This is Alessio, partner and CTO in residence at Decibel Partners, and I'm joined by my co-host Swyx, founder of Smol AI.

Swyx0:11

Hey, and today we're christening our new, uh, podcast studio in the Newton. And we have, um, Beyang and Steve from Sourcegraph. Welcome.

Steve Yegge0:20

Hey.

Beyang Liu0:20

Hey.

Steve Yegge0:20

Thanks for having us.

Swyx0:22

Uh, so this has been a long time coming. I'm very excited to have you. Uh, we also are just celebrating the one-year anniversary of ChatGPT yesterday.

Steve Yegge0:29

Yep.

Swyx0:29

Um, uh, but also we'll be talking about, uh, the GA of Cody, uh, later on today.

Steve Yegge0:34

Yeah.

Swyx0:34

But, but, um, we'll just do a quick intros of, of both of you. Obviously people can research you and check the show notes for more. Uh, but Beyang, you worked on computer vision at Stanford, and then you worked at Palantir.

Beyang Liu0:45

I did, yeah.

Swyx0:45

Uh, you also interned at Google, which is-

Beyang Liu0:47

I did back in the day, where I got to use Steve's, uh, system dev tool.

Swyx0:52

Right. Uh, what, what was it called?

Beyang Liu0:55

It was called Grok.

Swyx0:55

Yeah.

Beyang Liu0:55

Well, the, the end user thing was, uh, Google Code Search. That's what everyone called it, or just, like, CS.

Swyx1:00

Yeah.

Beyang Liu1:00

Uh, but the, the brains of it were really, uh, the kinda like trigram index and then Grok, which provided the reference graph.

Steve Yegge1:08

Mm-hmm. Today it's called Kythe, the open source Google one.

Swyx1:11

Sure.

Steve Yegge1:11

It's sort of like Grok V3.

Swyx1:13

On your podcast, which, uh, of... You've had me on.

Beyang Liu1:16

Yeah.

Swyx1:16

You've interviewed a bunch of other code search, uh, p- developers, including the current developer of Kythe, right? Or-

Beyang Liu1:22

No, we didn't have any Kythe people on, although we would love to-

Swyx1:26

Yeah

Beyang Liu1:26

... uh, if they're up for it. Um, we had, uh, Kelly Norton, who built a, a similar system at, uh, Etsy. Uh, it's an open source project called Hound. Uh, we also had, uh, Hanwen, uh, Ninhouse, uh, who created Zukt-

Swyx1:40

Yes

Beyang Liu1:41

... which is, uh-

Swyx1:41

That's the one I'm thinking about

Beyang Liu1:42

... uh, I think heavily inspired by the, the trigram index that powered, uh, Google's original code search-

Swyx1:49

Yeah

Beyang Liu1:49

... um, and that we also now use at Sourcegraph.

Swyx1:51

Yeah. So you teamed up with Quinn, uh, over 10 years ago-

Beyang Liu1:55

Yep

Swyx1:55

... to, um, to start Sourcegraph. And, um, I, I kind of view it like, uh, we- we'll talk more about this, like, you, you were indexing all of, all code on the internet.

Beyang Liu2:04

Yeah.

Swyx2:05

Um, and now you're, like, in the perfect spot to create a code intelligence, uh, startup.

Beyang Liu2:11

Yeah, yeah. I, I guess, like, the backstory was, you know, I, I used, uh, Google Code Search while I was an intern, and then after, uh, I left, uh, that internship and, you know, worked elsewhere, it was, like, the, the single dev tool that I missed the most.

I felt like, uh, my job was just a lot more tedious and, and much more of a hassle without it. Uh, and so when Quinn and I started working together at Palantir, he had also used various, like, code search engines in, in open source, uh, over the years.

Uh, and it was just a pain point that we both felt, uh, both working on code at Palantir and also working within Palantir's clients, which were a lot of, uh, you know, Fortune 500 companies, large financial institutions, uh, folks like that.

And if anything, like, the, the pains they felt, uh, i- in dealing with large complex code bases made, you know, our pain points feel, uh, small by comparison.

Swyx3:02

Mm.

Beyang Liu3:02

And so that was really the impetus for, for starting Sourcegraph.

Swyx3:05

Yeah. Excellent. Um, Steve, you famously worked at Amazon.

Steve Yegge3:11

Did. Yep.

Swyx3:12

And revealed, uh... And you, you've told many, many stories. Uh, I wanted, I want every single listener of Latent Space to check out Steve's YouTube, uh, because- ... he, he's, like, he effectively had a podcast that, like, no- like, no, you didn't even tell anyone about or something.

Steve Yegge3:24

Yeah, yeah, yeah.

Swyx3:25

Uh, you just started, you just hit record and-

Steve Yegge3:26

Yeah, yeah

Swyx3:27

... just went on, on a few rants.

Steve Yegge3:28

Mm-hmm.

Swyx3:29

Uh, I'm, I'm always here for a Steve rant. Uh, then you w- then you moved to Google, um, where you, where you also s- had some interesting thoughts on, on just the overall Google culture versus Amazon. You joined Grab as head of eng for a couple years.

I- I'm from Singapore, so I, uh, you know, have, I have actually personally used a lot of Grab's features.

Steve Yegge3:46

Nice.

Swyx3:47

And, uh, it's, it was very interesting to see you, uh, talk so highly of, of, uh, Grab's engineering and, and sort of overall prospects because, um-

Steve Yegge3:53

Because as a customer it sucked?

Swyx3:55

Yeah, no, yeah, it's just like we... Like, no. Uh, well, being from a smaller country, like, you never see anyone from our home country, like, being on, like, a global stage or talked about as, like, a good, um, a, a, a startup that people admire or look up to.

Like, on, on the league that, you know, you with all your legendary experience, um, would consider equivalent.

Steve Yegge4:17

Yeah. Yeah, no, it's absolutely. They actually, they didn't even know that they were as good as they were in a sense. They started hiring a bunch of people from Silicon Valley to come in and sort of, like, fix it, and we came in and we were like, "Eh, they, you know, they...

OE could've been a little better, operational excellence and stuff." But by and large, they're really sharp. And, and, uh, the only thing about Grab is that, uh, they get criticized a lot for being too Westernized.

Swyx4:39

Oh.

Steve Yegge4:40

Yeah.

Swyx4:40

Uh, by who?

Steve Yegge4:41

By Singaporeans who don't wanna work there.

Swyx4:44

Okay. Uh, well, um, I, I guess I'm biased 'cause I'm here, but, uh- ... I certainly don't see that as a problem. Um, and, and, like, if anything, they've, they've made, they have, they've had their success because they were more Westernized than the, the standard Singaporean tech company.

Steve Yegge4:58

I mean, they had their success because they are laser focused. They, they copied Amazon. I mean, they're just, uh, they're executing really, really, really well.

Swyx5:07

Yeah.

Steve Yegge5:07

I mean, for a giant. They had twen- we were... I was on a Slack with 2,500 engineers. It was like, uh, it was like this giant waterfall that you could dip your toe into. You'd never catch up with all the fe-

Swyx5:17

Mm.

Steve Yegge5:17

Actually, the AI summarizers would've been really helpful there.

But yeah, no, I think Grab is successful because they're just out there, like, with their sleeves rolled up just making it happen.

Swyx5:27

Yeah.

Steve Yegge5:28

Yeah.

Swyx5:28

Yeah. And for, for those who don't know, it's, it's not just like Uber of Southeast Asia. It's also a super app, um-

Steve Yegge5:33

Yeah

Swyx5:33

... in the way that-

Steve Yegge5:34

PayPal plus, um, yeah

Swyx5:35

... yeah, in the way that super apps don't exist in the West. Like, it's, it's, it's one of the greatest mysteries-

Steve Yegge5:40

Yeah

Swyx5:40

... enduring mysteries of B2C that super apps work in the East and don't work in the West. Like, we just don't understand it.

Beyang Liu5:45

Yeah, it, it is kinda curious.

Steve Yegge5:46

They, uh, they didn't work in India either, and it was primarily because of bandwidth reasons and, and-

Swyx5:51

Mm

Steve Yegge5:51

... smaller phones.

Swyx5:52

That should change now.

Steve Yegge5:53

Should.

Swyx5:54

Yeah.

Steve Yegge5:54

And maybe we'll see a super app here.

Swyx5:55

Yeah.

Beyang Liu5:56

Yeah.

Swyx5:57

Uh, you worked on... You, you retired, uh, ish.

Steve Yegge5:59

I did, yeah.

Swyx6:00

Uh, you worked on your own video game. Um-

Steve Yegge6:02

Mm, mm-hmm

Swyx6:03

... uh, which, uh, any, any fun stories about that? Any, any, um... A- and I think, and that's also where you discovered some need for code search, right?

Steve Yegge6:09

Uh, yeah. Mm-hmm, mm-hmm, sure. A need for a lot of stuff: better programming languages, better databases- ... better everything. I mean, I started in, like, '95, right? Where there was kind of nothing, so.

Beyang Liu6:20

Yeah.

Swyx6:20

Yeah.

Grok to Cody6:20

Beyang Liu6:20

I just wanna say, I remember when you first went to Grab because you wrote that blog post talking about why you were excited about it, about, like, the expanding Asian market, and our reaction was like, "Oh, man. Why didn't...

How, how did we miss, uh, Steve Yegge?"

Swyx6:32

Oh, hiring, hire you.

Beyang Liu6:33

Yeah, it was like, "Missed that."

Steve Yegge6:34

Wow.

Swyx6:34

How did that-

Beyang Liu6:35

And then-

Swyx6:35

Like, can we tell that story?

Beyang Liu6:36

Mm-hmm.

Swyx6:36

Like, so how did this happen, right? Like, so, uh, you, you were inspired by, by Groq, um-

Beyang Liu6:42

Yeah

Swyx6:42

... yeah.

Beyang Liu6:42

So, uh, I guess, like, you know, the backstory from my point of view is I had used CodeSearch and Groq while at Google, um, but I, I didn't actually know that it was connected to you, Steve. Like, I knew, I knew you from your blog posts, which were always, like, excellent, kinda like inside, very thoughtful takes on, uh...

From, from an engineer's perspective on, on some of the challenges facing, like, tech companies and, you know, tech culture and that sort of thing. Um, but my first introduction to you within the context of, like, code intelligence, code understanding, was I watched, uh, a talk that you gave, I think, at Stanford about Groq when you were first building it, and that was very, uh, eye-opening.

I was like, "Oh, like that guy. Like, the guy who, you know, writes the, the extremely thoughtful, ranty, like, blog posts also built that system."

Swyx7:26

Ooh.

Beyang Liu7:27

Um, and so that's, that's how, that's how I knew, you know, you were kind of in- in- involved in that. And then it was kind of like, uh, you know, we always kind of like wanted to hire you, uh, but never knew quite how to, uh, approach you or, you know, get that, get that conversation started.

Steve Yegge7:42

Mm. Well, uh, we got introduced by, uh, Max, right?

Beyang Liu7:46

Yeah.

Steve Yegge7:46

He was the head of, uh, um-

Beyang Liu7:48

Temporal

Steve Yegge7:48

... Temporal.

Beyang Liu7:48

Yeah.

Steve Yegge7:50

And, uh, yeah, I mean, it was a no-brainer. They called me up, and I had... I, I, I noticed when Sourcegraph had come out. Of course, when they first came out, I had this dagger of jealousy stab through me- ...

piercingly, which I remember because I am not a jealous person- ... by any means, ever. But boy, I was like, "Ar, ar, ar." But I was kind of busy, right? And, uh, and just one thing led to another.

I got sucked back into the ads vortex and whatever, so. Thank God SourceGraph actually kind of rescued me.

Swyx8:19

Here's a chance to build dev tools.

Beyang Liu8:20

Yeah.

Steve Yegge8:21

Yeah. That's the best.

Beyang Liu8:21

Yeah.

Steve Yegge8:21

Dev tools are the best.

Swyx8:23

Um, cool. Well, so, uh, that's the overall intro. I, I guess we, we can get into Cody. Uh, is there anything else that, like, people should know about you before we, we get started? Just-

Steve Yegge8:32

I mean, uh, everybody knows I'm a musician. Uh, so, um-

Swyx8:36

Yeah

Steve Yegge8:36

... um, I can juggle five balls.

Beyang Liu8:40

Five is good. Five is good.

Swyx8:41

Yeah.

Steve Yegge8:41

Right?

Swyx8:41

I've only ever managed three.

Steve Yegge8:43

Five is hard.

Swyx8:44

Yeah.

Steve Yegge8:45

And six a little bit.

Beyang Liu8:47

Wow. That's impressive. Um-

Swyx8:49

Um, so yeah, to jump into SourceGraph, this has been a company 10 years in the making, and a- as Sean said, uh, now you're at the right place-

Beyang Liu8:58

Phase two

Swyx8:59

... the... Now, now exactly. You spent 10 years collecting all this code, indexing, making it easy to surface it, and-

Beyang Liu9:05

Yeah

Swyx9:05

... uh, how-

Beyang Liu9:06

And also l- learning how to work with enterprises and having them trust you with their code bases.

Swyx9:10

Yeah, because initially you were only doing on-prem, right? Like VPC, um, a, a lot of, like, VPC deployments.

Beyang Liu9:16

So in the very early days, we were cloud only, but the first major customers we landed were all on-prem, self-hosted. Uh, and that was, I think, related to the nature of the problem that we're solving, which becomes just, like, a critical, unignorable pain point once you're above, like, 100 devs or so.

Swyx9:32

Yeah. And now Cody is gonna be GA by the time this releases, so congrats. Uh-

Beyang Liu9:38

Thanks.

Swyx9:38

... congrats to your future self for, for launching this in, in two weeks. Um, can you give a quick overview of just what Cody is? I think everybody understands that it's a AI coding agent, but a lot of companies say they have a AI coding agent.

Beyang Liu9:52

Yeah.

Swyx9:52

So, uh, yeah. What does Cody do? How do people interface with it?

Beyang Liu9:56

Yeah, so basically, you know, like, how is it different from the, like, several dozen other AI coding agents that exist in the market now? Um, I think our take... When, when we thought about building, uh, a coding assistant that would do things like code generation and question answering about your code base, I think we came at it from the perspective of, you know, we've spent the past decade building the world's best code understanding engine for human developers, right?

So, like, it's kinda your, your, uh, guide as a human dev if you wanna go and dive into a large, complex, uh, code base. And so our intuition was that a lot of the context that we're providing to human developers would also be useful, uh, context for AI developers to consume.

And so in terms of the feature set, uh, Cody is very similar to a lot of other assistants. It does inline auto-completion. It does code base-aware chat. Uh, it does specific commands that automate, you know, tasks that you might rather n- not, not wanna do, like generating unit tests or, uh, adding detailed-

Swyx10:59

Mm-hmm

Beyang Liu10:59

... documentation. Um, but we think the, the, the core differentiator is, is really the quality of the context, uh, which is hard to kinda describe succinctly. It's, uh, it's a bit like saying, you know, "What's the difference between Google and AltaVista?"

Um, there's not, like, a quick checkbox list of features that you can rattle off, but it really just comes down to all the attention and detail that we've paid to making that context, uh, work well and be high quality and fast for human devs.

We're now kind of plugging into the AI coding assistant as well.

Steve Yegge11:29

Yeah. I mean, just to add, just to, um, add a l- my own perspective onto what Beyang just, just described, um, I'd say, uh, RAG is kind of like a consultant that the LLM has available, right? That knows about your code.

Rag- RAG provides basically a bridge to a, a lookup system for the LLM.

Swyx11:48

Mm-hmm.

Steve Yegge11:48

Right? Whereas fine-tuning would be more like, uh, you know, um, on-the-job training for, for somebody. If the LLM's a person, you know, and you send them to a new job, and you do on-the-job training, that's what fine-tuning is like, right?

Swyx11:59

Mm-hmm.

Steve Yegge11:59

So tuned to a specific task. You're always gonna need that expert, even if you get the on j- on-the-job training, because the expert knows your particular code base, your task, right? And, uh, that expert has to know your code, and there's a chicken and egg problem because, right?

You know, we're like, "Well, I'm gonna ask the LLM about my code, but first I have to explain it," right? It's this chicken and egg problem. That's where RAG comes in, and we have the best, uh, consultant, right?

Alessio12:25

Mm.

Steve Yegge12:25

The best assistant who knows, uh, knows your code. And, uh, uh, and so when you sit down with Cody, right? What Beyang said earlier about going to Google and using code search and then starting to feel like without it his job was super tedious, yeah?

Um, once you start using these... Do you guys use coding assistants?

Alessio12:43

Mm-hmm.

Steve Yegge12:44

Yeah, right? I mean, like, it's get- we're getting to the point very quickly, right? Where, uh, you feel like you're kinda like, almost like you're programming without the internet, right? Or something, you know. It's like you're programming back in the '90s without the coding assistant, yeah?

So, um, hopefully that helps for people who have, like, no idea about coding assistants- ... or what they are.

Alessio13:03

Yeah. And I mean, going back to using them, we had a lot of them on the podcast already. We had Cursor, we have Codium and Codeium, um- ... very similar names.

Steve Yegge13:13

Yeah.

Alessio13:13

Um-

Steve Yegge13:13

Replit, Cr- uh, Phind.

Alessio13:15

Yeah, yeah.

Steve Yegge13:15

And then of course there's Copilot.

Beyang Liu13:16

Tabnine.

Cody vs Copilot13:18

Alessio13:18

Oh, RIP. Um-

Steve Yegge13:19

No, uh, Kythe is the one that died, right?

Alessio13:21

Oh, right.

Steve Yegge13:21

Yeah, yeah.

Alessio13:22

I don't know. It's hard to keep track. Um, so, uh, you had a Copilot versus Cody blog post, and, um, I think it really shows the context improvement. So you had two examples that stuck with me. One was, what does this application do?

And the Copilot answer was like, "Oh, it uses JavaScript and NPM and this," and it's like, but that's not what it does.

Steve Yegge13:42

Yeah.

Alessio13:42

You know? That's what it's built with.

Steve Yegge13:44

Yeah.

Alessio13:44

Versus Cody was like, "Oh, these are, like, the major functions, and, like, these are the functionalities and things like that." Um, and then the other one was, how do I start this up? And Copilot, like you just said, "NPM start," even though there was, like, no start command in the-

Steve Yegge13:59

Right

Alessio13:59

... in the package JSON.

Steve Yegge13:59

Oof.

Alessio13:59

But, you know, mode collapse, right?

Steve Yegge14:01

Yeah.

Alessio14:02

Most projects use NPM start, so-

Steve Yegge14:03

Yeah, yeah

Alessio14:04

... maybe this does too. Um, how do you think about, um, open source models, um, and kinda, like, private... Because Copilot has their own private thing-

Steve Yegge14:13

Yeah

Alessio14:13

... and I, I think you guys use StarCoder, uh, if I remember right.

Beyang Liu14:16

Yeah, that's correct.

Alessio14:17

Um-

Beyang Liu14:17

I think Copilot uses some variant of Codex. They're kinda cagey about it. I don't think they've, like, officially announced what, what-

Steve Yegge14:24

Yeah

Beyang Liu14:24

... model they use.

Steve Yegge14:25

And I think they use a range of models d- based on what you're doing.

Beyang Liu14:27

Uh, yeah. So everyone uses a range of model. No- like, no one uses the same model for, like, inline completion versus, like, chat, because the, the latency requirements for-

Alessio14:34

Fill in the middle. Oh, okay.

Beyang Liu14:35

Well, there's fill in the middle. There's also, like, the, like, what the model's trained on. So, like, we actually had completions powered by Claude Instant for a while and... But you had to kinda, like, prompt hack your, your way to, to get it to output just the code and not like, "Hey, you know, here's the code you asked for."

Like that, that sort of text.

Alessio14:51

Mm-hmm.

Beyang Liu14:52

Um, so, like, everyone uses a range of models. Um, we've kind of designed Cody to be, uh, like, especially model, uh, n- not agnostic, but, like, uh, pluggable. So, uh, one of our kind of design considerations was, like, as the ecosystem evolves, we wanna be able to integrate the best-in-class models, whether they're proprietary, uh, or, or open source, uh, into Cody, because the pace of innovation in the space is just so, so quick.

Um, and I think that's been to our advantage. Like, today Cody uses StarCoder for inline completions, and with the benefit of the context that we provide, uh, we actually show, uh, like, comparable completion acceptance rate metrics. Uh, it's kinda like the standard metric that folks use to evaluate inline completion quality.

It's like, if I show you a completion, what's the chance that you actually accept the completion versus you re- reject it?

Alessio15:43

Mm-hmm.

Beyang Liu15:43

And so we're, we're at par with Copilot, which is at the head of the, the industry right now, and we've been able to do that with the StarCoder model, which is open source, and the benefit of the, the context fetching stuff that we provide.

And, of course, you know, a lot of, like, prompt engineering and, and other stuff, uh, along the way. Um, yeah.

Alessio16:01

A- and Steve, you've wrote a, a post called Cheating Is All You Need, uh-

Steve Yegge16:05

Mm

Alessio16:05

... about what you're building, and one of the points you made is that everybody's fighting on the same axis, which is better UI in the IDE, maybe, like, a better chat response, but data modes are kind of the most important thing, and you guys have, like, a 10-year-old, uh, mode with all the data you've been collecting.

How do you kinda think about what other companies are doing wrong, right? Like, why is nobody, uh, doing this, uh, in terms of, like, really focusing on RAG? I feel like y- you see so many people, "Oh, we just got a new model, and it's, like, a bits human eval," and it's like- ...

"Well, but maybe, like, that's not what we should really be doing," you know? Like, do you think most people underestimate the importance of, like, the actual RAG in code?

Steve Yegge16:47

Yeah, I mean, uh, I think that people weren't doing it much. It wasn't... It's kind of at the edges of AI. It's not in the center. I know that when, uh, ChatGPT launched, so within the last year, I've heard a lot of, you know, rumblings from inside of Google, right?

Context is King16:49

Steve Yegge17:00

Because they're, they're undergoing a huge transformation to try to, you know, of course, get into the, the new world. And, uh, I heard that they told, you know, a bunch of teams to go and train their own models or fine-tune their own models, right?

Both. And, uh, you know, it was a shit show, right? Because nobody ha- nobody knew how to do it, and, and, and they launched, uh, they launched two coding assistants. One was called Codey with a E-Y. Uh, and then there was, uh...

I don't know what happened to that one. And then there's Duet, right? Google loves to compete with themselves, right? They do this all the time. And, uh, they had a paper on Duet, like, from a year ago, and they were doing exactly what Copilot was doing, which was, um, just pulling in the local context, right?

But fundamentally, uh, just, I, I thought of this because we were talking about the splitting of the models. It's... In the, in, in the early days, it was the LLM did everything.

Alessio17:44

Mm-hmm.

Steve Yegge17:44

And then we realized that for, uh, for certain use cases, like completions, that a different, smaller, faster model would, would be better. And, uh, and that, that fragmentation of models, actually, we, we expect it to continue and proliferate, right?

Because we are fundamentally, we're a recommender engine right now. Yeah, we're recommending code to the LLM. We're saying, "May I interest you in this code right here-

Alessio18:05

Mm-hmm

Steve Yegge18:05

... so that you can answer my question?" Yeah? And, uh, and being good at recommender engine, I mean, who, who are the best recommenders, right? There's YouTube and Spotify and, you know-

Alessio18:13

Netflix

Steve Yegge18:13

... and Amazon, whatever, right?

Alessio18:14

Yeah.

Steve Yegge18:14

Yeah, and they all have many, many, many, many, many models, right? For all fine-tuned for very specific, you know... And that's where we're headed in code, too. Absolutely, you know?

Alessio18:24

Yeah. The-- We just did an episode we released on Wednesday, which, uh, we said RAG is like REX's or like LLMs. You're basically just suggesting good, good content.

Beyang Liu18:34

It's like what?

Alessio18:35

Recommendations.

Swyx18:35

Recommendations.

Beyang Liu18:36

Oh, got it. Yeah.

Alessio18:37

Yeah. REXs.

Swyx18:37

Um...

Beyang Liu18:38

Yeah.

Swyx18:38

So, like, uh, the, the naive implementation of RAG is you, you embed everything through in a vector database. You embed your query, and then you, you find the nearest neighbors, and that's your RAG.

Beyang Liu18:46

Yeah.

Swyx18:47

Uh, but actually, you need to rank it, and actually, uh, you need to make sure there's, like, uh, sample diversity and that kind of stuff.

Beyang Liu18:52

Yeah.

Swyx18:52

And then, then you're, like, slowly gradient descenting yourself towards rediscovering, uh, proper REX, REXs, which has been-

Beyang Liu18:59

Yeah

Swyx18:59

... traditional ML for a long time, but, like, approaching it from an LLM perspective.

Beyang Liu19:03

Yeah. It's, it's, uh, I almost think of it as, like, a generalized search problem 'cause it's a lot of the same things. Like, you want your layer one to have high recall and, you know, get all the potential things that could be relevant, and then there's, uh, typically, like, a layer two re-ranking, uh, mechanism that-

Swyx19:18

Yeah

Beyang Liu19:19

... bumps up the precision, tries to get the relevant stuff to the top-

Swyx19:22

Yeah

Beyang Liu19:22

... of, of the, the re- results list.

Swyx19:25

Have you discovered that ranking matters a lot? So, so-

Beyang Liu19:27

Oh, yeah

Swyx19:27

... the context is that, um, I think a lot of research shows that, like, one, context utilization matters, um, b-based on model. Like, GPT uses the top of the context window, and then a-apparently Claude uses the bottom better.

Beyang Liu19:40

Yeah, yeah.

Swyx19:40

But then, and it's lossy in the middle.

Beyang Liu19:42

Yeah.

Swyx19:42

Um, so ranking matters.

Beyang Liu19:43

No, it really does. The skill with which models are able to take advantage of context is always gonna be dependent on how that factors into, uh, the impact on the training loss, right? So, like, if you want long context window models to work well, then you have to have a ton of data where it's like, "Here's, like, a billion lines of text, and I'm gonna ask a question about, like, something that's, like, you know, embedded deeply into it, and, like, give me the right answer."

Uh, and unless you have that training set, then, of course, you're gonna have variability in terms of, like, where it attends to. And in most kinda, like, naturally occurring data, the thing that you're talking about right now, the thing I'm asking you about, is gonna be something that we talked about recently.

Swyx20:20

Yeah.

Steve Yegge20:21

Did you really just say gradient descenting yourself? Actually, I love that it's entered the casual lexicon.

Swyx20:27

Yeah, yeah, yeah. Uh, my, my favorite ver-version of that is, um, you know how you have to p-hack papers? So, um, you know, when you throw humans at the problem, it's-- that's called graduate student descent.

Chomsky & Norvig20:33

Steve Yegge20:39

That's great.

Swyx20:39

Uh, yeah, it's, it's really, it's really awesome.

Alessio20:43

Um, I, I think the other interesting thing that you have is this, um, inline assist, um, UX that is, uh, I, I wouldn't say async, but, like, it works while you can also do work. So you can ask Cody to make changes on a code block, and you can still edit the same file at the same time.

Beyang Liu20:59

Yeah.

Alessio20:59

Um, how, how do you see that in the future? Like, do you see a lot of Codys running together at the same time? Like-

Beyang Liu21:05

Mm

Alessio21:05

... how do you, how do you validate also that they're not messing each other up as they make changes- ... in, in the code? And maybe what are the limitations today, and what do you think-

Beyang Liu21:13

Yeah

Alessio21:13

... about where the tech is going?

Steve Yegge21:15

I wanna start with a little history, and then I'm gonna turn it over to Beyang, all right? So we actually had this feature in the very first launch back in June. Dominic wrote it. It was called Non-Stop Cody.

Alessio21:25

Mm-hmm.

Steve Yegge21:25

And you could, uh, have multiple, uh, basically LLM requests in parallel modifying your source file, and he wrote a bunch of code to handle all of the diffing logic, and you could see the regions of code that the LLM was going to change, right?

And, uh, and he was showing me demos of it, and it just felt like it was just a little before its time, you know? But, uh, a bunch of that stuff, that scaffolding got-- was able to be reused for, uh, for where, where inline's sitting today.

Alessio21:52

Mm-hmm.

Steve Yegge21:52

Where would you, where-- how would you characterize it today?

Beyang Liu21:54

Yeah, so that interface has really evolved from a, like, hey, general purpose, like, you know, request anything inline in the code and have the code update, to really, like, targeted features like, you know, fix the bug that exists, uh, at this line or request a very specific change.

And the reason for that is I think the challenge that we ran into with inline fixes, and we do wanna get to the point where you could just fire and forget and have, you know, half a dozen, dozen of these running in, in parallel.

But, uh, I think we ran into the challenge early on that a lot of people are running into now, uh, when they're trying to construct agents, which is, um, the re- the reliability of, uh, you know, working code generation is just not quite there yet in today's, uh, language models.

Uh, and so, um, that kind of constrains you to an interaction where the human is always, like, in the inner loop, like checking the output of, uh, each response.

Alessio22:51

Mm-hmm.

Beyang Liu22:51

And if you want that to work in, in a, a way where you can be asynchronous, you kinda have to constrain it to a domain where today's language models can generate reliable code well enough. So, you know, generating unit tests, that's, like, a well-constrained problem, or fixing a bug, uh, that shows up in, uh, as, like, a compiler error or a test error, that's, that's a well-constrained problem.

But the more general, like, "Hey, write me this class that does X, Y, and Z using the libraries that I have," um, that is not quite there yet, um, even with the benefit of, of really good context. Like, it, it definitely moves the needle a lot, but it-- we're not quite there yet to the point where you can just fire and forget.

And I actually think that this is something that people don't broadly appreciate yet because I think the... Like, everyone's chasing this, this dream of, uh, agentic execution, and if, if we were to really define that down, I think it implies a couple things.

You have, like, a multi-step process where each step is fully automated, where you don't have, have to have a human in the loop every time, and there's also kinda, like, an LLM call at each stage, or nearly every stage in, in that chain.

Um, and based on all the work that we've done, you know, with the inline interactions, um, with, uh, you know, kind of, like, general co- uh, Cody features for, for implementing longer chains of thought, we're actually a little bit, I think, more bearish than, uh, the average, you know, AI hypefluencer out there, uh, on the feasibility of agents with, with, uh, purely kinda, like, transformer-based models.

To your original question, like, the inline, uh, interactions with Cody, we've actually constrained it to be more, uh, targeted, like, you know, fix the current error or make this quick fix. And I think that, that does differentiate us from a lot of the other tools on the market because a lot of people are going after this, like, snazzy, like, inline edit interaction, whereas- I think where we've moved, and, and this is based on the user feedback that we've gotten, it's like that's, that sort of thing, it demos well, but when you're actually coding day-to-day, you don't wanna have like a long chat conversation in line with the code base.

That's a waste of time. Uh, you'd rather just have it write the right thing and then move on with your life or not have to think about it, and that's what we're, we're trying to work towards.

Steve Yegge24:57

I mean, yeah, we're not going in the agent direction, right? I mean, I'll, I'll believe in agents when somebody shows me one that works. Yeah. Instead, we're working on, um, you know, sort of solidifying our strength, which is bringing the right context in.

Uh, so new context sources, ways for you to plug in your own context, ways for you to control or influence the context, you know, the mixing that happens before the request goes out, et cetera, right? And, uh, there's just so much low-hanging fruit left in that space that, you know-

Beyang Liu25:23

Yeah

Steve Yegge25:23

... agents seems like a little bit of a boondoggle.

Beyang Liu25:25

It-- Just to dive into that a little bit further, like I think, you know, at a very high level, what do, what do people mean when they say agents? They really mean like greater automation, fully automated. Like the dream is like, "Here's an issue, go implement that, and I don't have to think about it as a human."

And I think we, we are working towards that. Like that is the eventual goal. I think it's specifically the approach of like, "Hey, can we have, uh, a transformer-based LLM alone be the kinda like backbone or the orchestrator of these agentic flows?"

Where we're a little bit more, uh, bearish, uh, to-today.

Swyx25:57

You want a human in the loop.

Beyang Liu25:58

Uh, I mean, you kinda have to.

Swyx25:59

Yeah.

Beyang Liu25:59

It's just a reality of, uh, the behavior o-of, of language models that are purely like transformer-based, and I think that's just like a reflection of reality, and I don't think people realize that yet. Because, um, if you look at the way that a lot of other AI, uh, tools have implemented context fetching, for instance.

Um, like you see this in, in the Copilot approach, where if you, if you use like the @workspace thing that supposedly provides like code-based level context, um, it has like an agentic approach where, uh, y-you kinda look at how it's behaving and it, it, it feels like it, they're making multiple requests to the LLM being like, "What would you do in this case?

Would you search for stuff? What sort of files would you gather? Uh, go and read those files." And it's like a multi-hop step, so it takes a long while. Uh, it's also non-deterministic, because any sort of like LLM invocation, it's like a, a dice roll.

Um, and then at the end of the day, the context it fetches is, is not that good. Whereas our approach is just like, "Okay, let's do some code searches that make sense, and then maybe like, you know, crawl through the, the, the reference graph a little bit."

That is fast. That doesn't require any sort of LLM, uh, invocation at all, and we can pull in much better context, you know, uh, very quickly. So it's faster, it's more reliable, it's deterministic, and it yields better context quality.

And so that's what we think. Like we, we just don't think you should cargo cult or, or naively go like, "You know, agents are the future. Let's just try to like implement agents on top of the, the LLMs, uh, that exist today."

I think there are a couple of other technologies or approaches that need to be refined first before we can get into these kinda like multi-stage, fully automated workflows.

Swyx27:40

You know, we're very fo- very much focused on developer inner loop right now.

Beyang Liu27:43

Mm-hmm.

Swyx27:44

But you, you do see things eventually moving towards developer outer, outer loop.

Beyang Liu27:47

Yeah.

Steve Yegge27:47

Yeah.

Swyx27:47

Um, so would you basically say that they're tackling the agents problem that you don't wanna tackle?

Beyang Liu27:54

Um, no. I, I would say at a high level, we are after, uh, maybe like the same high-level problem, which is like, "Hey, uh, I want some code written. I wanna develop some software, uh, and can, can a, an automated system go build that software, uh, for me?"

Um, I think the, the approaches might be different. Um, so I think the analogy in my mind is, think about like the AI chess players, right? Like is, uh, like coding in some sense is, I mean, it's similar and, and dissimilar to chess.

Uh, I think one question I ask is like, "Do you think producing code is, is more difficult than playing chess or less difficult than, than playing chess?"

Swyx28:32

More?

Beyang Liu28:33

Uh, I think more, right?

Swyx28:33

Yeah.

Beyang Liu28:34

And, and if you look at like the, the best AI chess players, like yes, you can use an LLM to play chess. Like people have showed demos where it's like, oh, like yeah, uh, GPT-4 is actually a pretty decent like chess move suggestor, right?

Um, but you would never build like a best-in-class, uh, chess player off of GPT-4, uh, a-alone, right? Like the way that people design, uh, chess players is you have kinda like a search space, uh, and then you have, uh, a way to explore that search space, uh, efficiently.

There's a bunch of search algorithms essentially, where you, where you're doing tree search in various ways, and, uh, you can have heuristic functions which might be powered by an LLM, right? Like you might use a-an LLM to ge-generate proposals in that space that you can efficiently, uh, uh, explore.

Um, but the backbone is still this kinda more formalized, uh, uh, tree search-based, uh, approach rather than the, the LLM itself. And so like our... I think my high-level intuition is that like the way that we get to this like more reliable multi-step workflows that can do things beyond, you know, generate unit test, um, is, is it's really gonna be like a search-based approach, where, where you use an LLM as kinda like an advisor or a pr- a proposal function, um, sort of your heuristic function in like the A- A* search, uh, um, algorithm.

Uh, but it's probably not gonna be the thing that is the backbone, because I guess it's not the right, uh, tool for that.

Swyx30:01

Yeah.

Beyang Liu30:02

Yeah.

Swyx30:02

Yeah. Um, you, you, uh, you also have, um, you... I can see yourself kinda thinking through this but not saying the words, uh, the sort of philosophical Peter Norvig, uh, type discussion. Maybe you wanna sort of introduce, uh, those, those two, that divide-

Normsky30:06

Beyang Liu30:15

Yeah

Swyx30:16

... in software.

Beyang Liu30:16

Yeah, definitely. So, uh, I mean, the, so your, your listeners are, are savvy. They're probably familiar with the classic like Chomsky versus Norvig, uh, debate.

Swyx30:24

Oh, no, actually I wanted... I was-

Beyang Liu30:25

Oh, okay. So-

Swyx30:26

I was prompting you to introduce that, just 'cause

Beyang Liu30:27

Oh, got it. Uh, so I mean, if you look at the history of artificial intelligence, right, uh, you know, it goes way back to, I don't know, uh, it's probably as old as modern computers, like '50, '60s, '70s.

People are debating on like what is the path to producing a sort of like general human level of intelligence.

Swyx30:44

Yeah.

Beyang Liu30:45

Kind of two schools of thought that emerged. Uh, one is the Norvig school of thought, uh, which, you know, roughly speaking includes large language models, uh, you know, regression, SVE- uh, basically any model that you kind of like learn from data and is like data-driven, uh, machine learning.

M-most of machine learning would fall under this umbrella. And, and, and, and that school of thought says like, you know, uh, just learn from the data. That's the approach to reaching intelligence. Um, and then the Chomsky approach is, is more things like compilers and parsers and, uh, formal systems.

So basically like let's, let's think very carefully about how to construct a formal, precise system, uh, and, and, and that will be the approach to how we build a truly intelligent system. Um, Lisp, uh, for, for instance, was like a, a, originally like an attempt to...

I think Lisp was invented to, so that you could create like rules-based systems that you would call AI. As a language, yeah. Yeah. And, and for a long time there was like this debate, like there were certain like AI research labs that were more like, you know, in the Chomsky camp, and others that were more in the Norvig camp.

Uh, and it's, it's a debate that rages on today, and I feel like the consensus right now is that, you know, Norvig definitely has the, the upper hand right now with the advent of, of LMs and diffusion models and all, all the other recent progress, uh, in machine learning.

Um, but the Chomsky, uh, based stuff is still really useful in, in my view. I mean, it's like parsers, compilers. Basically a lot of the stuff that pr- provides really good context. It provides kind of like the knowledge graph backbone- Mm-hmm ...

that you wanna explore, uh, with your AI dev tool. Like, that will come from kind of like Chomsky-based tools, like pa- compilers and parsers. It's, it's a lot of what we've invested in, in the past decade at Sourcegraph, just like...

And, and what you built at- Mm ... uh, with, with Groq. Yeah. Yep ... basically like these formal systems that construct these very precise knowledge graphs, uh, that are great context providers and great kind of guardrails enforcers and, uh, uh, kind of like safety checkers for the output of a more kind of like data-driven, fuzzier system that, that uses like the Norvig, uh, based models.

Steve Yegge32:55

Beyang was talking about this stuff like it happened in the Middle Ages. It feels really old. Like, okay, so when I was in college-

Beyang Liu33:02

Way back when.

Steve Yegge33:02

... okay, I was in college learning Lisp and Prolog and planning and all the deterministic Chomsky approaches to AI.

Beyang Liu33:07

Yeah.

Steve Yegge33:08

Uh, and I was there when, uh, when Norvig basically declared it dead. I was there 3,000 years ago- ... when Norvig and Chomsky fought on the volcano.

Beyang Liu33:17

When, when did he declare it dead? What did, what did, what do you mean he declared it dead?

Steve Yegge33:19

It was like late, late '90s. Uh, yeah, when I went to Google, Peter Norvig was already there. Um, and, uh, he had basically like... I forget exactly where. It was some... He's, he's got so many famous short posts, you know, amazing ones.

Beyang Liu33:32

He had a famous talk, uh, The Unreasonable Effectiveness of Data.

Steve Yegge33:35

Yeah, maybe that was it. But at some point, basically he basically convinced everybody that the determi- deterministic approaches had failed, and that heuristic-based, you know, data-driven statistical approaches, stochastic, were, were better, yeah? The primary reason, I can tell you this 'cause I was there- ...

was that, uh

was that, uh, well, the steam-powered engine... No.

The reason was that it didn't sc- the deterministic stuff didn't scale.

Beyang Liu34:01

Yeah.

Steve Yegge34:02

Right? There reason Prolog, man, your constraint systems and stuff like that. Well, that was a long time ago, right? Today, actually, these, these Chomsky-style systems do scale, and that's in fact exactly what Sourcegraph has built, yeah? And so we have a very unique...

I, I love the framing that Beyang's made, the, the, the, you know, the, the marriage of the Chomsky and the Norvig, you know, sort of models, you know, conceptual models, because we, you know, we have both of them, and they're both really important.

And in fact, there, there's this really interesting like, um, kind of overlap between them, right? Where like the AI or our graph or our search engine could potentially provide the right context for any given query, which is of course why ranking is important.

But it... What it, what we've really signed ourselves up for is an extraordinary amount of testing. Yeah? Because, uh, you know, like, like you were saying, Swets, you were saying that, you know, GPT-4 tends to the front of the context window, and maybe other elements to the back, and maybe, maybe LLaMA more in the middle.

Beyang Liu34:53

It's just an emergent property.

Steve Yegge34:54

Yeah. And so that means that, you know, if we're actually like, you know, verifying whether we, you know, some change we've made has, has improved things, we're gonna have to test putting it at the beginning of the window and at the end of the window, you know, and maybe make the right decision based on the LLM that you've chosen.

Which some of our competitors, that's a problem that they don't have, but we meet you, you know, where you are.

Beyang Liu35:11

Yeah.

Steve Yegge35:12

And we're, and we're, just to finish, we're writing thousands, tens of thousands. We're generating tests, you know, fill-in-the-middle type tests and things, and then using our graph to basically, um, you know, fine, sort of fine-tune Cody's behavior there.

Yeah.

Beyang Liu35:24

Yeah. I, I also wanna add, like I have like an internal pet name that I'm, for this like kinda hybrid architecture that-

Steve Yegge35:31

Sure

Beyang Liu35:31

... I'm trying to make catch on. Uh, maybe I'll just say it here.

Steve Yegge35:35

Do it.

Beyang Liu35:35

'Cause saying it publicly kinda makes it more real. But like I ca- I call it, I call the architecture that we developed, uh, the Normsky, uh-

Steve Yegge35:42

Oh

Beyang Liu35:42

... architecture.

Steve Yegge35:43

Yep.

Beyang Liu35:44

Uh, and it's kinda like, uh, I mean, it's obviously a portmanteau of, of, uh, um-

Steve Yegge35:49

Yeah

Beyang Liu35:49

... Norvig and Chomsky, but the, the acronym, it stands for, uh, non-agentic- ... rapid multi-source code intelligence. So non-agentic because- Oh, rolls right off the tongue.

Context Techniques36:00

Steve Yegge36:01

Wow.

Beyang Liu36:02

And Normsky. Yeah. Uh, uh, yeah. Um, but like it's, it's non-agentic in the sense that like we're not trying to like pitch you on kinda like agent hype, uh, right? Like it's... The things it does are really just use developer tools developers have been using for decades now, like parsers and, and really good search indexes and, and things like that.

Um, rapid because we place an emphasis on speed. We don't wanna sit there waiting for kinda like multiple LLM requests to, to return to complete a simple user request. Multi-source because we really think we're, we're thinking broadly about, you know, what, what pieces of information and knowledge are useful context.

So obviously starting with things that you can search in your code base, and then you add in the reference graph, which kinda like allows you to crawl outward from those initial, uh, results. But then even beyond that, you know, sources of information like, uh, there's a lot of knowledge that's embedded in, uh- docs, in, uh, PRDs or product specs, um, in your production logging system, uh, uh, in, in your chat, you know, in, in your, in your Slack channel, right?

Like, there's so much context is embedded there, and when you're a human developer and you're trying to, like, be productive in your code base, you're gonna go to all these different systems to collect the context that you need to figure out what you, what code you need to write.

And I don't think the AI developer will be any different. It will need to pull context from all these different sources. So we're, we're thinking broadly about how to integrate these into Cody, um, we hope through kind of like a- an open protocol that, like, others can extend-

Steve Yegge37:37

Ooh

Beyang Liu37:37

... and, and implement, and this is something else that should be, uh, I- I guess like, a- accessible by December 14th in, in kind of like a preview stage. Um, but that's really about, like, broadening this notion of the code graph beyond, you know, your Git repository to all the, the other sources where technical knowledge and valuable context can, can live.

Steve Yegge37:56

Yeah, it becomes an artifact graph, right?

Beyang Liu37:58

Yeah.

Steve Yegge37:58

It can link into your logs and your wikis and, you know, any, any data source, right?

Alessio38:03

How, how do you guys think about the importance of... It's almost like data pre-processing in a way, which is bring it all together, tie it together, make it ready. Um, yeah, any thoughts on how to actually make that good What some of the innovation you guys have made?

Steve Yegge38:18

We talk a lot about the context fetching, right? I thought... I mean, there's a lot of ways you could answer this question, but-

Beyang Liu38:24

Yeah

Steve Yegge38:24

... you know, we've spent a lot of time just in this, in this, uh, podcast here talking about context fetching, but stuffing the context into the window is also an... You know, the bin packing problem, right? Because the window's not big enough and you've got more context than you can fit, and you've got a ranker maybe.

But, uh, you know, w- what is, what is that context? Is it a function that was returned by an embedding or a graph call or something? Uh, do you need the whole function, or can you... Do you just need, you know, the top part of the function, this expression here, right?

You know, so that art, the golf game of trying to, you know-

Beyang Liu38:55

Mm-hmm

Steve Yegge38:55

... get each piece of context down into its smallest state-

Beyang Liu38:58

Mm-hmm

Steve Yegge38:58

... possibly even summarized by another model, right, before it even goes to the-

Beyang Liu39:01

Yep

Steve Yegge39:01

... LLM, uh, becomes this is the game that we're in, yeah.

Beyang Liu39:05

Mm-hmm.

Steve Yegge39:05

And so, you know, recursive summarization and all the other techniques that you gotta use to, like, stuff stuff into that context window become, you know, critically important, and, uh, you have to test them across every configuration of models that you could possibly need.

Beyang Liu39:17

I think data, data pre-processing is probably the, like, unsexy, way underappreciated secret to a lot of the cool stuff, uh, that people are shipping today, whether it's... Whether you're doing, like, RAG or fine-tuning or, uh, pre-training. Like, the, the pre-processing step matters so much because, uh, it, uh, uh, it's basically garbage in, garbage out, right?

Like the... If you, if you're feeding in garbage to the model, then it's gonna output garbage. Um, concretely, you know, uh, for, uh, code RAG, um, if you're not doing some sort of, like, pre-processing that takes advantage of a parser and is able to, like, extract the key components of, uh, a particular file of code, you know, separate the function signature from the body, from the docstring, what are you even doing?

Like, that's like table stakes. Uh, you know, and it, it, it allows you... It opens up so much more possibilities, uh, with which you can, um, kinda like tune your system to take advantage of, uh, the signals that come from those different parts of the code.

Like, we've had a tool, you know, since computers were invented that understands the structure of source code-

Alessio40:23

Mm-hmm

Beyang Liu40:23

... to, you know, 100% precision. Like, the, the compiler knows everything there is know- to know about the code in terms of, like, structure. Uh, like, why would you not wanna use that in, in a system that's trying to generate code, answer questions about code?

You shouldn't throw that out the window just 'cause now we have really good, you know, data-driven models, uh, that can do other things.

Steve Yegge40:45

Yeah. When I called it a data moat, you know, in my, in cheating post, um, a lot of people were confused about, uh, you know, because data moat, uh, sort of sounds like data lake because there's data and water and stuff.

I don't know. And so they thought that we were sitting on this giant mountain of data that we had collected, but that's not what our data moat is. It's really a data pre-processing engine that can very quickly and scalably, like, basically dissect your entire code base in, in a very small, fine-grained, you know, semantic units and, uh, and then serve it up, yeah?

And so it's really, it's not a data moat, it's a data pre-processing moat, I guess.

Beyang Liu41:20

Yeah. If anything, we're, like, hypersensitive to customer data privacy requirements, so it's not like we've taken a bunch of private data and, like-

Steve Yegge41:27

Mm-hmm

Beyang Liu41:27

... you know, trained a, a generally available model. I- in fact, exact the, exactly the opposite. A lot of our customers are choosing Cody over Copilot and other competitors because we have an explicit guarantee that we don't do any of that, and that we've done that from day one.

Yeah. I, I think that's a very real concern in, in today's day and age, because, like, if your proprietary IP gets, finds its way into the training set of, of any model, uh, it's very easy both to, like, extract that, that knowledge from the model and also use it to, you know, build systems that kind of work on top of the institutional knowledge that you've, you've built up.

Alessio42:02

About a year ago, I wrote a post on LLMs for developers, and one of the points I had was maybe the death of, like, the DSL. I spent most of my career writing Ruby and- ... um, I love Ruby.

It's so nice to use. Uh, but you know, it's not as performing, but it's really easy to read, right? And then-

Beyang Liu42:17

Yeah

Alessio42:18

... you look at other languages, maybe they're faster, but, like, they're more verbose, you know? And when you think about efficiency of the context window, that, that actually matters.

Beyang Liu42:26

Yeah.

Alessio42:27

Um, but, but I haven't really seen a DSL for models, you know? I haven't seen, like, code being optimized to, like, be easier to put in, in a model context and-

Beyang Liu42:37

Mm

Alessio42:37

... it seems like your pre-processing is kind of doing that. Do you see in the future, like, the way we think about, yeah, DSL and APIs and kind of, like, service interfaces be more focused on being context friendly, where it's like maybe it's less...

It's, it's harder to read for the human, but, like, the human is never gonna write it anyway. Like, we, we were talking on the Hex podcast, there are, like, some data science things, like spin up this pandas, like humans are never gonna write again because the models can just do very easily.

Um, yeah, curious to hear your thoughts.

Steve Yegge43:08

Well, so DSLs are, um, you know, they, they involve, you know, writing a g- a grammar and a, and a parser and, and, uh, it- you know, they're, they're like little languages, right? And, uh, uh, we do them that way because, you know, we need them to, you know, compile, and humans need to be able to read them and so on.

Um, the LLMs don't need that level of structure. You can throw any pile of crap at them, you know, more or less unstructured, and they'll deal with it. So I think that's why a DSL hasn't emerged for sort of like communicating with the LLM or packaging up the context or anything.

Maybe it will at some point, right? We've got, you know, tagging of context and things like that, that are sort of peeking into DSL territory, right? But your point, uh, on, uh, do users, you know, do people have to learn DSLs, like regular expressions or, you know, pick your favorite, right, XPath?

I think you're absolutely right that the LLMs are really, really good at that, and I think you're gonna see a lot less of, uh, people having to slave away learning these things. They just have to know the broad capabilities, and the LLM will take care of the rest.

Beyang Liu44:07

Yeah, I, I'd agree with that. I, I think we will see kind of like a revisiting of like... B- basically, like, the value prop of a DSL is that it makes it easier to work with a, a lower-level language, but at the expense of introducing an abstraction layer.

Steve Yegge44:21

Mm.

Beyang Liu44:22

Uh, and in, in many cases today, you know, without the benefit of AI code generation, like that, that's d- that's like totally worth it, right? Um, with the benefit of, of AI code generation, I mean, it's... I don't think all DSLs will go away.

I think there's still, you know, places where that trade-off is, is gonna be worthwhile. But it, it's kind of like, you know, how much, how much of source code do you think is gonna be generated through natural language prompting in the future?

'Cause in a way, like any programming language, it's just a DSL on top of assembly.

Steve Yegge44:52

Yeah.

Beyang Liu44:52

Uh, right? And so if people can do that, then yeah, like, uh, maybe for a large portion of the code that's written, people don't actually have to understand the DSL that is Ruby or Python or basically any other programming language that exists today.

Steve Yegge45:07

I mean, seriously, do you guys ever write SQL queries now without using a model-

Beyang Liu45:12

Right

Steve Yegge45:12

... of some sort?

Beyang Liu45:12

At least a draft.

Steve Yegge45:13

Ever?

Beyang Liu45:14

Yeah.

Steve Yegge45:15

Yeah, right? And so I mean, we're-- we have kind of like, you know, passed that bridge, right?

Beyang Liu45:18

Yeah. Yeah, I, I think, like, to me, the, the long-term thing is like, is there ever gonna be you don't actually see the code, you know? It's like, hey... The basic thing is like, "Hey, I need a function to sum two numbers," and that's it.

I, I don't need you to generate the code, you know? Like-

Steve Yegge45:33

And the follow-up question: Do you need the engineer or the paycheck?

Beyang Liu45:38

I mean, right? That's kind of the agents discussion in a way, where like- Yeah ... you, you cannot automate the agents, but like s- slowly you're getting more of the atomic units of the work- Yeah, yeah, yeah ...

kind of like, uh, done. I kind of think of it as like, you know, do you need a punch card operator to answer that for you? And so, like, I think we're, we're still gonna have people in the role of a software engineer, but the th- the, the portion of time they spend on these kind of like low-level, tedious tasks, uh, versus the, the higher level, more creative tasks is, is gonna shift.

Steve Yegge46:07

No, I haven't used punch cards. He looks over at me like, "Have you?"

Beyang Liu46:12

Yeah. Uh, no, I've been, I've been talking about, like... So we- we've kind of made this podcast about the sort of rise of the AI engineer. Mm-hmm. Um, and like the first step is the AI-enhanced engineer that, uh, that is that software developer that is not- Mm ...

Protocols & Graphs46:15

Beyang Liu46:25

no longer doing these routine, boilerplatey-type tasks, 'cause they, they're just enhanced by tools like yours. And so you, you mentioned, uh, your OpenCodeGraph. I mean, that, that is a kind of, uh, DSL maybe. Y-y-y... and, um, because we're releasing this, uh, as you, as you go GA, um, you hope to, uh, for other people to, to take advantage of that?

Oh, yeah. I, I would say... So OpenCodeGraph is not a DSL. It's more of a protocol. It's basically like, "Hey, if you want to make, uh, your system, whether it's, you know, chat or logging or whatever, accessible to, um, an AI developer tool like Cody, um, here is kind of like the, the, the schema, uh, by which you can provide that context and offer hints."

Yeah. Um, so I would... You know, comparisons, like LSP obviously did this for, uh, kind of like standard code intelligence. Mm-hmm. It's kind of like a lingua franca for providing finder references and go definition. There's kind of like analogs to that.

There might be also analogs to, uh, kind of the, the original OpenAI, kind of like plugins, uh, API, where it's like, "Hey, you know, here's, here..." There's all this, like, context out there that might be useful for, uh, an LM-based system to consume.

Uh, and so at a high level, what we're trying to do is, uh, define a, a common language, uh, for context providers to provide context to other tools in the software development life cycle. Yeah. Do you have any critiques of LSP, by the way, since, like, this is very much very close to home?

Steve Yegge47:48

One of the authors wrote a really good critique recently, yeah.

Beyang Liu47:51

Oh.

Steve Yegge47:51

Uh, how could have been better.

Beyang Liu47:52

I don't think I saw that.

Steve Yegge47:53

Yeah, yeah. How LSP, LSP could've been better. It just came out a couple weeks ago. It was a good, good article.

Beyang Liu47:57

We'll have to look for that. Yeah. I, I don't, I don't know if I... Like, I think LSP is great. Like, it, for what it did for the developer ecosystem, it was... It's absolutely fantastic. Like, n- nowadays, like, it's, it's very easy, eas- it's much easier now to get c- uh, code navigation up and running in, uh- A bunch of editors ...

in a bunch of editors- Yeah ... uh, by speaking this protocol. I think maybe the interesting question is, like, looking at the different design decisions made, comparing LSP basically with, with Kythe, uh, because Kythe has more of a, um...

I don't know. How would you describe it? I don't wanna-

Steve Yegge48:29

Storage format.

Beyang Liu48:31

I think the critique of LSP from a, a, a Kythe point of view would be like, with LSP, you don't actually have l- an actual model, uh, symbolic model of, of the code.

Steve Yegge48:40

That's right.

Beyang Liu48:40

It's not like LSP models like, "Hey, this function calls this other function." LSP is all, like, range based. Like, "Hey, your token is at-"

Steve Yegge48:47

Yeah

Beyang Liu48:47

... "like, line 32... Your, your cursor's at line 32, column one."

Steve Yegge48:51

Yeah.

Beyang Liu48:51

And that's the thing you feed into-

Steve Yegge48:54

Yeah

Beyang Liu48:54

... the, the language server, and then it's like, "Okay, here's where... Here's the range that you should jump to if you click on that range."

Steve Yegge48:59

Yeah.

Beyang Liu48:59

So it kind of is ex- intentionally ignorant of the fact that there's a, a thing called a reference underneath your cursor, and that's linked to a symbol definition.

Steve Yegge49:07

Well, actually, that's, that's the worst example you could've used. You're, you're, you're right. But, but there... That's the one thing that it actually did bake in, is, is following references, but-

Beyang Liu49:15

Sure

Steve Yegge49:15

... but it, it's sort of hardwired.

Beyang Liu49:17

Yeah.

Steve Yegge49:17

Yeah, yeah.

Beyang Liu49:18

Where- whereas Kythe attempts to model, like, all these things explicitly, and so, uh-

Steve Yegge49:23

Well, so Kythe also... So LSP's a protocol, right? And, uh, so Google's internal protocol is gRPC based, and, uh, it, uh, it is, uh, it's a different approach than LSP. It's, uh, um, basically you make a, a heavy query to the back end and you get a lot of data back, and then you render the whole page, you know?

Um, so we've looked at LSP and we think that, uh, it's just, uh, it's, you know, it's a little long in the tooth, right? I mean, it's a great protocol, you know, lots and lots of support for it, but we need to push into the, the domain of exposing the intelligence-

Beyang Liu49:54

Yeah

Steve Yegge49:55

... through the protocol.

Beyang Liu49:56

Yeah. And so I would say, um, I mean, we ha- we have a, we've developed a protocol of our own called Skip, which is, uh, I think at a very high level trying to take some of the good ideas from LSP and from, from Kythe and, and merge that into a system that in the near term is useful for SourceGraph, but I think in the long term we hope will be useful for the ecosystem.

And I would say, like, the... Okay, so here's what LSP did well. LSP, by virtue of being, like, intentionally dumb, dumb in air quotes, 'cause I, I'm not, like, ragging on it.

Steve Yegge50:24

Yeah, that's good. Yeah.

Beyang Liu50:24

Um, but what it allowed it to do is it, it, uh, allowed language servers developers to kind of, like, bypass the hard problem of, like, model- modeling language semantics precisely. So, like, if all you want to do is jump to definition, you don't have to come up with, like, a universally unique naming scheme for each symbol, which is actually quite challenging, 'cause you th- have to think about like, okay, what's the top scope of, of this name?

Is it the get- the source code repository? Is it the, the package? Uh, you know, does it depend on, like, what package, uh, um, server you're fetching this from? Like, whether it's the public one or the one inside your g- anyways, like, naming is hard, right?

Um, and by just going from a kind of like a location, a location, uh, based approach, you basically just, like, throw that out of the window. All I care about is jumping to definition. Just make that work, and you can make that work without having to deal with, like, all the, the, the co- complex gl- global naming things.

The limitation of that approach is that it's harder to build on top of that, uh, to build, like, a true knowledge graph. Like, if you, if you actually want a system that says like, "Okay, here's the web of functions and here's how they reference each other, and I want to incorporate that, like, semantic model of how the code operates or, or how the code relates to each, each other at, like, a static level," you can't do that with LSP 'cause you have to deal with l- line ranges.

And, like, concretely, the pain point that we found in using LSP for SourceGraph is, like, in order to do, like, uh, a find references and then jump to definition, it's like, it's like a multi-hop process, 'cause, like, you had to jump to the range and then you have to find the symbol at that range.

And it just adds a lot of latency and complexity to these operations, where as a human you're like, "Well, this thing clearly references this other thing. Why can't you just jump me, uh, to that?"

Steve Yegge52:03

Yeah.

Beyang Liu52:03

And I think that's the thing that Kythe does well, but then I think the issue that Kythe has had with, with adoption is because it is more, uh, it, it's a, it's a more sophis- sophisticated schema, I think.

And so there's b- basically more things that you have to implement to get, like, a, a Kythe implementation up and running. I, I-

Steve Yegge52:20

Yeah

Beyang Liu52:20

... want to note, like-

Steve Yegge52:21

No, no

Beyang Liu52:21

... correct me if I'm, uh, I'm wrong about any of this.

Steve Yegge52:23

No. 100%. 100%. Uh, Kythe, Kythe also has a problem, all these systems have the problem, even Skip, uh, or at least the way that we implemented the indexer, is that they have to integrate with your build system-

Beyang Liu52:33

Mm-hmm

Steve Yegge52:33

... in order to build that knowledge graph, right? Because you have to basically compile the code in a special mode to generate artifacts instead of binaries. And I would say... By the way, earlier I was saying that, uh, uh, X refs were in LSP, but it's actually, I was thinking of LSP plus LSIF.

Beyang Liu52:50

Yeah.

Steve Yegge52:50

Ugh, LSIF.

Beyang Liu52:52

That's another-

Steve Yegge52:53

Which, uh, which, which is actually ba- we can say that's bad, right?

Beyang Liu52:56

I don't know. Never, never heard of it.

Steve Yegge52:57

LSIF is not good. Um, it's like, it's like Skip or Kythe. It's, it's, it's-

Beyang Liu53:01

Yeah

Steve Yegge53:01

... supposed to be sort of a model f- uh, you know, a serialization, you know, for the code graph, but it's, uh-

Beyang Liu53:05

Well, uh, yeah, yeah

Steve Yegge53:06

... it basically just does what LSP needs, the bare minimum.

Beyang Liu53:08

L- LSIF, LSIF is basically if you took LSP and turned that into a serialization format.

Steve Yegge53:11

Yeah. Yeah.

Beyang Liu53:12

So, like, you build an index for language servers to kind of, like, quickly bootstrap from cold start.

Steve Yegge53:16

But it's a graph model with all of the inconvenience of the API-

Beyang Liu53:19

Mm-hmm

Steve Yegge53:20

... without an actual graph.

Beyang Liu53:22

Yeah.

Steve Yegge53:22

And so, uh-

Beyang Liu53:23

Yeah

Steve Yegge53:23

... yeah, it's, it's not great.

Beyang Liu53:25

So, like, one of the things that we try to do with Skip is try to capture the best of both worlds. So, like, make it easy to write an indexer and make the schema simple, um, but also model some of the more symbolic characteristics of the code that would allow us to essentially construct this knowledge graph that we can then make useful for both the human developer through SourceGraph and through the AI developer through Cody.

Steve Yegge53:45

So anyway, we, uh, uh, just to f- uh, finish off the, uh, graph, uh, comment, is we've, we've, uh, we've got a new graph, yeah, that's Skip based. Uh, yeah, we call it BFG internally, right? For, um, Beautiful something Graph.

Beyang Liu54:00

Big Friendly Graph. Big Friendly Graph.

Steve Yegge54:02

But blaz- it's blazin' fast.

Beyang Liu54:03

Blazing fast.

Steve Yegge54:03

Blazing fast.

Beyang Liu54:03

Blazing fast graph.

Steve Yegge54:05

And it is blazin' fast, actually. It's really, really interesting. I, I, uh, I should probably have to do a blog post about it to walk, walk you through exactly how they're doing it.

Beyang Liu54:12

Oh, please, yes.

Steve Yegge54:13

But it's a very AI-like, uh, iterative, you know, experimentation sort of approach, where, uh, we're, we're, we're building a code graph based on, you know, all of our 10 years of knowledge about building code graphs, yeah? But we're building it quickly with zero configuration, and it doesn't have to integrate with your build system, and, uh, through some magic tricks that we have.

And so it, it just, well, just happens when, when you install the, the plugin that, that it'll be there and indexing your code and providing that knowledge graph in the background without all that build system integration. This is a bit of secret sauce that we haven't really, like, um, I don't know, we haven't, you know, advertised it very much lately, but I am super excited about it.

Because what they do is they say, "All right, you know, let's tackle function parameters today. Uh, Cody's not doing a very good job of completing function call arguments or function parameters in the definition," right? Yeah, we generate those thousands of tests, and then we can actually reuse those tests for the AI context as well.

So, uh, fortunately, things are kind of converging on. We have, you know, half a dozen really, really good context sources, um, and, uh, and we mix them all together. So anyway, BFG, you're gonna hear more about it, um, probably, mm- I would say probably in the holidays

Beyang Liu55:24

Yeah. I, I think it'll be... I, I think it'll be online, uh, for December 14th. We'll probably mention it. I d- BFG is probably not the public name we're gonna go with. Um, I think we might call it like con- uh, Graph Context or s- or something like that.

Steve Yegge55:37

We're officially calling it BFG.

Swyx55:39

You heard it, you heard it here first.

Beyang Liu55:40

Uh, BFG was just kinda like the working name. And, and it's interesting, like, so the impetus for, for BFG was, like, i- if you look at, like, current AI inline code completion tools, um, and the errors that they make, a lot of the errors that they make, even in kinda like the easy, like, single line case, are essentially, like, type errors, right?

Like, you're trying to complete a, a function call, uh, and it suggests a variable that you defined earlier, but that variable's the wrong type. And that's the sort of thing where it's like, well, like a, a f- a first-year, like, freshman CS student would not make that error, right?

So, like, why does the AI, uh, make that error? And the reason is, I mean, the AI is just suggesting things that are plausible without the context of the types or, uh, you know, any other like, you know, broader files in the code.

Um, and so the kinda intuition here is, like, why don't we just do th- the, the basic thing that, like, a, a, any baseline intelligent human developer would do, which is, like, click Jump to Definition, click some Find References, and pull in that, like, graph context, uh, in- into, into the context window, uh, and then have it, uh, generate the completion.

So, like, that's sort of like the MVP of what BFG was, and turns out that works really well. Like, you can eliminate a, a lot of, uh, type errors, um, that, that AI coding tools, uh, make just by pulling in that context.

Steve Yegge57:07

Yeah, but the graph is definitely our Chomsky side.

Beyang Liu57:10

Yeah, exactly. So, like, this, like, Chomsky Norvig thing, I think pops up in a, a bunch of different layers, and I think it's just a, a very useful and also kinda, like, nicely nerd- nerdy way to describe-

Swyx57:20

Yeah, absolutely

Beyang Liu57:21

... the, the system that we're trying to build.

Steve Yegge57:23

By the way, I remember the, uh, the, uh, point I was trying to make earlier to your question, Alessio, about is AI gonna replace programmers, and I was talking about how compilers, w- they thought, "Oh, are compilers gonna replace programming?"

And what it did was it just changed kinda what programmers have to focus on.

Swyx57:37

Mm-hmm.

Steve Yegge57:37

And I think AI is just gonna level us up again, right? So we're, programmers are still gonna be building stuff and, you know, un- until agents come along, but I don't believe. And so, and so yeah, that's where we're at.

Beyang Liu57:48

Yeah. I mean, to be clear, again, like, with, with the agent stuff at a high level, I think we will get there. I think that's still the, the kinda long-term target. And I think also with Cody, it's like you can have Cody, like, draft up an execution plan.

Uh, it's just not gonna be the sort of thing where, uh, you, you can't attend to what it's doing.

Swyx58:06

Mm-hmm.

Beyang Liu58:06

Like, we, we think that, like, with Cody, it's like you ask Cody, like, "Hey, I have this bug. Help me solve it," it will, it would do a reasonable job of fetching context and saying, like, "Here are the files you should modify," and if you prompt it further, it can actually suggest, like, code changes to make to those files.

And that's, that's a very nice way to, like, resolve issues, 'cause you're kinda, like, on the rails for most of the time, but then, you know, now and then you have to intervene as a human. I just think that, like, if we're trying to get to complete automation where it's, like, the, the sort of thing where, like, a non-software engineer, like someone who has no technical expertise can just, like, speak a non-trivial feature into existence- ...

um, you know, that is still, uh, I think several key innovations away from happening right now, and I don't think the pure, like, transformer-based LLM orchestrator modeled agents that, that, that is kinda, like, dominant today is gonna get us there.

Swyx58:58

Yeah. Yeah, just to, uh, what, what you're talking about triggered a thread I've been working on for a little bit, which is, you know, we- we're, we're very much reacting to developments in models on a month-to-month basis. Um, you had a, you had a post about, uh, um, yeah, we're gonna need a b- bigger moat, which is great Jaws reference- ...

for tho- for those that didn't, didn't catch it, uh, about how quickly-

Steve Yegge59:21

I forgot all about that

Swyx59:22

Quickly, quickly, um, how quickly models are evolving. But I think if you, like, kinda look out, um, I actually caught, uh, Sam Altman on a podcast yesterday talking about GPT-10. Uh,

Beyang Liu59:32

Ooh.

Swyx59:33

I know.

Beyang Liu59:33

Wow. Things are accelerating.

Swyx59:36

Um, and actually, uh, there's a pretty good cadence from GPT-2, 3, and 4, uh, that you can... if you project out. Um, so for, uh, 4 is, uh, based on George Hotz's, um, uh, concept of, like, a 20 petaFLOPS being, like, a human, a human year, a human's worth of compute.

Um, GPT-4 took about 100 years, uh, uh, in terms of human years to, to train in, in terms of the, the amount of compute. So that's one, that's one living person, and every generation of GPT, um, increases two orders of magnitude.

Beyang Liu1:00:06

Mm-hmm.

Swyx1:00:07

So 5 is, uh, you know, 100, 100 people, uh, and if you just project it out, uh, 9 is, um, every human on Earth.

Beyang Liu1:00:15

Mm-hmm.

Swyx1:00:15

And, uh, 10 is every human ever.

Beyang Liu1:00:18

Mm-hmm.

Swyx1:00:19

Uh, s- and he, and he thinks, he thinks he'll, he'll reach there by the end of the decade. So I-

Beyang Liu1:00:23

George Hotz does?

Swyx1:00:23

Uh, no, Sam Altman.

Beyang Liu1:00:24

Oh, Sa- Sam Altman. Okay.

Swyx1:00:25

Yeah. So I, I just like setting those, like, high-level- ... like, you know, you have dots in a line. It... Like, we're, we're like-

Beyang Liu1:00:32

Yeah

Swyx1:00:32

... we're, we're at the start of the curve, it, like, with, uh, with Moore's law.

Beyang Liu1:00:36

Yeah.

Swyx1:00:36

Like, Moore's Law, George Moore I think thought it would last, like, 10 years.

Beyang Liu1:00:40

Yeah.

Swyx1:00:40

And he, he just like kept going for, like, another 50.

Beyang Liu1:00:43

Yeah, yeah.

Swyx1:00:44

And I, I think we're, we're... we have all these data points, and we're just, like, trying to draw, extrapolate the curve-

Beyang Liu1:00:48

Yes, yes

Swyx1:00:48

... out to where this, where this goes. Um, so all I'm saying is, like, you know, this, this agent stuff that we doubt might come here by, like, 2030, and, like, I, I, I don't know how you plan when, um, things are not possible today and you're like, "Ah, it's not..."

Like, it's like, "It's not worth doing." But, like, you know, I mean, we're gonna be here in 2030 and like

Beyang Liu1:01:07

Yeah, yeah.

Swyx1:01:10

And what do we do then?

Beyang Liu1:01:12

So is the question, like, you know-

Swyx1:01:14

There's no que- It's a que-

Beyang Liu1:01:15

Oh, okay.

Swyx1:01:15

It's a, it's a, like, sharing of a comment, uh, just because, like, um, at the back of my head, anytime I, anytime we, we hear things like, uh, things are not practical today-

Beyang Liu1:01:23

Yeah

Swyx1:01:23

... I'm just like, "All right, but how do we-"

Beyang Liu1:01:26

I... So here, here's, here's, like, a, a question maybe. Like, I, I get the qu- whole, like, scaling argument, and I, I do think that there will be something like a Moore's law for, um, AI inference. I mean, d- definitely I think at, like, the, the hardware level, like GPUs.

Um, I, I think it get a little s- it gets a little fuzzier the higher you move up in the stack.

Swyx1:01:43

Sure.

Beyang Liu1:01:43

Um, but for instance, like, going back to the chess analogy, right? Uh, at, at what point do we think that, you know, GPTX or whatever, you know, a pure, a transformer-based LLM model, uh, will be, like, state-of-the-art or outperform the best, like, chess-playing algorithm today?

Tech Stack & Models1:02:00

Beyang Liu1:02:05

'Cause I think that is one milestone on-

Swyx1:02:07

Where you, where you completely overlap, uh, symbolic... Uh, or search-

Beyang Liu1:02:10

Yeah, exactly

Swyx1:02:11

... and symbolic models

Beyang Liu1:02:11

'Cause I think that would be... Uh, I mean, just to put my cards on the table, I think that would kind of disprove the thesis that I just stated, which is, you know, kind of like the pure transformer, just scale the transformer-based approach.

Uh, that would be a proof point where, like, "Hey," like, maybe that is the right approach versus-

Swyx1:02:26

Yeah

Beyang Liu1:02:26

... "Oh, we actually have to think, take a step back and think." You get what I'm saying, right?

Swyx1:02:29

Yeah, yeah.

Beyang Liu1:02:29

Like, i- is, is the transformer gonna be, like... Is that the end-all be-all of architectures and it's just a matter of scaling that?

Swyx1:02:35

Yeah.

Beyang Liu1:02:35

Or are there other algorithms and, and, like, that is gonna be one piece of, uh, a, a system of intelligence that's gonna take advantage, that, that will have to take advantage of, like, many other algorithms and, and approaches.

Swyx1:02:47

Yeah, we shall see. Maybe John Carmack will find it.

Beyang Liu1:02:52

Yeah.

Swyx1:02:54

Um, all right. So sorry for that digression.

Beyang Liu1:02:55

No, no.

Swyx1:02:56

I, I'm just, uh, very curious. Uh, so one thing, uh, I, I did actually want to check in on, uh, because you... We talked a little bit about code graphs and, like-

Beyang Liu1:03:03

Mm-hmm

Swyx1:03:03

... reference graphs and all that. Do you actually use a graph database? No, right?

Beyang Liu1:03:06

No.

Swyx1:03:07

Isn't that very-

Beyang Liu1:03:08

Well, I mean, like, it... How would you find a graph database?

Swyx1:03:11

We, yeah- ... we use Postgres.

Beyang Liu1:03:12

Yeah.

Steve Yegge1:03:12

And, uh, yeah, I saw a paper actually right after I joined Sourcegraph. There was some joint study between IBM and some other company that basically showed that Postgres was performing as well as most of the graph databases for most graph workloads.

Swyx1:03:23

Wow.

Steve Yegge1:03:24

Yeah.

Beyang Liu1:03:24

In v0 of Sourcegraph, we're like, "We're building a co- code graph. Let's use a, a graph database."

Swyx1:03:29

Yes. Yeah.

Beyang Liu1:03:30

Uh, I won't name the database 'cause, I mean, it was like 10 years ago, so they're probably much better now. But, like, we basically tried to dump, um, like, a, a non-trivially sized, like, data set, but also, like, you know, not, not the whole universe of code, right?

Like, it was a relatively small, uh, data set compared to what we're indexing now into the database, and it was just... It... We let it run for, like, a week, and it, I think it, like, segfaulted or something.

And we're like, "Okay, uh, let's try a- another approach. Um, let's just put everything in Postgres." And these days, like, the graph, uh, data, I mean, it's, it's partially in Postgres. It's partially just, uh, I mean, you can store them as, like, flat files.

Steve Yegge1:04:08

Yep.

Beyang Liu1:04:08

Uh, I mean, at the end of the day, all a database is, like, just get me the data I want.

Swyx1:04:12

instance you need.

Beyang Liu1:04:12

Like, answer the queries- ... that I need, right? Like, if all your queries are, like, you know, uh, single hops, uh, in, in this, in this, uh-

Steve Yegge1:04:20

Which they will be if you denormalize in, in-

Beyang Liu1:04:22

Yeah

Steve Yegge1:04:23

... your use case.

Beyang Liu1:04:23

Exactly. Um, and-

Swyx1:04:25

Interesting.

Beyang Liu1:04:26

So yeah.

Swyx1:04:27

S- ... seventh normal form is just a bunch of files on your SD.

Beyang Liu1:04:29

Yeah, yeah. And I don't know. Like, I feel like there's a bunch of stuff like that where it's like, if you look past the marketing and think about, like, the actual, uh, query load or, like, the traffic patterns or the end user use cases you need to serve, um, y- just go with, like, the tried and true, kinda, like, dumb classic tools over kinda, like, the new age stuff.

Swyx1:04:51

2.0 technology, yeah.

Beyang Liu1:04:52

I mean, there, there's a bunch of stuff like that in the search domain too, ri- especially right now with, like, you know, embeddings and, and vector search and, and, and all that. Uh, but-

Swyx1:05:01

Yeah

Beyang Liu1:05:01

... you know, like, classic search techniques still go very far and, um-

Swyx1:05:05

Yeah

Beyang Liu1:05:05

... I don't know. I, I think in the next year or two maybe as, like, the... As, as we get past, like, the peak AI hype, we'll, we'll start to see, uh, the, the, the gap emerge or become more obvious to, to more people about, like, how, how, how many of, like, the newfangled techniques actually work in practice and, and yield a better product experience day to day.

Swyx1:05:25

Yeah. Uh, so speaking of which, like, you know, obviously there's a bunch of other people trying to build AI tooling.

Beyang Liu1:05:30

Yep.

Swyx1:05:30

Uh, what can you say about your, your AI stack? Um, what do... Like, obviously you build a lot proprietary in-house, uh, but, like, what approaches, uh, you know... Like, so prompt eng- prompt engineering, do you have a prompt engineering management tool?

You know, what, what approach is there? Do you, do you do, um, prepro- preprocessing orchestration? Like, do you use Airflow? Do you use something else? Like, you know, that kind of stuff.

Beyang Liu1:05:54

Yeah. Uh, ours is very, like, duct taped, uh- ... together at the moment. Um, so in, in terms of stack, uh, I mean, it's, it's essentially, uh, Go and TypeScript, uh- ... and now Rust. Um, there's the, the knowledge graph, the code knowledge graph that we built, which is using indexers, uh, many of which are open source, um, that speak the Skip protocol.

Uh, and, uh, we have the code search backend. Um, you know, traditionally we supported regular expression search and, uh, uh, string literal search with, like, a trigram index, and we're also building more, like, fuzzy search on top of that now, uh, kinda like natural language or keyword-based search on top of that.

Um, and we use a variety of open source and proprietary models. We, we try to be, like, pluggable with respect to different models, so we can easily kinda, like, swap-

Swyx1:06:45

Yeah

Beyang Liu1:06:45

... swap the latest model in and out, uh, as they come online. Um-

Swyx1:06:50

I'm just hunting for like-

Beyang Liu1:06:51

Yeah

Swyx1:06:51

... are there anything... Is there anything out there that you're like, "These are... These guys are really good. You should... You... Everyone else should-

Beyang Liu1:06:56

Oh, I see

Swyx1:06:56

... check them out"? So for example, you talked about recursive summarization-

Beyang Liu1:06:59

Yeah

Swyx1:07:00

... which is something that LangChain and LlamaIndex do. I presume you wrote your own. Uh, I presume you-

Beyang Liu1:07:04

Yeah, we wrote our own. I, I think, like, the stuff that, uh, LlamaIndex and, and LangChain are, are doing are, like, super interesting. I think from our point of view, it's like we're still in the application, like, end user use case discovery phase.

Swyx1:07:17

Yeah.

Beyang Liu1:07:17

And so adopting, like, a, an external, um, infrastructure or, or middleware, uh, kind of- ... tool just seems like overly constraining right now, 'cause like we're-

Swyx1:07:28

Yeah, we need full control.

Beyang Liu1:07:29

Yeah, we need full control 'cause we need to be able to iterate rapidly up and down the stack.

Swyx1:07:32

Yeah.

Beyang Liu1:07:32

Um, but maybe at some point there'll be like a, a convergence and we can actually merge some of our stuff into theirs and turn that into a common resource. Um, in terms of like other, uh, vendors that we use, I mean, obviously like, uh, nothing but good things to say about Anthropic and OpenAI, uh-

Swyx1:07:47

Ah

Beyang Liu1:07:47

... which we both, uh, kinda partner with and, and use.

Swyx1:07:50

Yeah.

Beyang Liu1:07:50

Um, also a plug for Fireworks as an inference platform.

Swyx1:07:54

Mm-hmm.

Beyang Liu1:07:54

Um, uh, their team was kinda like ex- ex-Meta people who, uh, basically know all like the, the, the bag of tricks for making inference fast.

Swyx1:08:03

Yeah, I met, I met Lynn. So sh- she was apparently-

Beyang Liu1:08:04

Lynn is great

Swyx1:08:05

... the, uh, she was like with Sumith. She, she was like the co-manager-

Beyang Liu1:08:08

Yep

Swyx1:08:08

... of PyTorch for five years.

Beyang Liu1:08:09

Yeah, yeah, yeah. Um, and-

Swyx1:08:10

So, but like is their main thing that w- we just do fastest inference on earth? Is, is that like, is that what it is, or?

Beyang Liu1:08:16

I think that's the pitch.

Swyx1:08:17

Okay.

Beyang Liu1:08:17

Um, and it keeps getting faster, uh, somehow. Like we run StarCoder, uh, on, on top of Fireworks, and that's made it so that we don't, just don't have to think about, uh, building up a, an inference stack. And so that's great for us 'cause it allows us to fo- focus more on, uh, the, the kinda like data fetching, the knowledge-

Swyx1:08:35

Yeah

Beyang Liu1:08:35

... graph, and, uh, model fine-tuning, which we've also, uh, invested a bit in. Um-

Steve Yegge1:08:40

That's right. We've got multiple AI work streams in progress now because we hired a head of AI finally.

Beyang Liu1:08:46

Yay.

Steve Yegge1:08:46

We spent close to a year, actually. I think I talked to probably like 75 candidates.

Beyang Liu1:08:52

Yeah.

Steve Yegge1:08:52

And, uh, our, the guy we hired, Rishab, is, uh, um, absolutely world-class, and he started, immediately started, uh, multiple work streams, including he's fine-tuned StarCoder already. Uh, he's, uh, he's got prompt engineering work stream. He's got, uh, uh, the embeddings work stream.

He's got evaluation and experimentation. Benchmarking, wouldn't it be nice if Cody was on huggy f- Hugging Face with a, you know, with a, uh, a benchmark that we could just... Anybody could say, "Well, we'll run, we'll run against the benchmark, or, or we'll make our own benchmark if we don't like yours."

But we'll, we'll be forcing people into the sort of quantitative, you know, comparisons. Yeah. And that's, that's all happening under the AI program that he's building for us.

Swyx1:09:29

Yeah. Uh, I sh- I should mention, by the way, I've heard that there's a V2 of StarCoder g- coming on. Uh, so you guys should talk to Hugging Face.

Beyang Liu1:09:36

Cool. Awesome.

Steve Yegge1:09:37

Great.

Swyx1:09:38

Uh, I actually visited their offices in Paris, which is where I heard it.

Beyang Liu1:09:41

That's awesome.

Steve Yegge1:09:41

I mean, can you guys believe how amazing it is that the open source-

Beyang Liu1:09:44

Free models

Steve Yegge1:09:45

... models are like-

Beyang Liu1:09:45

Yeah

Steve Yegge1:09:46

... competitive with, you know, GPT and Anthropic?

Beyang Liu1:09:49

Yeah.

Steve Yegge1:09:49

I mean, it's nuts, right? I mean, uh, that one Googler that was predicting that, right, open source would catch up.

Beyang Liu1:09:55

Yeah.

Steve Yegge1:09:55

At least he wa- he was right for completions.

Beyang Liu1:09:58

Yeah, I mean, for completions, open source is, is state-of-the-art right now.

Swyx1:10:01

Yeah, you, you, you were on OpenAI, then you went to Claude, and now you've ripped it out.

Beyang Liu1:10:04

Yeah. Yeah, for completions.

Swyx1:10:06

Yeah.

Beyang Liu1:10:06

I mean, we still use, uh, Claude and, and GPT-4 for, uh, chat and, and also commands.

Swyx1:10:11

Yeah.

Beyang Liu1:10:12

Um, but, you know, they- there'll be... Like, the ecosystem's gonna continue to, to evolve, uh, evolve. We obviously love the, the open source ecosystem, and like you shout out to, to Hugging Face, and, and also, like, Meta Research.

Uh, we love the work that they're doing in, in kind of driving the ecosystem forward.

Swyx1:10:27

Yeah, you didn't mention Code Llama.

Beyang Liu1:10:29

We're not using Code Llama currently. Um, it's k- always kinda like a constant evaluation process. So like, I don't wanna come out and say like, "Hey, this model's the best 'cause we chose it." It's basically like we did a bunch of like tests for the sorts of like context that we're fetching now, and given the way that our prompt's constructed now.

And at the end of the day, it was like a judgment call, like StarCoder seemed to work the best, and that's why we adopted it. Um, but it's sort of like a continual process of revisitation. Like, if someone comes up with like a neat new, like, context fetching mechanism, and we have a couple coming online, uh, soon, then it's always like, okay, let's try that against the, the kind of like array of models that are available and see, you know, uh, how this moves the needle across, uh, that set.

Swyx1:11:10

Yeah. What do you wish someone else built?

Steve Yegge1:11:14

Uh, what did we have to build that we wish we could've used?

Beyang Liu1:11:18

Is that, is that the question? Uh, interesting.

Swyx1:11:20

This is a request for startups.

Beyang Liu1:11:24

I mean, if someone could just provide like a very nice, clean data set of, uh, uh, both naturally occurring and synthetic, uh, code data out there.

Steve Yegge1:11:35

Yeah, can someone please give us their data moat?

Beyang Liu1:11:38

Well, not even the data moat. It's just like, I feel like most models today, they still use like combination of like the stack and the pile, uh, as like, uh, their, their training corpus. Um, but you can only stretch that so far.

At some point, we need more data. Uh, and um, I don't know. I, I think there's still more alpha in like synthetic data. Like, we have a couple efforts where like we think fine-tuning some models on specific coding task will yield alpha, will yield more kinda like reliable, uh, code generation of the sort where it's like reliable enough that we can, we can fully automate it, at least like the one-hop-

Swyx1:12:13

Mm

Beyang Liu1:12:13

... uh, thing. Um, and synthetic data is, is playing a part of that. But I mean, if there were like a synthetic data provider. I, I don't think you could construct a provider that has access to like some proprietary, uh, code base.

Like, no company in the world would, would be able to like sell that to you. But like anyone who's just like providing clean data sets off of the, the publicly available data, uh-

Swyx1:12:33

Yeah

Beyang Liu1:12:34

... that would be nice.

Swyx1:12:36

Yeah. Uh, yeah.

Beyang Liu1:12:36

I don't know, I don't know if there's a business around that, but like that's something that we'd definitely like love to use.

Steve Yegge1:12:40

Oh, for sure. My God. I mean, but that's, that's also like the secret weapon, right? For any AI, you know, is, is the, the data that you've curated. So I doubt people are gonna be- Oh, we'll see. You know?

But, but we can maybe contribute, you know, if we wanna have a benchmark of our own. You know?

Beyang Liu1:12:57

Yeah.

Swyx1:12:57

Yeah. Uh, I would say like that, that was, that would be the bull case for Replit, uh, that like you, you want to be a coding platform where you also offer bounties, um, and, and like then you eventually bootstrap your own proprietary set of co- coding data.

I, I don't, I don't think they'll ever share it, and it's, uh... The, the rumor is, and this is from no- nobody of, uh, at Replit that, that I'm, that I'm hearing, but like also like they're, they're just not leveraging that, uh, actively.

Like, they're actually just betting on OpenAI, OpenAI to do a lot of that.

Beyang Liu1:13:25

Mm.

Steve Yegge1:13:25

Mm.

Swyx1:13:25

Which would... Be- betting on OpenAI- ... I think, uh, you know, has been a winning strategy so far.

Beyang Liu1:13:31

Yeah. They're, they're definitely great at executing and-

Steve Yegge1:13:34

Executing their CEO.

Beyang Liu1:13:36

Ooh.

Steve Yegge1:13:37

And then bringing him back in four days.

Beyang Liu1:13:39

Yeah, yeah.

Steve Yegge1:13:40

He went-

Beyang Liu1:13:40

That was a whole, like, uh-

Steve Yegge1:13:42

He went

Beyang Liu1:13:42

... I don't know

Swyx1:13:42

Did, did you guys, like, uh... Yeah, was, was the company, like, just obsessed by the drama? Like, uh, like we were unable to work. I just walked in after, after it happened, and this whole room, in the new room, was just like everyone's just staring at their phones.

Beyang Liu1:13:54

Yeah, I mean, it was, it was a bit difficult to ignore. Um, I mean, it would have real implications for us too, 'cause, like, we're using them, and so-

Swyx1:14:01

Yeah

Beyang Liu1:14:02

... there, there's a very real question of like, uh, do we have to, like, do a quick-

Steve Yegge1:14:04

Yeah, do you... Yeah, Microsoft, like you just move to Mic- Microsoft, right?

Beyang Liu1:14:07

Yeah, I mean, that would, that would've been, like, the break glass, uh, plan. If, uh, the worst case played out, then I think we'd have a lot of customers, um, you know, the day after being like, "You know, how can you guarantee the reliability of your services, uh, if, if the company itself isn't, isn't stable?"

Swyx1:14:22

Yeah.

Beyang Liu1:14:23

But I'm really happy they got things sorted out and, uh, things are stable now, 'cause, like, they build really cool stuff, and we love using their, their tech.

Steve Yegge1:14:30

Yeah. Awesome.

Swyx1:14:31

So we kind of went through everything, right? Sourcegraph, Cody, uh, why agents don't work, why, uh, inline completion- ... uh, is better, all, all of these things. How does that bubble up to who manages the people, right? Because as engineering managers and, uh, you know, I never...

For Managers1:14:35

Swyx1:14:51

I, I didn't write much code. I was mostly helping people write their own code, you know? So-

Beyang Liu1:14:55

Yeah

Swyx1:14:55

... even if you have the best inline completion, it doesn't help me do my job.

Beyang Liu1:14:59

Yeah.

Swyx1:14:59

Um, w- what's kind of the, the future of Sourcegraph in the engineering org?

Beyang Liu1:15:04

Yeah, so that's a really interesting question. Um, and I think it's... It sort of gets at this, like, issue, which is, I think, uh, basically, like, every AI, uh, dev tools creator or producer these days, I think us included, um, we're kind of, like, focusing on the wrong problem in a way.

Um, because, like, the, the, the real problem of modern software development, I think, is, is not how quickly can you write more lines of code. It's really about managing the emergent complexity, uh, of code bases as they evolve, uh, and grow, and how to get...

How to make, uh, like, efficient development tractable again. Because, uh, the bulk of your time becomes more about understanding how the system works, uh, and how the pieces fit together currently so that you can update it in a way, uh, that gets you your added functionality, um, doesn't break anything, and doesn't introduce a lot of additional complexity that will slow you down in the future.

Um, and if anything, like, the inner loop developer tools that are all about, like, generating lines of code, uh, yes, they help you get your feature, uh, done faster. They, they generate a lot of boilerplate for you. But they might make this problem of, like, managing large complex code bases, uh, more challenging, just because now you'll, instead of having, like, you know, a pistol, you'll have a machine gun in terms of-

Swyx1:16:31

Mm-hmm

Beyang Liu1:16:31

... like being able to write, write code. And there's going to be a, a bunch of, like, natural language prompted code that is generated in the future that was produced by someone who doesn't even have, like a, uh, an understanding of, of source code.

And so, like, how are you going to verify the quality of that and, and make sure it, it not only checks the kind of like low level boxes, but also fits architecturally, uh, in a, in a way that's sensible into your code base.

And so I think as we look forward to the future of the next year, we, we have a lot of ideas around how to make code bases, as they evolve, more, uh, understandable and manageable to the people who really care about the code base as a whole, uh, you know, tech leads, engineering leaders, folks like that.

And that... It, it is kind of like a return to, uh, our, our ultimate mission at Sourcegraph, which is to make code accessible to all. It's not really about, you know, enabling people to write code. And if anything, like, the original version of Sourcegraph was a rejection of like, "Hey, let's stop trying to build, like, the next best editor," um, because, like, there's already enough people doing that.

Um, the, the real problem that we're facing, I mean, Quinn, myself, and you, Steve, uh, at, at Google, was like, how do we make sense of the code that exists so we can understand enough to know what code needs to be written?

Steve Yegge1:17:46

Yeah. Well, I tell you what customers want, right, and what they're going to get. What they want is for Cody to have a monitor for developer productivity. And any developer who falls below a threshold, a button lights up where the admin can fire them.

Pow. Or, or Cody will even press that button for you- ... if enough time passes. But, uh, what... I'm kind of only half tongue in cheek here.

Beyang Liu1:18:06

Yeah.

Steve Yegge1:18:06

We've got some, some prospects who are kind of like sniffing down that, that avenue, and we're like, "No." Um, but, uh, what they're going to get is, uh, g- much like, like Beyang was saying, much greater whole code base understanding, which is actually something that, that Cody is, I, I would argue, the best at today in the coding assistant space, right?

'Cause of our search engine and, and the techniques that we're using. And that whole code base understanding is so important, you know, for, uh, you know, for any sort of a manager who just wants to get a feel for the architecture or potential security vulnerabilities, or whether, you know, people are writing code that's well tested and et cetera, et cetera, right?

And, um, just, uh, solving that problem is tricky, right? This is the... This is not the developer inner loop or outer loop. It's like the manager inner loop? No, outer loop. The inner... The manager inner loop is staring at your belly button, I guess.

So in any case, uh-

Beyang Liu1:18:54

Waiting for the next Slack message to arrive.

Steve Yegge1:18:57

Yes. What they really want is a batch mode for these assistants, where you can actually take the coding assistant and shove its face into your code base, you know, and, and six billion lines of code later, right, it's told you all the security vulnerabilities.

That's what they really actually want. It's an insanely expensive proposition, right? You know, just the, the GPU cost-

Beyang Liu1:19:16

Mm-hmm

Steve Yegge1:19:16

... especially if you're doing it on a regular basis. So it's better to do it at the point the code enters the system. And so now we're starting to get into developer outer loop stuff, and I think that's, that's where a lot of the...

To your question, right? A lot of the admins and managers and so- you know, the decision-makers, anybody who just, like, kind of isn't coding but is involved-

Beyang Liu1:19:32

Uh, they're gonna have, uh, I, I think, uh, well, a set of, a set of tools, right? A set of, of... Just like with code search today. Code search, our code search actually serves that audience as well.

Swyx1:19:44

Mm-hmm.

Beyang Liu1:19:44

The CIO types, right?

Swyx1:19:45

Mm-hmm.

Beyang Liu1:19:45

You know, 'cause they're just like, "Oh, hey, I wanna see how we do, you know, SamLoft," and they use our-

Swyx1:19:49

Yeah

Beyang Liu1:19:49

... search engine and they go find it, you know?

Swyx1:19:50

Yeah.

Beyang Liu1:19:50

And AI's just gonna make that so much easier for them.

Swyx1:19:54

Yeah. Uh, I have a... This is the, my perfect place to put my anecdote of how I used Cody yesterday. Um, I was actually trying to build this sort of Twitter scraper thing, and, uh, Twitter's notoriously very challenging to work with, um, because they don't want to work with you- ...

with anyone. Um, and there's a, there's a repo that I wanted to, to inspect. It, it was, it was really big that, that had a Twitter, Twitter scraper thing in it. Um-

Beyang Liu1:20:16

Mm-hmm

Swyx1:20:17

... and I pulled it into Copilot, didn't work. Uh, and, and, uh, but then I, I noticed that on your landing page you had a web version. Like, I, I typically think of-

Beyang Liu1:20:26

Yeah

Swyx1:20:26

... Cody as a Chrom- uh, VS Code extension.

Beyang Liu1:20:28

Yeah.

Swyx1:20:28

But you have a web version where you can just plug in any repo in there and just, uh, talk to it, and, uh, that's what I used to, to figure it out. So.

Beyang Liu1:20:34

Yeah.

Steve Yegge1:20:35

Wow, a Cody web.

Beyang Liu1:20:36

Cody web.

Steve Yegge1:20:36

In the wild.

Swyx1:20:37

Yeah.

Beyang Liu1:20:38

I, I mean, we've do- we've done a very poor job of, uh, making-

Swyx1:20:42

I found it-

Beyang Liu1:20:42

... the existence of that feature

Swyx1:20:43

... it's not, it's not, it's not easy to find.

Beyang Liu1:20:44

It's not easy to find.

Swyx1:20:44

But if you go through, like, the search thing, it's like, oh, this is old SourceGraph.

Beyang Liu1:20:47

Yeah.

Swyx1:20:47

You don't wanna look at old SourceGraph anymore. You can use SourceGraph, all the AI stuff. Uh, old SourceGraph has AI stuff-

Beyang Liu1:20:52

Yeah

Swyx1:20:52

... and it's Cody Web, and like, that's-

Beyang Liu1:20:54

Yeah, yeah. There's a little, like, Ask Cody button that's-

Swyx1:20:56

Yeah

Beyang Liu1:20:56

... kinda, like, hidden in the upper right-hand corner.

Swyx1:20:58

Yeah, yeah.

Beyang Liu1:20:58

Uh, we should make, we should make that more visible. It's, it's definitely one of those, like, aha moments when you can ask a question of, of the base.

Swyx1:21:03

Of any repo, right?

Beyang Liu1:21:04

Yeah, yeah.

Swyx1:21:04

Because you already indexed it.

Beyang Liu1:21:06

Yeah.

Swyx1:21:06

Well, you didn't embed it, but you indexed it.

Beyang Liu1:21:07

Yeah.

Steve Yegge1:21:08

Mm.

Beyang Liu1:21:08

And, and there's actually some use cases that have emerged among your power users where they kinda do... Like, you're familiar with, like, v0, uh-

Steve Yegge1:21:15

Yeah

Beyang Liu1:21:15

... net dev?

Steve Yegge1:21:15

Mm-hmm.

Beyang Liu1:21:16

Like, you can kind of replicate that, but for, like, arbitrary frameworks and, and libraries with, with Cody Web. 'Cause there's, there's also, like, an equally hidden toggle which you may not have discovered yet-

Steve Yegge1:21:24

Yeah

Beyang Liu1:21:24

... where you can actually tag in multiple repositories as context.

Steve Yegge1:21:26

Yeah.

Beyang Liu1:21:27

And so you can do things like, like we have a, a demo path where it's like, okay, let's say you wanna build, like, a stock ticker, uh, that's React-based but uses this, like, one, like, tick data fetching API.

It's like you tag both repositories in, you ask it. It's, like, two sentences. Like, build a stock tick app, track the tick data of, like, Bank of America, Wells Fargo over the past week. Uh, and it generates the code.

You can paste that in, and it just, it, it works, uh, magically.

Steve Yegge1:21:53

Yeah.

Beyang Liu1:21:53

Um, we'll probably invest in that more just because, like, the, the wow factor of that is, is just pretty incredible.

Steve Yegge1:21:58

Yeah.

Beyang Liu1:21:58

It's like, what if you can speak apps into existence that use, like, the frameworks and packages that, like-

Steve Yegge1:22:04

God

Beyang Liu1:22:04

... you wanna use?

Steve Yegge1:22:05

Yeah.

Beyang Liu1:22:05

Um.

Swyx1:22:07

It's not even fine-tuning. It's just taking advantage of your RAG pipeline.

Beyang Liu1:22:09

Yeah, it's just RAG.

Swyx1:22:10

Yeah.

Beyang Liu1:22:10

Uh, RAG is all you need for-

Swyx1:22:12

It's all you need

Beyang Liu1:22:12

... many things.

Steve Yegge1:22:14

Uh, it's not just RAG. It's RAG, right? RAG's good. Not a, not a fallback.

Beyang Liu1:22:21

Yeah. But I guess, like, getting back to the, the original question, I think there's, there's a couple things I think would be interesting for engineering leaders. One is, is the use case that you called out, is, like, all the stuff that you currently don't do that you really ought to be doing with respect to, like, ensuring code quality or updating dependencies or, um, it, like, keeping, keeping things up to date.

Like, the things that, like, humans find toilsome and tedious and just, like, don't wanna do, but, like, would really help up level the quality, security, and robustness of your code base. Like, now we potentially have a way to do that with machines.

I think there's also this other, um, thing, and this gets back to the, the, the point of, like, you know, how do you measure developer productivity? It's, like, the, the perennial age-old question. Like, every CFO in the world would love to, to, to do it in the same way that you can measure, you know, marketing or, or sales or other parts of the organization.

And I think, like, like, what is the, like, actual way you would do this that, that is good, and if you had all the time in the world? I think as, like, an engineering manager or an engineering leader, what you would do is you would go read through the Git log.

Like, maybe, like, line by line be like, "Okay, you know, you, uh, you know, Sean, these are the features that you built over the past, you know, six months or, or year. Um, these are the things that deliver, that you helped drive.

Here's the, the stuff that you, you did to help your teammates. Um, here, here are the reviews that you did that helped ensure that we have, uh, maintain a, a, you know, coherent, um, and high-quality, uh, code base."

Um, now connect that to the things that matter to the business. Like, what were we trying to drive, uh, this? Was it, like, engagement? Was it revenue? Was it, you know, adoption of some new product line? And really, like, weave that story together.

Like, the work that you did had this impact on the metrics that moved the needle for the business, and ultimately show up in, you know, revenue or stock price or whatever it is that's, you know, at the very top of any for, for-profit organization.

And, like, you could, in theory, do all that- ... today if you had all the time in the world.

Steve Yegge1:24:24

Yeah.

Beyang Liu1:24:24

But as an engineering leader-

Swyx1:24:26

It's too busy building.

Beyang Liu1:24:26

Yeah, you're too busy building. You're too busy with a d- bunch of other stuff. Plus, it's also, like, tedious. Like, reading through, you know, a Git log and trying to, like, understand, like, what a change does and summarizing that-

Steve Yegge1:24:36

Yeah

Beyang Liu1:24:36

... um, it's just r- it's, it's, it's not the most exciting work in the world. Um, but with the benefit of, of, of AI, I think we, y- you could conceive of a system that actually does a lot of the tedium and, and helps you actually tell that story, and I think that is maybe the ultimate answer to how, how we get at, like, developer productivity in, in a way that, like, a CFO would be like, "Okay," like, "I can buy that," right?

Like, the work that you did, uh, impacted these core metrics because, you know, these features were tied to those, and therefore, you know, we can a- afford to invest more in this part of the organization, and that's what we really wanna drive towards.

I think that's-

Steve Yegge1:25:13

Yeah

Beyang Liu1:25:13

... that's what we've been trying to build all along in a way with Sourcegraph. It's this kinda, like, code-based level of understanding, um, and, and the availability of, you know, LLMs and, and AI now just, like, puts that, uh, much sooner in reach, I think.

Steve Yegge1:25:26

Yeah. But, I mean, you know, we have to focus also. Small company, you know? And so, uh, you know, our, our short-term focus is lovability, right?

Beyang Liu1:25:35

Yeah.

Steve Yegge1:25:35

I mean, we just, we absolutely have to make Cody, like, this is... Everybody wants it, right? Uh, but, uh, absolutely, Sourcegraph is all about enabling, uh, m- all of the d- non-engineering, you know, roles, decision-makers, and, and so on.

And, uh, and as Beyang says, I mean, I think there's just a lot of opportunity there once we've built a lovable Cody.

Alessio1:25:55

Awesome. Um, we wanna jump into lightning round?

Swyx1:25:58

Lightning round.

Alessio1:25:59

Okay. Which we always forget to send the, the questions ahead of time. Um, so we usually have three, one a- around acceleration, exploration, and then a final takeaway. So the, the acceleration one is, what's something that already happened in AI that is possible today that you thought would take much longer?

Lightning Round1:26:00

Beyang Liu1:26:17

I mean, just LMs and, and, uh, how good the, the vision models are now. Like I, I got my start-

Swyx1:26:23

Oh, vision. Okay.

Beyang Liu1:26:24

Yeah. Well, I mean, like, uh, uh, back in the day, um, like I, I got my start in machine learning, uh, in computer vision, but circa like 2009, 2010. Uh, and in those days, everything was like statistical based.

N- neural nets had not yet made their, their comeback. Uh, and so nothing really worked. And so I was very bearish after that experience on, on the future of computer vision. But like, man, the progress that's been made just in the past like, uh, three, four years, has just been absolutely, uh, astounding.

So yeah, it, it came up faster than I expected it to.

Steve Yegge1:27:00

Yeah, multimodal in general I think is, uh, um... I think there's a lot more capability there that, that we're not tapping into, uh, potentially even in the coding assistant space. And, uh, you know, honestly, I think that the form factor that coding assistants have today is probably not the steady state that we're seeing, you know, long term.

I mean, you'll, you'll always have completions, and you'll always have chat and commands and so on, but I think we're gonna discover a lot more. And I think multimodal potentially opens up some kinda new ways to, you know, get your, get your stuff done.

So yeah, I think the capabilities are there today. And they're just... It's just shocking. I mean, like, I still am astonished when I sit down, you know, and I have a conversation with the LLM with the context, and, and, and, a- and it's like I'm talking to a, you know, a senior engineer or an architect or somebody, right?

And I can, I can bounce ideas off it. And I think that people have very different working models with, with these assistants today. You know, some people are just completion, completion, completion. That's it. And if they want some code generated, they write a comment, and then, then-

Swyx1:27:56

Mm

Steve Yegge1:27:57

... you know what I mean? You're telling it what to do. But I truly think that there are other modalities that we're gonna stumble across, and, uh, just, just kind of latently, you know, uh, you know, inherently built into the LLMs today, that we just haven't found them yet.

They're more of a discovery than an invention, you know?

Swyx1:28:13

Like other usage patterns.

Steve Yegge1:28:14

Absolutely.

Swyx1:28:15

Mm.

Steve Yegge1:28:15

I mean, the, the one that we talked about earlier, Nonstop Cody, is one-

Swyx1:28:18

Mm-hmm

Steve Yegge1:28:18

... right? Where you can just kick off a whole bunch of, you know, requests to refactor and so on. But, uh, you know, there could be any number of others. You know, we talk about agents, you know, that's kinda out there, but I think there are kinda more inner loop type ones to be found.

And, uh, and see, we, we haven't looked at all at multimodal yet.

Swyx1:28:35

Yeah. Uh, for sure, like there's a, um... There, there's two that come to mind just, just off the top of my head. Um, one, which is, um, effectively architecture diagrams and entity relationship diagrams.

Steve Yegge1:28:46

Mm-hmm.

Swyx1:28:46

Um, you probably... It's probably... There's probably more alpha in like synthesizing them for management to see.

Steve Yegge1:28:52

Ooh.

Beyang Liu1:28:53

Yeah.

Swyx1:28:53

Which is, like, you don't need AI for that. You can, you can just use your reference graph.

Beyang Liu1:28:57

Yeah.

Swyx1:28:57

Uh, but then also doing it the other way around when, like someone draws stuff on a whiteboard-

Beyang Liu1:29:01

Yeah

Swyx1:29:01

... and actually generating code.

Steve Yegge1:29:02

Well, you can, you can, you can, you can generate the, the diagram, and then, you know, e- explanations as well.

Swyx1:29:07

Yeah. Uh, and then the other one is, uh, there was a demo that went pretty vir- viral like a, uh, two, three weeks ago about, uh, how someone just had an always on script just screenshotting and sending it to GPT Vision-

Steve Yegge1:29:19

Mm

Swyx1:29:20

... um, every, like on, on some kind of time interval, and it would just au- autonomously suggest stuff.

Beyang Liu1:29:24

Yeah.

Steve Yegge1:29:25

Mm.

Swyx1:29:25

Um-

Steve Yegge1:29:25

Great stuff

Swyx1:29:25

... so, like, no trigger, just, just watching your screen-

Beyang Liu1:29:27

Yeah

Swyx1:29:28

... and just, like, um, being a, being a real copilot rather than having- ... you initiate with, with a chat.

Beyang Liu1:29:33

Yeah, yeah.

Swyx1:29:33

Um, so there's some, there's some-

Beyang Liu1:29:35

It's like the return of Clippy, right?

Swyx1:29:37

Return of Clippy.

Beyang Liu1:29:37

But, but actually good.

Swyx1:29:39

Um, we actually... So, uh, the reason I know this is that we actually did a hackathon where, um, we did, we had, we, we wrote that, that project, but it roasted you while you did it. So it's like, it's like, "Hey," it's like, "You're, you're, you're on Twitter right now.

You should be coding." Like... Uh, and that, that can be a fun copilot-

Beyang Liu1:29:55

Yeah

Swyx1:29:55

... thing as well.

Beyang Liu1:29:55

Yeah, yeah.

Swyx1:29:56

Um, okay, so I'll, I'll, I'll jump on. Um, exploration, uh, what do you think is the most interesting unsolved question in AI?

Beyang Liu1:30:03

I mean, I think-

Steve Yegge1:30:03

It used to be scaling, right? With CNNs and RNNs-

Beyang Liu1:30:06

Yeah

Steve Yegge1:30:06

... and Transformer solved that.

Beyang Liu1:30:08

Yeah.

Steve Yegge1:30:08

So what's the next big hurdle that's keeping GPT-10 from emerging?

Beyang Liu1:30:12

I mean, do, do you mean that like-

Swyx1:30:13

Ooh, that seems like a safety-ist argument.

Beyang Liu1:30:15

I, I feel like... Do you mean the, like the, the pure model, like AI layer, or-

Swyx1:30:19

No, it doesn't have to be. You can just like-

Beyang Liu1:30:20

For me, personally, it's like how, how do you get reliable, like first try working code generation?

Swyx1:30:25

Mm.

Beyang Liu1:30:26

Uh, uh, even like a single hop, like write a function that does this, because I think, like in order... If you wanna get to the point where you can actually be truly agentic or, or like multi-step automated, uh, a necessary part of that is, like the single step has to be robust and reliable.

Uh, and so I, I think that's the problem that, like, we're focused on solving right now, because once you have that, it's a building block that you can then compose in, into longer chains.

Alessio1:30:55

Um, and just to wrap things up, what's one message, takeaway that you want people to, to remember and think about?

Beyang Liu1:31:06

Um, I mean, I, I think for me, it's just like, uh, the best dev tools in the future are gonna have to leverage many different forms of intelligence. Uh, you know, calling back to that like Normsky, uh, architecture, which trying to make catch on.

Swyx1:31:20

You should call it something cool, like, like S-Star or RSI.

Beyang Liu1:31:23

Yes, yes, yes.

Swyx1:31:24

You know? Uh, just, just one letter, and then just let people speculate what it is.

Beyang Liu1:31:27

Yeah, yeah. What could he mean? Um- But I don't know. Like in terms of like trying to describe what we're building, we, we try to be a little bit more like down to earth and, and like straightforward, and, and I think like Normsky kinda like encaps- encapsulates like the, the, the, the two big, like, technology areas that we're investing in that we think will be, uh, very important for producing really good, uh, dev tools, and I think that that's a big differentiator, uh, that we view that, that Cody has right now.

Steve Yegge1:31:57

Yeah. And my- mine would be, uh, I know for a fact that not all developers today are using coding assistants, yeah? And, uh, that's probably because they tried it, and, uh, it didn't, you know, immediately write a bunch of beautiful code for them, and they were like, "Ah, too much effort," and they, and they left, right?

Swyx1:32:16

Mm.

Steve Yegge1:32:16

Well, my big takeaway from this talk would be, if you're one of those engineers, man, you be- you better start like, you know, planning another career. Okay? Because this stuff is the future, and in its hon- honestly, it takes some effort to actually make coding assistants work today, right?

You have to... You know, just like talking to GPT, they'll give you the runaround, just like doing a Google search sometimes. But if you're not putting that effort in and learning the sort of footprint, you know, and the characteristics of how LLMs behave under different, you know, query conditions and so on, if you don't, if you're not getting a feel for the coding assistant, then you're letting this whole train just like pull out of the station and leave you behind.

Beyang Liu1:32:53

Yeah.

Alessio1:32:53

Cool.

Swyx1:32:54

Absolutely.

Alessio1:32:55

Yeah. Thank you guys so much for-

Beyang Liu1:32:56

Thanks for coming on

Alessio1:32:57

... for coming on and being the first guests in the new studio.

Steve Yegge1:32:59

Our pleasure.

Beyang Liu1:33:00

Thanks for having us.

Steve Yegge1:33:01

Thanks for having us.