LALatent SpaceNov 3, 2023· 1:18:54

Beating GPT-4 with Open Source Models - with Michael Royzen of Phind

Michael Royzen, co-founder and CEO of Phind, explains how his company built a GPT-4-beating open-source model for developer Q&A. Royzen recounts founding SmartLens in high school, shifting to NLP after a Hugging Face demo, and creating an internet-scale LLM-powered RAG system in January 2022. He details Phind’s pivot to programmers, the Hacker News launch that gave it 1,500 points, and Paul Graham’s role in naming the company and introducing Ron Conway, who then connected Phind to NVIDIA for GPU access. Royzen argues that Phind’s model, fine-tuned from Code Llama 34B with extra data, closes the gap with proprietary models—especially on code reasoning—and that open-source will win the enterprise because the delta to GPT-5 will be small. He shares how Phind handles multi-step conversations via a pair programmer mode where users can pin messages, and reveals plans for reinforcement learning to reduce hallucination and improve correctness.

  1. 0:00Origin Story
  2. 5:10Building Search
  3. 20:58Product Vision
  4. 26:28Use Cases
  5. 30:00GPT-4 Era
  6. 39:24Custom Model
  7. 47:00UI Features
  8. 51:18PG, Ron & Hardware
  9. 1:04:41Local Models
  10. 1:07:06Autodidact AI
  11. 1:11:51Lightning Round

Powered by PodHood

Transcript

Origin Story0:00

Alessio0:07

Hey everyone, welcome to the Latent Space Podcast. This is Alessio, partner and CPO of Residence and Decibel Partners, and I'm joined by my co-host Swyx, founder of Smol AI.

Swyx0:17

Hey, and today we have in the studio Michael Royzen from Phind. Welcome.

Michael Royzen0:21

Thank you so much. It's great to be here.

Swyx0:22

Yeah, we are recording this in a surprisingly hot October in San Francisco, and, uh, I mean, uh, sometimes the studio works, but, uh-

Alessio0:30

The, the Blue Angels are flying by right now.

Swyx0:31

And the Blue Angels are flying over, so sorry about the noise. I don't think they can hear it. And, and yeah, we, we have enough damping. Anyway, so, um, so welcome. Uh, I've seen Phind blow up this year mostly, I think since your, your launch in Feb and, uh, V2 and then, uh, your, your g- your, um, Hacker News posts.

Um, I- we tend to like to introduce our guests, but then obviously you, you can fill in the blanks with the origin story. Um, so you actually were a high school entrepreneur. You started SmartLens, uh, which is a computer vision startup, uh, in 2017.

Michael Royzen1:01

That's right, yeah. So, um, I remember when like TensorFlow came out and like people started talking about... I-- Oh, obviously at the time the, um, after Alex Net, the deep learning revolution was already, you know, in flow and, um, good computer vision models were a thing.

And what really made me interested in deep learning was, um, I, um, got invited to go to Apple's WWDC conference as a student scholar, um, 'cause I was really into making iOS apps at the time. Um, and so I go there, and I go to this talk where they are, um, they added an API that let people run computer vision models, um, on the device using far more efficient like GP primitives.

Um, and after seeing that, I was like, "Oh, this is cool. Um, this is gonna like have like a big explosion of, you know, different like computer vision models running like locally on the iPhone." Um, and so I had this crazy idea, um, where it was like, what if I could just make this model that could recognize just about anything and have it run on the device?

And that was the genesis for what eventually became SmartLens. Um, I took, um, this dataset called ImageNet 22K. So most people when they think of ImageNet, think of ImageNet 1K, but the int- the full ImageNet actually has, I think, 22,000 different categories.

Yeah. So I took that, um, filtered it, pre-processed it, um, and then did a, did a massive fine-tune on, um, Inception v3, which was I think the state-of-the-art, uh, deep convolutional computer vision model at the time. Um, and to my surprise, it actually worked like insanely well.

I had no idea what would happen if I, it, like I give a single model, um, I think it ended up being 17,000 categories approximately that I collapsed them into. Um, it actually ended up working so well. It worked so well that, um, it actually worked better than Google Lens, um, which released its V1 around the same time.

Um, and so-- And on top of this, the model ran on the device, so it didn't need an internet connection. Um, a big part of the issue with Google Lens at the time was that, um, connections were slower.

You know, 4G was around, but like it wasn't nearly as fast, and so there was a noticeable lag having to upload an image to a server and get it back. Um, but just processing it locally, even on the iPhones of the day in like 2017, much faster.

Um, and so, um, it was a cool little project. Uh, it got some traction. TechCrunch wrote about it, and th- there was kind of like one big, you know, spike in usage, and then over time it tapered off.

Um, but people still pay for it, which is wild.

Swyx3:45

That's awesome. Oh, it's like a monthly or annual subscription?

Michael Royzen3:47

Yeah, it's like a monthly, monthly subscription.

Swyx3:49

Even though you don't actually have any servers.

Michael Royzen3:50

Even though we don't have any servers. That's right. I was in high school. I wanted to make a little bit of money. I was like, "Yeah." Um-

Swyx3:56

Oh, that's awesome. The, the modern equivalent is kinda Be My Eyes.

Michael Royzen4:00

Right.

Swyx4:00

And, uh, they actually disclosed in the GPT-4 Vision system card recently that the usage was surprisingly like not that frequent. Like y- the extent to which, uh, like all three of us have a sense of sight, I would think that if I lost my sense of sight, I would use Be My Eyes all the time.

The average usage of Be My Eyes per day is 1.5 times.

Michael Royzen4:19

Exactly. And I was, I was thinking about this as well, where, um, I was also look, looking into image captioning, where like you give an, a model an image, and then it tells you what's in the image. But it turns out that what people want is the exact opposite.

People want to give you a description-- Well, people want to give a description of an image and then have the AI generate the image. And so we see-

Swyx4:37

Oh, the other way.

Michael Royzen4:38

Exactly.

Swyx4:38

Okay.

Michael Royzen4:38

And so, you know, at the time, um, there were-- I think there were some GANs, and video was working on this like back in 2019, 2020. They had some like impressive like I think face GANs, where they had, um, this model that would produce these really high-quality portraits.

Um, but it wasn't able to take a natural language description the way Midjourney or DALL-E 3 can and, um, just generate you an image with like exactly what you described in it.

Swyx5:08

Awesome. Uh, and how did that get into NLP?

Building Search5:10

Michael Royzen5:10

I released the SmartLens app, and that was around the time I was a senior in high school. I was applying to college. Um, college rolls around. I'm still sort of working on updating the app in college. Um, but you know, I start thinking like, "Hey, like, what if I make like an enterprise version of this as well?"

Um, at the time, there was Clarifai that provided some computer vision APIs. Um, but I thought, you know, this massive classification model like works so well, and it's so small and so fast, you know, might as well like build an enterprise product.

And I didn't even like talk to users or do any of those things that you're supposed to do. I was just mainly interested in like just like building a type of back end I've never built before. So I was mainly just doing it for myself just to learn.

Um, and so I build this enterprise classification product, and as part of it, um, I'm also building a like invoice processing product, um, where like- Using some of the aspects that I built previously, although obviously it's like very different from classification, um, I wanted to be able to just extract a bunch of structured data from an unstructured invoice, um, through our API.

Um, and that's what led me to Hugging Face for the first time, because that involves some natural language components. Um, and so I go to Hugging Face, and with, um, various e-encoder models, uh, that were around at the time, I think, uh, I used, I used s- the standard BERT and also Longformer, um, which came out around the same time.

Um, and Longformer was interesting because it allowed-- it had a much bigger context window than those models at the time. Like BERT, all of like the f- first gen encoder-only models, they only had a, um, a context window of five hundred twelve tokens.

And it's fixed. There's none of this like Alibi or RoPE that we have now, you know, where we can basically massage it to be longer. It was-- they, they're fixed at five twelve absolute encodings. And so Longformer at the time was the only way that you can fit, say, like a sequence length or ask a question about like four thousand tokens worth of text.

Um, and so, um, implemented Longformer, it worked super well. Um, but like nobody really kind of used the enterprise product. Um, and that's, and that's kind of what I expected, because at the end of the day, um, it was COVID.

I was building this kind of mostly for me, mostly just kind of to learn. Um, and so nobody really used it, and my heart wasn't in it, and like, I kind of just like shelved it. But a little later, I went back to Hugging Face, and I saw this demo that they had, and this is in the summer of twenty-twenty.

They had this demo, um, made by this researcher, Yasin Jernite, and, uh, he called it, um, long-form question answering. Um, and basically, it was this self-contained notebook demo where, um, you can ask a model, uh, a question the way that we do now with ChatGPT.

Um, it would do a lookup into some database, um, and it would give you an answer. And it absolutely blew my mind. Um, the demo itself, it used, I think, BART as the model, and in the notebook, uh, it had support for both, um, an Elasticsearch, um, index of Wikipedia, as well as a dense index, um, powered by Facebook's face-- Vice, I think that's how you pronounce it.

Um, it had both, and it was very iffy, but when it worked, I think the, the question in the demo was, "Why are all boats white?" When it worked, it blew my mind, um, that instead of doing this few-shot thing like people were doing with GPT-3 at the time, which was all the rage, you could just ask a model a question, um, provide no extra context, and it would know what to do and just give you the answer.

Um, it blew my mind to such an extent that I couldn't stop thinking about that. Um, and I started thinking about ways to make it better. Um, I tried, uh, training or f- doing the fine-tune with a, uh, a larger BART model.

Um, a- and this BART model, yeah, it was fine-tuned on this Reddit dataset called, um, Eli Five. So basically-

Swyx9:21

The subreddit.

Michael Royzen9:22

Yeah. The subreddit, yeah. Uh, someone had scraped... I think I, I forget who did it, but someone had scraped a subreddit, um, and put it into like a well-formatted, relatively clean dataset of like human questions and human answers.

So we're bootstrapping this model from, from Eli Five, um, and that made it pretty good at at least getting the right format, um, when doing this RAG retrieval from these databases and then generating the final answer. And so Eli Five actually turned out to be a good dataset, uh, for training these types of question answering models because, um, the question's written by a human, the answer's written by a human, um, and at least helps the model get the format right, even if, you know, the model is still very small and it can't really think super well.

At least it gets the format right. Um, and so it ends up acting as kind of a glorified summarization model, where if it's fed in high-quality context from the retrieval system, it's able to have a reasonably high-quality output.

And so once I made the model as big as I can, just fine-tuning on BART large, um, I started looking for ways to, um, improve the index. So in the demo, in the notebook, uh, it was, um, there were instructions for how to make an Elasticsearch index just for Wikipedia.

And I was like, "Why not do all of Common Crawl?" So I downloaded Common Crawl, and thankfully I had like ten or fifteen thousand dollars worth of AWS credits left over from the SmartLens project. Um, that's what really allowed me to do this, because like there's no other funding.

I was still in college, you know, um, not a lot of money. Um, and so I was able to spin up a bunch of instances and just process all of Common Crawl, which is massive. So it's roughly like, um, it's terabytes of text.

Um, and so I whitelisted... I went to Alexa, um, to get like the top thousand websites or ten thousand websites in the world, and then filtered only by those websites, um, and then indexed those websites. 'Cause the, the webpages were already included in dump, so I just-

Swyx11:29

You mean to supplement Common Crawl or to filter Common Crawl?

Michael Royzen11:32

Filter Common Crawl.

Swyx11:32

Oh, okay.

Michael Royzen11:33

Yeah. So we filtered Common Crawl, filtered Common Crawl just by, uh, yeah, the top I think ten thousand, uh, just to limit this, because obviously like there's this massive long tail of small sites that are really cool actually, and there's, there's other projects like, um, shout out to Marginolion New, um, which is a search engine specialized on the long tail.

I think they actually exclude like the top ten thousand-

Swyx11:56

That's what they do.

Michael Royzen11:57

Yeah.

Swyx11:58

I've seen them around. I just kn- don't really know what their pitch is.

Michael Royzen12:00

Yeah, yeah, yeah. So, so they, they exclude all the top stuff. So the long tail's cool, but for this, that was kind of out of the question, and that was most of the data anyway. So we've- Remove that.

Um, and then I indexed the remaining approximately three hundred and fifty million web pages through, um, Elasticsearch. Um, so I built this index, uh, running on AWS with these web pages. Um, and it actually worked quite well. Like, you can ask it like general common knowledge, history, politics, current events questions, um, and it would be able to do a fast lookup in the index, feed it into the model, um, and, and it would give like a surprisingly good result.

And so when I saw that, I thought that this is definitely doable. And like, it kind of shocked me that like no one else was doing this. And so this was now the fall of twenty twenty. Um, and, um, yeah, I was kind of shocked no one was doing this, but it cost a lot of money to keep it up.

I was still in college. There are other things going on. I got bogged down by classes, and so I ended up shelving this for almost a full year, actually. Um, and I returned to it in fall of twenty twenty-one when, um, Big Science released T0.

Um, when Big Science released the T0 models, that was a massive jump in the reasoning ability of the model, and, um, it was better at reasoning, it was better at summarization. It was still a glorified summarizer, basically.

Swyx13:34

Was this a precursor to Bloom? Because Bloom's the one that I know.

Michael Royzen13:37

I think Bloom ended up, uh, actually coming out in twenty twenty-two, but Bloom had other problems where I think for whatever reason, the Bloom models just were never really that good, which is so sad because I really wanted to use them.

But I think they didn't train on that much data. Um, I think they used like the original-- They were trying to replicate GPT-3, so they just used those numbers, which we now know are like far below Chinchilla Optimal.

And even Chinchilla Optimal, which we can like talk about later, like what we're currently doing with the fine model goes, yeah, it goes way beyond that. Um, but they weren't sharing enough data. I'm not sure how that data was cleaned, but it probably wasn't super clean.

And then they didn't really do any fine-tuning until much later. Um, so T0 worked well because they took the, uh, T5 models, which were, um, closer to Chinchilla Optimal, because I think they were trained on also like three hundred something billion tokens similar to GPT-3, but the models were much smaller.

Um, uh, so the models, yeah, they were pre-trained better, and then they were fine-tuned on, um, this-- I, I think T0 is the first model that did large-scale instruction tuning, um, from diverse data sources in the fall of twenty twenty-one.

Um, this is before InstructGPT. Um, this is before FLAN-T5, which came out in twenty twenty-two. This is, I think, the very, very first, um, at least well-known example of that. Um, and so it came out, and then I did-- on top of T0, I also did the Reddit, um, Eli5 fine-tune.

Um, and that was the first model and system that actually worked well enough to where I didn't get discouraged like I did previously, because the failure cases of like the BART-based system was so egregious. Like, sometimes it would just misinterpret your answers so, or questions so horribly, um, that like it, it was just extremely discouraging.

But for the first time, like it was working reasonably well. Um, al-also using a much bigger model. I think the BART model is like eight hundred million parameters, but T0, we were using 3B. So it was T0 3B, you know, bigger model.

Um, and that was the very first iteration of Hello. Um, so ended up doing a show HN on Hacker News in January twenty twenty-two of that system. Our fine-tuned T0 model connected to our Elasticsearch index of those, um, three hundred and fifty million top ten thousand Common Crawl websites.

Um, and to the best of my knowledge, I think that's the first, um, example that I'm aware of, of a, um, LLM search engine model that's effectively connected to like a large enough index that I would consider like an internet scale.

Um, so, so I think we were the, we were the first, um, to, to release like an internet scale LLM-powered RAG search system, um, in January twenty twenty-two. And around the time, uh, me and my future co-founder, Justin, we were like, you know, we really-- why not do this full-time?

Like, this, this seems like the future. This is really cool. Um, I couldn't really sleep even. Like I would- I was going to bed, and I was like, I was thinking about it. Like I would stay up until like two thirty AM, like reading papers on my phone in bed, go to sleep, wake up the next morning at like eight and just be super excited to keep working.

Um, and I was also doing my thesis at the same time, um, my senior honors thesis at UT Austin, um, about something very similar. Um, we were researching factuality, um, in abstractive question answering systems, um, so a lot of overlap with this project.

Um, and the conclusions of my research actually kind of helped guide the development path of, of Hello in that in the research, we found that LLMs don't, um, they don't know what they don't know. So the conclusion was, is that you always have to do a search, um, to ensure that the model actually knows what it's talking about.

Um, and my favorite example of this even today is kind of with, um, ChatGPT browsing, where you can ask ChatGPT browsing, um, "How do I run llama.cpp?" And ChatGPT browsing will think that llama.cpp is some file on your computer that, you know, you can just compile with GCC and you're all good.

Um, it won't even bother doing a lookup, even though I'm sure somewhere in their internal prompts, you know, they have something like, "If you're not sure, do a lookup." Like, that's not, that's not good enough. So models don't know what they don't know.

You always have to do a search. Um, and so we approached LLM-powered- Question answering from the search angle. Um, we pivoted to make this for programmers in June of 2022, around the time that, um, we were getting into YC.

We realized that, like, what we're really interested in is the case where the models actually have to think. 'Cause up until then, the models were kind of more glorified summarization models. Like, we really thought of them like, um, the Google featured snippets, but on steroids.

And so we, like, we saw a future where the simpler questions would get commoditized. Um, and I still think that's going to happen with, like, Google SGE and... Like, it's-- nowadays, it's really not that hard, um, to, like, answer the more basic kinda like summarization, like current events questions with lightweight models.

And that'll only continue to get cheaper over time. And so we kind of started thinking about this trade-off where LLM models are gonna get both better and cheaper over time. Um, and that's gonna force people who run them to make a choice.

Either you can run a model of the same intelligence that you could previously for cheaper, or you can run a better model for the same price. And so someone like Google, once the price kind of falls low enough, they're going to deploy, and they're already doing this with SGE, they're gonna deploy, um, a relatively basic kind of glorified summarizer model that can answer very basic questions about, like, current events, like who won the Super Bowl, like, you know, what's going on on Capitol Hill, like those types of things.

Um, and the, the flip side of that is, like, more complex questions where, like, you have to reason, and you have to solve problems and, like, debug code. Um, and we realized, like, we were much more interested in kind of, um, going along the bleeding edge of that frontier case.

And so we've optimized everything that we do for that. Um, and, and that's a big reason of why we've built Phind specifically for programmers, as opposed to saying, like, you know, we're kind of a search engine for everyone.

Because, um, as these models get more capable, we're very interested in seeing kind of what the emergent properties are, um, in terms of reasoning, in terms of being able to solve complex multi-step, multi-step problems. Um, and I think that some of those emergent capabilities, like, we're starting to see, but we don't even fully understand.

So as... I think there's always an opportunity for us to become more general if we wanted. Um, but we've, uh, we've been along this path of, like, what is, what is the best, most advanced reasoning engine that's connected to your code base, that's connected to the internet that we can just provide?

Alessio20:58

What is Phind today pragmatically from a product perspective? How do people interact with it?

Product Vision20:58

Michael Royzen21:03

Yeah.

Alessio21:03

Where does it plug into your workflow?

Michael Royzen21:05

Yeah. So Phind is really a system. Um, Phind is a system for programmers when they have a question or when they're frustrated or when something's not working.

Alessio21:14

When they're frustrated. Oh.

Michael Royzen21:15

Yeah, for them to get unblocked. The most abstract pitch for Phind is like, if you're experiencing really any kind of issue as a programmer, we'll solve that issue for you in fifteen seconds as opposed to fifteen minutes or longer.

And so Phind has an interface on the web. Um, it has an interface in VS Code and more IDEs to come. Uh, but ultimately, it's just a system where a developer can paste in a question or paste in code that's not working, um, and Phind will do a search on the internet, or they will find other code in your code base perhaps that's relevant.

Um, Phind will find the context that it needs to answer your question and then feed it to a reasoning engine powerful enough to actually answer it. So that's really the philosophy behind Phind. It's a system for getting developers the answers that they're looking for.

Um, and so right now, from a product perspective, this means that, um, we're really all about getting the right context. So, um, the VS Code extension that we launched recently is a big part of this, uh, 'cause you can just ask a question, um, and it knows where to find the right code context.

Um, in your code, it can do an internet search as well, so it's up to date, um, and it's not just reliant on what the model knows. Um, and it's able to, like, figure out what it needs by itself, um, and answer your question based on that.

Um, and if, you know, it needs some help, you can also, like, yourself kind of just... There's opportunities for you yourself to put in all that context in. Um, and but the issue is also, like, not everyone, um, wants to use VS Code.

Um, some people, like, you know, are real Neovim sticklers, um, or, you know, they're using, like, PyCharm or other IDEs, um, JetBrains. Um, and so for those people, um, they're actually, like, okay with switching tabs, at least for now, if it means them getting their answer.

Um, 'cause really, like, there's been an explosion of all these, like, startups doing code, uh, doing search, et cetera, uh, but really who everyone's competing with is ChatGPT, um, which only has, like, that one web interface. And, like, ChatGPT is really the bar.

Um, and so, and then so that's what we're, what we're, we're up against.

Alessio23:26

And so your idea, you know, we have Aman from Cursor on the podcast, and they've gone through the we need to own the IDE-

Michael Royzen23:33

Yep

Alessio23:33

... thing. Yours is more like in order to get the right answer, people are happy to, like, go somewhere else basically.

Michael Royzen23:40

Yeah.

Alessio23:40

They're happy to get out of their IDE and...

Michael Royzen23:43

That was a great podcast, by the way. Uh, but, but yeah. So, so part of it is that, um, people sometimes perhaps aren't even in, in an IDE. So, like, c- like, the whole task of software engineering goes way beyond just writing code, right?

There's also, like, a design stage. There's a planning stage. A lot of this happens, like, on whiteboards. It happens in notebooks. Um, and so the web product also exists for that, where you're not even coding yet, and you're just trying to get, like, a more conceptual understanding of what you're trying to build first.

Um, but some-- I-- The podcast with Aman was great, but somewhere where I disagree with him is that, um, you actually need to own the IDE. Um, I think- In the long... Sorry, sorry. Uh, let's cut that. Yeah, so I thought the podcast with Aman was great, but somewhere where I disagree with him is that you need to own the IDE.

Um, I think, like, he made kind of some good points about, you know, not having platform risk in the long term, but some of the, you know, features that were mentioned, like, um, suggesting GIFs, for example, um, those are all doable with an extension.

Um, we haven't yet seen, um, with VS Code in particular, um, any functionality that we'd like to do yet in the IDE that we can't either do through directly supported VS Code functionality or something that we kind of hack into there, uh, which we've also done a fair bit of.

Um, and so I think it remains to be seen, um, where that goes. But I think what we're looking to be is, like, we're not trying to just be in an IDE or be an IDE. Like, Phind is a system that goes beyond the IDE and, like, is really meant to cover the entire, um, life cycle of a developer's thought process in going about, like, "Hey," like, "I have this idea, and I want to get from that idea to a working product."

And so then that's what the long-term vision of Phind is really about, is starting with that, where, like, in the future, you know... Uh, in the future, I think programming is just gonna be, um, really just the problem-solving.

Like, you come up with an idea, you come up with, like, the basic design for the algorithm in your head, and you just tell the AI, "Hey," just like, "just do it. Just make it work." Um, and that's what we're building towards.

Swyx26:05

Fantastic. Um, I, I, I think we might want to give people, um, and some impression about, like, the type of traffic that you have. Um, because when you present it with a text box, you could type in anything, and I don't know if you have some mental categorization of, like, what are, like, the top three use cases that people tend to coalesce around.

Michael Royzen26:27

Yeah, that's a great question. Um, so the two main types of searches that we see are how-to questions, like how to do X using Y tool. Um, and this historically has been our bread and butter because, uh, with our embeddings, like, we're really, really good at just going over a bunch of developer documentation and figuring out exactly the part that's relevant and just telling you, "Okay," like, "you can use this method."

Use Cases26:28

Michael Royzen26:53

But as LLMs have gotten better, and as we've really transitioned to, um, using GPT-4 a lot in our product, um, people organically just started pasting in code that's not working and just said, "Fix it for me."

Swyx27:06

Fix this.

Michael Royzen27:06

Yeah. And what really shocks us is that, um, a lot of the people who do that, um, they're coming from ChatGPT. So they tried it in ChatGPT with ChatGPT-4. It didn't work. Uh, maybe it required, like, some multi-step reasoning.

Maybe it required, um, to, like, some internet context or something found in either a Stack Overflow po- Flow post or some documentation to solve it. Um, and so then they paste it into Phind, and then Phind works. Um, so those are really, those two different cases.

Like, how can I build this conceptually or, like, remind me of this one detail that I need to, to build this thing, or just like, "Here's this code, fix it." Um, and so that's what a big part of our VS Code extension is, is, like, enabling a much smoother, "Here, just, like, fix it for me" type of workflow.

That's, that's really its main benefit. It's like it's in your code base. It's in the IDE. It knows how to find the relevant context to answer that question. Um, but at the end of the day, like I said previously, that's still a relatively...

Not to say it's a small part, but it's a limited part of the entire kind of mental life cycle of, of a programmer.

Swyx28:17

Yep. When you launched in-- So you launched in Feb, and then you launched V2 in August. You had a couple other pretty impactful post/feature launches. Um, the web search one was, was massive.

Michael Royzen28:29

Yeah.

Swyx28:30

Um, and you, you were m- so you were mostly a GPT-4 wrapper.

Michael Royzen28:36

We were for a long time.

Swyx28:37

For a long time until recently.

Michael Royzen28:38

Yeah, until recently. Um, and so-

Swyx28:40

So, like, people coming over from ChatGPT were saying, asking, querying the same model-

Michael Royzen28:43

Yep.

Swyx28:44

Uh, w- well, with your version of web search, would that be the primary value proposition?

Michael Royzen28:48

Basically, yeah. And so what we've seen is that any model plus web search is just significantly better than that model itself.

Swyx28:55

Do you think that's what you got right in April? Like, um-

Michael Royzen28:57

Yeah

Swyx28:57

... so you got 1,500 points on Hacker News in, in April, which is un- like, un- if you live on Hacker News a lot-

Michael Royzen29:02

Yeah

Swyx29:02

... that is unheard of for someone so, so early on in your, in your journey.

Michael Royzen29:06

Yeah. W- uh, super, super grateful for that. Definitely was not expecting it.

Swyx29:10

Yeah.

Michael Royzen29:10

So what we've done with Hacker News is we've just kept launching.

Swyx29:13

Yeah.

Michael Royzen29:13

Like, uh, what they don't tell you is, like, you can just keep launching. So, so that's what we, we've been doing. So we launched the very first version of Phind, um, in its current incarnation, um, after, like, the previous demo connected to our own index.

Like, once we got into YC, we scrapped our own index because it was, it was too cumbersome at the time. Um, we moved over to using Bing as kind of just the raw source data, and, um, we launched as Hello Cognition.

And over time, every time we, like, added some intelligence to the product to better model, we just ke- keep launching. And every additional time we launched, h- we got way more traffic. So we actually silently reba- rebranded to Phind-

Swyx29:55

Yeah

Michael Royzen29:56

... in late December of last year. But, like, we didn't have that much traffic. Nobody really knew who we were.

GPT-4 Era30:00

Swyx30:00

How'd you pick the name out of it?

Michael Royzen30:01

Paul Graham actually picked it for us.

Swyx30:03

All right, tell the story.

Michael Royzen30:03

Yeah, so- Oh, boy. Yeah, where do I start? So this is the biggest side. Should I, should we go for, like, the full, the full program story or just-

Swyx30:11

Do you want to do it now or you want to do it later? I'll, I'll give you a choice.

Michael Royzen30:14

Hmm. I think, okay, let's, let's just start with the name for now-

Swyx30:18

Yes

Michael Royzen30:18

... and then we can do the full Paul Graham story later.

Swyx30:20

Okay.

Michael Royzen30:20

Um, but basically, um, Paul Graham, when we were lucky enough to meet him, he saw our name and our, our domain was at the time sayhello.so. And he's just like, "Guys, like, come on. Like, like, like, like what, like what is this?"

You know? Like, and, um, and we were like, "Yeah, but like when we bought it, you know, we were just kind of broke college students." "Like, we didn't have that much money." And, like, we really liked Hello as a name because, um, it was-

Swyx30:46

Starting point, yeah

Michael Royzen30:46

... the first, like, conversational search engine, and that's kind of... that's the angle that we were approaching it from. And so we had sayhello.so, and he's like, "There's so many problems with that. Like- ... like, like the say hello, like what does that even mean?

And like .so, like it's gotta be like a .com." We did some time just, like, with Paul Graham in the room. We just, like, looked at different domain names, like different things that, like, popped into our head. Um, and one of the things that popped into, like Paul Graham said, was Phind.

Like with the P-H-I-N-D spelling in particular.

Swyx31:16

Yeah. Which is not typical naming advice, right?

Michael Royzen31:18

Yes.

Swyx31:18

Like, because it's not... When, when people hear it, they don't spell it that way.

Michael Royzen31:21

Exactly. It's, it's hard to spell, and also it's like very '90s. And so at first-

Swyx31:26

Yeah

Michael Royzen31:27

... like we didn't like... I was like, like, "Uh," like, "I don't know." But over time, like it kind of, it kept, it kept growing on us. And, um, and eventually we're like, "Okay, you know, we like the name."

Um, it's owned by this elderly Canadian gentleman, uh, who, who we got to know, and he was willing to sell it to us. And so we bought it, and, uh-

Swyx31:48

Good buy

Michael Royzen31:48

... and we, and we changed the name. Yeah.

Swyx31:50

All right. Cool.

Michael Royzen31:51

Um, but anyways, where were we? We were-

Swyx31:53

I, I had to ask. I mean, you know-

Michael Royzen31:54

Yeah

Swyx31:54

... every- everyone who looks at you looks, is wondering.

Michael Royzen31:56

A lot of people a- and a lot of people actually pronounce it Phind. Um- ... which, you know, and by now is kinda, you know, uh, it's, it's, it's part of the game. But eventually we wanna buy P-H-I-N-D.com and then just have that redirect to P-H-I-N-D.

Swyx32:11

Yeah.

Michael Royzen32:11

So Phind is like definitely the right spelling.

Swyx32:13

Yeah.

Michael Royzen32:13

And like we'll just, yeah, we'll have all the cases addressed.

Swyx32:16

So Bing web search, uh, and then, and then in August you launched V2. Could you, um... Is, is, is V2 the sy- the Phind as a system pitch, or have you mo- evolved since then?

Michael Royzen32:26

Yeah, so I don't, I, I, like the V2 moniker, like I don't really think of it that way in my mind. There's like, there's the version we launched during, last summer during YC, which was, um, the Bing version directed towards programmers.

Um, and that's kinda like, that's why I call it like the first incarnation of what we currently are, 'cause it was already directed towards programmers. We had like a code snippet search built in as well, 'cause at the time, you know, the models we were using weren't good enough to generate code snippets.

Even GPT, like the Text Davinci 2, which avail- was available at the time, wasn't that good at generating code, and it would generate like very, very short, very incomplete, um, code snippets. And so, um, we launched that last summer, got some traction, but really like we were only doing like, I don't know, maybe like 10,000 searches a day.

Like some people knew about it. Some people used it, which is impressive, 'cause looking back, the product like was not that good. Um, and yeah, every time we've like made an improvement to, um, the way that we retrieve context, uh, through better embeddings, more intelligent like HTML parsers, um, and importantly, like better underlying models.

Um, yeah, I would really consider every kind of iteration after that when we-- Every major version after that was when we introduced a better underlying answering model. Like in February, we launched, um, we... It took, it took a, it's, we had to swallow a bit of our pride when we were like, "Okay, our own models aren't good enough.

We have to go to OpenAI." Um, and that actually, that did lead to kind of like our first like decent bump of traffic, um, in February. Um, and people kept using it. Like our retention was way better too.

Um, but we were still kind of running into problems of like more advanced reasoning. Some people tried it, but people were leaving because even like GPT 3.5, uh, both turbo and non-turbo, like still not that great at doing like code-related reasoning beyond, uh, like the how do you do X, like documentation search type of use case.

Um, and so it was really only when GPT-4 came around in April that we were like, "Okay, like this is like the f- our first real opportunity to really make this thing like the way that it should have been all along."

Um, and having GPT-4 as the, the brain, um, is what led to that Hacker News post. Um, and so what we did was we just let anyone use GPT-4 on Phind for free without a login, um, which I actually don't regret.

So it was very expensive obviously, but like at that stage, all we needed to do was show like, we just needed to like show people, "Here's what Phind can do." That was the main thing, and so that worked.

That worked. Like we got a lot of users. Um, um, do you know Fireship-

Swyx35:26

Yeah

Michael Royzen35:26

... the YouTube channel?

Swyx35:27

The YouTube, Jeff Delaney.

Michael Royzen35:28

Yeah. He made a, uh, a short about Phind.

Swyx35:31

Oh.

Michael Royzen35:31

And that's... And on top of the Hacker News post.

Swyx35:33

Yeah, yeah.

Michael Royzen35:34

And that's what like really, really made it blow up. It got millions of views in days, and, and he, he's just funny. Like what I love-

Swyx35:40

Yeah

Michael Royzen35:40

... about Fireship is like he, like you guys-

Swyx35:42

Short, punchy

Michael Royzen35:43

... yeah, yeah. You, like humor-

Swyx35:44

No, yeah

Michael Royzen35:44

... humor goes a long, a long way- ... um, towards like really grabbing people's attention, and so that blew up.

Swyx35:50

So some- something I would be anxious about as a founder during that period, so obviously we all remember that pretty closely. There were a couple of people who had access to the GPT-4 API doing this, which was unrestricted access to GPT-4, and I have to imagine YC, uh, uh, OpenAI wasn't that happy about that.

Michael Royzen36:08

Um-

Swyx36:09

You know what I mean? 'Cause it was like kind of de facto access to GPT-4 before they released it.

Michael Royzen36:14

ChatGPT-4 was in ChatGPT from day one, I think. Um, OpenAI ac- A- AI actually came to our support because what happened was we had people building unofficial APIs around Phind to-

Swyx36:27

Yeah, basically, right

Michael Royzen36:28

... yeah-

Swyx36:28

This is exactly gonna happen

Michael Royzen36:28

... to try to get free access to it. Um, and- I think OpenAI actually has the right perspective on this, where they're like, "Okay, people can do whatever they want with the API if they're paying for it, like they can do whatever they want, but it's like not okay if, you know, paying customers are being exploited by these other actors."

So they actually got in touch with us and they helped us like, um, set up better Cloudflare bot monitoring controls, um, to effectively like crack down on those unofficial APIs. Um, which, yeah, which, you know, we're very happy about.

Um, but, but yeah, so we, so we launched GPT-4. A lot of people come to the product. Um, and yeah, for a long time, we're just- we're figuring out, like how do we... Like what do we make of this, right?

Like how do we, A, make it better, but also deal with like our costs, which have just like massively, massively ballooned. And I think, um, over time, it's, I think it's become more clear with the release of Llama 2 and Llama 3 on the horizon that we will once again see a return to, um, vertical applications running their own models, as was true last year and, and before.

Um, I think that GPT-4, my hypothesis is that the jump from 4 to 4.5 or 4 to 5 will be smaller than the jump from- 3 to 4 ... 3 to 4. And the reason why is because there were a lot of different things.

Like there was two plus, effectively two, two and a half years of research that went into going from 3 to 4. Like more data, bigger model, all of like the instruction tuning techniques, RLHF. Um, all of that is known.

And like Meta, for example, and now there's all these other startups like Mistral too, like there's a bunch of very well-funded open source players that are now working on just like taking the recipe that's now known and scaling it up.

So I think that even if a delta exists in 2024, the delta between proprietary and open source won't be large enough that a, um, startup like us with a, a lot of data that we've collected can take the data that we have, fine-tune an open source model, and like be able to have it be better than whatever the proprietary model is at the time.

That's, that's my hypothesis. That we'll once again see a return to these verticalized models. Um, and that's something that we're super excited about 'cause, um, yeah, that brings us to kind of the Phind model because, um, the, the plan from kind of the start was to be able to return to that, if that makes sense.

And I think now we're definitely at a point where it does make sense, um, because we have requests from users who like they want longer context in the model, basically. Like they want to be able to, um, ask questions about their entire code base.

Custom Model39:24

Michael Royzen39:24

Um, they want... And w-without, you know, context and retrieval and taking a chance of that, like I think it's generally been shown that if you have the space to just put the raw files inside of a big context window, that is still better than chunking and retrieval.

It just, it just is. So there's various things that we could do with longer context, faster speed, lower cost. Um, super excited about that, and that's the direction that we're going with the Phind model. Um, and our big hyposi- hypothesis there is, is precisely that we, we can take a really good open source model, um, and then just train it on absolutely all of the high-quality data that we can find.

Um, and there's a lot of various, you know, interesting ideas for this. We have our own techniques that we're kind of playing with internally. One of the very interesting ideas that, that I've seen is Octopack from, um, from BigCode.

I don't think that it made that big waves when it came out, I think in August. But the idea is that, um, they have this dataset that maps, um, GitHub commits, um, to a change. So basically, there's all this really high quality, like human-made, human, human-written diff data out there on every time someone makes a commit in some repo.

Um, and you can use that to train models. Uh, you take the file state before and like given a commit message, what should that code look like in the future? Got it. Do you think your human eval is any good?

No, unfortunately. So, so we ran this experiment. We trained the Phind model. Um, and if you go to the BigCode Leaderboard, um, as of today, October, um, 5th, um, all of our models are at the top of the, um, BigCode Leaderboard by far.

It's not close, particularly in languages other than Python. Um, we have a 10-point gap between us and the next best model, uh, on Java, JavaScript, uh, I think C#, um, multilingual. Um, and what we kind of learned from, from that whole experience, um, releasing those models is that human eval doesn't really matter.

Um, not just that, but GPT-4 itself has been trained on human eval. And we know this because GPT-4 is able to predict the exact docstring in many of the problems. Um, I've seen it predict like the specific example values in the docstring, which is extremely improbable for it to just, you know- Yeah ...

know. Um, so I think there's a lot of dataset contamination, and it only captures a very limited subset, um, like what programmers are actually doing. Um, what we do internally for evaluations are, um, we have GPT-4, um, score answers.

Uh, GPT-4 is a really good evaluator. I mean, obviously it's-- By really good, I mean it's the best that we have. I'm sure that, you know, a couple months from now, next year, we'll be like, "Oh, you know, like GPT-4.5, GPT-5, it's so much better.

Like GPT-4 is terrible." But like right now, it's the best that we have short of humans. Um, and what we found is that when doing like temperature zero, um, evals, um, it's actually mostly deterministic, GPT-4, um, across runs, um, in assigning scores to, to different answers.

So- We found it to be a very useful tool in comparing our model to, say, GPT-4. Um, but yeah, on our like internal, like real world, here's what people will be asking this model dataset. Um, and the other thing that we're running is just like releasing the model to our users and just seeing what, what they think.

Uh, because that, that's like the only thing that really matters is like releasing it for the application that it's intended for and then seeing how people react. Um, and for the most part, the incredible thing is, is that people don't notice a difference between our model and GPT-4 for the vast majority of s- searches.

There's some reasoning problems, um, that GPT-4 can still do better. We're working on addressing that. Um, but in terms of like the types of questions that people are asking on Phind, um, yeah, like there's, there's not that much difference.

And in fact, like I've, I've been running my own kinda side-by-side comparisons. Um, shout out to God Mode, by the way. And I've like myself have kind of confirmed this to be the case. And even sometimes it gives a better answer, um, perhaps like more concise or just like better implementation than GPT-4.

Which, that's what surprises me. Um, and, and so we-- like by now we kind of have like this reasoning is all you need kind of hypothesis where we've seen emerging capabilities in the Phind model whereby training it on high-quality code, it can actually like reason better.

Um, it went from not being able to solve, um, like w- world problems where like riddles where like with like temporal, um, and like lo- like placement of objects and moving and stuff like that, uh, that GPT-4 can do pretty well.

We, we went from not being able to do those at all to being able to do them just by training on more code, which is wild. Um, so, so we're already like starting to see like these emerging capabilities.

Swyx44:34

Yeah. So I just wanted to make sure that we have the, I guess like the, the, the, the model card in our heads. So you started from Code Llama.

Michael Royzen44:43

Yes.

Swyx44:43

Uh, 65? 34?

Michael Royzen44:45

34.

Swyx44:46

34.

Michael Royzen44:46

So unfortunately there's no Code Llama 7B.

Swyx44:48

So yeah.

Michael Royzen44:48

If there was, that would be super cool-

Swyx44:50

Right

Michael Royzen44:50

... but there's not.

Swyx44:50

34, and then, uh, which, which by... which in itself was Llama 2, uh, which was on 2 trillion tokens and they added 500 billion code tokens.

Michael Royzen44:58

Yes. And they als-

Swyx44:59

And they just added a bunch more.

Michael Royzen45:00

Yeah. And they al- they d- also did a couple of things. So they did-- I think they did 500 billion like general pre-training, and then they did an extra 20 billion long context pre-training. So they actually, um, ra- increased the like max position, um, tokens to 16K up from 8K.

Um, and then they changed the, um, the theta parameter for the RoPE embeddings as well, um, to give it theoretically better long context support up to 100K tokens. Uh, but yeah, but otherwise it's like basically Llama 2.

Swyx45:31

So you just took, took that and just added data.

Michael Royzen45:33

Exactly. So we-

Swyx45:34

You didn't do any other fundamental-

Michael Royzen45:35

Yeah. So we-

Swyx45:36

... changes

Michael Royzen45:36

... so we didn't actually... We, we haven't yet done anything with the model architecture and we just trained it on like many, many more billions of tokens-

Swyx45:42

Yeah

Michael Royzen45:43

... um, on our own infrastructure. Um, and something else that we're taking a look at now is, um, using reinforcement learning for correctness. Um, one of the interesting pitfalls that we've noticed with the Phind model is that in cases where it gets stuff wrong, sometime is capable of getting the right answer.

It's just there's a big variance problem. It's wildly inconsistent. Um, like there are cases when it is able to get the right chain of thought and able to arrive at the right answer, um, but not always. And so like one of our hypotheses and something that we're gonna try is that like we can, we can actually do reinforcement learning on like for a given problem, generate a bunch of completions, and then like use the correct answer as like a loss basically to try to get it to, um, um, be more correct.

And I think, uh, there's a high chance I think of this working because it's very similar to the like RLHF method where you basically show, um, pairs of completions for a given question, um, except the criteria is like which one is like, you know, um, less harmful.

Uh, but here, you know, we have different criteria, but it-- if the, if the model's already capable of getting the right answer, which it is, which it is, we're just-- we just need to cajole it into being more consistent.

UI Features47:00

Alessio47:00

There were a couple things that I noticed in the product that were not strange but unique.

Michael Royzen47:05

Mm.

Alessio47:05

So first of all, the model can talk multiple times in a row. Like most other applications-

Michael Royzen47:11

Yeah

Alessio47:11

... it's like human model, human model. And then you had outside of the thumbs up, thumbs down, you have things like have the LLM prioritize this message in its answers or then continue from this message to like go back.

How does that change the flow of the user and like in terms of like prompting it, um, yeah, what are like some tricks or learnings to that?

Michael Royzen47:32

Yeah, that's a great question. Um, so, so yeah, that's specifically in our pair programmer, uh, mode, which is a more conversational mode, um, that, um, also like asks you clarifying questions back if it doesn't fully understand what you're doing and it kind of, it, it holds your hand a bit more.

Um, and so from user feedback, we, we had requests to make more of an auto GPT where you can kind of give it this problem that might take multiple searches or multiple different steps, like multiple reasoning steps to solve.

Um, and so that's the, um, that's the, um, impetus behind building that product. Uh, being able to do multiple steps and also be able to handle really long conversations. Like people are really trying to use the pair programmer to go from like sometimes really from like basic idea to like complete working code.

And so what we noticed was, is that we were, we were having like these very, very long threads sometimes with like 60 messages, like 100 messages. And like those become really, really challenging to manage like the appropriate context window of what should go inside of the, um, inside of the context, um, and how to preserve the context so that the model can continue or the product can continue giving good responses Um, even if you're like 60 messages deep in a conversation.

Um, so that's where the, the prioritized user messages are like, comes from is like, uh, we... People have asked us to just like let them pin messages that they, they want to be left in the conversation. Um, and, and yeah, and, and then that seems to have like really gone a long way towards solving that problem.

Swyx49:07

Yeah. And then you have a run in Replit thing.

Michael Royzen49:10

Yes.

Swyx49:10

Are you planning to build your own repl, like learning some people trying to run the wrong code- ... unsafe code?

Michael Royzen49:16

Yes. Yes. So I think like in ter- in, in the long term vision of like being a place where people can go from like idea to like fully working code, having a code sandbox, uh, like a natively integrated code sandbox makes a lot of sense.

Um, and Replit is great, and people use that feature. Um, but, but yeah, I think there's more we can do in terms of like having something, um, a bit closer to code interpreter, where it's able to run the code and then like recursively-

Swyx49:45

Iterate on the code

Michael Royzen49:45

... iterate on it. Exactly.

Swyx49:46

I think Replit is working on, um, APIs to enable you to do that.

Michael Royzen49:49

Yep.

Swyx49:50

So Amjad has specifically told me in person that he's, he wants to enable that for people. At the same time, he's also working on his own models-

Michael Royzen49:56

Right

Swyx49:56

... and Ghostwriter and, you know, all the other stuff.

Michael Royzen49:58

Right. Yeah.

Swyx49:58

So it's gonna get interesting. Like he wants to power you but also compete with you.

Michael Royzen50:03

Yeah. And like, and we love Replit. Um, I think that a lot of these, like a lot of the companies in our space, like we're all going to converge to solving a very similar problem but from a different angle.

So like Replit, um, approaches this problem from the IDE side. Like they started as like this IDE that you can run in the browser. Um, and they started for like from that side making coding just like more accessible.

And we're approaching it from the side of, um, like an LLM that's just like connected to everything that it needs to be connected to, which includes your code context, so that's why like we're kind of making, you know, inroads into IDEs.

Uh, but we're kind of, we're approaching this problem from different sides, and I think it'll be interesting to see where things end up. Um, but I think that, you know, in the long, long term, we have an opportunity to also, um, just have like this general kind of like technical reasoning engine product, um, that's, you know, potentially also not just for, not just for programmers, and it's also powered in this web interface, like where there's, there's potential, I think other, um, things that we will build that eventually might go beyond like our current scope.

PG, Ron & Hardware51:18

Swyx51:18

Exciting. We'll, we'll look forward to that.

Michael Royzen51:20

Thank you.

Swyx51:20

Uh, we're gonna zoom out a little bit into sort of, um, AI ecosystem stories, but, but first we gotta get the Paul Graham- ... Ron Conway story.

Michael Royzen51:29

Yeah. So, uh, flashback to last summer, we're in the YC batch, um, and, uh, we're doing the summer, summer batch, summer '22. So the summer batch runs from June to September, approximately. And so this was late July, early August, um, right around the time that many like YC startups start like going out, like figuring out h- here's how, you know, we're gonna pitch investors and everything.

Um, and at the same time, me and my co-founder, Justin, we were planning on moving to New York. Um, so, um, for a long time, actually, we were thinking about building this company in New York, um, mainly for personal reasons, actually, 'cause like during the pandemic, pre-ChatGPT, pre-last year, pre the AI boom, um, SF unfortunately really kinda, you know, like-

Swyx52:18

So dead

Michael Royzen52:18

... lost its luster. Yeah, like no one was here. Um, it was far from clear, like if there would be an AI boom, if like A- SF would be like the AI-

Swyx52:29

Back.

Michael Royzen52:29

Yeah, exactly. If SF would be so back as everyone is saying these days. Um, it was far from clear. And so, and all of our friends from c- we were graduating college, um, 'cause like we happened to just graduate college and immediately start YC.

Like we didn't even have... I think we had a week in between. Um, so it was just-

Swyx52:47

You didn't bother looking for jobs. You were just like, "This is what we want."

Michael Royzen52:50

Um, well, actually, both me and my co-founder, we had jobs that we secured in 2021 from previous internships, but we both, like we... Funny enough, um, I, when I spoke to, uh, my boss's boss at the company at which like at the, where I reneged my offer, I told him we got into YC.

Uh, they actually said, "Yeah, you should do YC."

Swyx53:10

Wow. That's very selfless. That's great.

Michael Royzen53:13

Yeah. That, that was really great that they did that.

Swyx53:14

In San Francisco, they would have offered to invest as well.

Michael Royzen53:17

Yes.

Swyx53:17

Yeah.

Michael Royzen53:17

Yes, they would have. Uh, but yeah, we were both planning to be in New York. Um, and all of our friends were there from college. And, um, so like at this point, like we have this whole plan. We're like, on August 1st, we're gonna move to New York, and we had like this Airbnb for the month in New York.

We're gonna stay there, and we're gonna work and like all of that. Um, the day before we go to New York, I called Justin, and I, I just, I tell him like, "Why are we doing this?" Like, like, "Why are we doing this?"

'Cause in our batch, like by the time that August 1st rolled around, all of our mentors at YC were saying like, "Hey, like you should really consider staying in SF."

Swyx53:54

It's the hybrid batch, right?

Michael Royzen53:56

Yeah. It was the, it was the, it was the hybrid batch. Um, but like there were already signs that like something was kind of like afoot in SF-

Swyx54:03

Yeah

Michael Royzen54:03

... even if like we didn't fully wanna admit it yet. Um, and so we were like, "No," like, "Uh, I don't know." Um, and so the day bef- but like, I don't know, something kinda clicked when the rubber met the road and it was time to go to New York.

We were like, "Why are we doing this?" And like we didn't have any good reasons for staying in New York at that point beyond like our friends are there. So we still go to New York 'cause like we have the Airbnb.

Like we don't have any other kind of place to go for the next few weeks. We're in New York. Um, and New York is just unfortunately too much fun. Like all of my other friends from college who are just, you know, a- like basically starting their jobs, starting their lives as adults, um, you know- They just got stepped into these jobs, they're making all this money, and they're like partying, and like all these things are happening.

And, and like, yeah, it's just a very distracting place to be. And so we were just like sitting in this like small, you know, like cramped apartment, terrible posture, trying to get as much work done as we can.

Too many distractions. And then we get this email, um, from YC saying that Paul Graham is in town, in SF, um, and he is doing office hours with a certain number of startups in the current batch. Um, and whoever signs up first gets it.

And I happened to be super lucky. I was about to go for a run, but I just, I saw the email notification come across the screen, I immediately clicked on the link. And like immediately, like h- half the spots were gone, but somehow, um, the very last spot was still available.

And so I picked the very, very last time slot at 7:00 PM semi strategically. Um, you know, so we would have like time to go over. Um, and also because like, yeah, I didn't really know how we were gonna get to SF yet.

And so we made a plan that we're gonna fly from New York to SF and back to New York in one day, and do like the full round trip, and we're gonna meet with PG, um, at the YC Mountain View office.

Um, and so we go there, we do that. We meet PG, you know, we, we tell him about the startup. Um, and one thing I love about PG is that he gets like, he gets so excited. Like when he gets excited about something, like you can see his eyes like really light up, and he...

And he'll just start asking you questions. In fact, it's a little challenging sometimes to like finish kinda like the rest of like the description of your pitch, 'cause like he'll just, he'll just like start like, you know, asking all you these, these, all these questions about how it works, and like, you know, what's going on.

Um-

Swyx56:32

And you, like what was the most challenging question that he asked you?

Michael Royzen56:35

Um, I think that like he was asking us a lot of questions about like, like really how it worked. 'Cause like as soon as like we told him like, "Hey, like we think that the future of search is answers, not links," um, like, like we could really see like the gears turning in his head.

Um, I think we were like the first-

Swyx56:55

And you're like-

Michael Royzen56:55

... demo of that that he saw

Swyx56:56

... you had like 10 minutes with him, right?

Michael Royzen56:57

We had a, a, like 45 mostly.

Swyx57:00

Oh, okay.

Michael Royzen57:00

Yeah, it was... Yeah, we had a dec-

Swyx57:01

That's decent

Michael Royzen57:01

... we had a decent chunk of time. Yeah. And so we tell him how it works. Um, like he's very excited about it, and I just like, I just blurt it out, I just like ask him to invest.

And he hasn't even seen the product yet . I just ask him to invest. Um, and he says yeah. Um, and like we're super excited about that. Um-

Swyx57:19

And you're st- like you haven't started your batch?

Michael Royzen57:22

No, no, no. This is like, uh-

Swyx57:23

Oh, this is after your batch

Michael Royzen57:24

... yeah, this is about, uh, halfway through the batch. Or two, two third, no, two thirds of the way through the batch.

Swyx57:29

Which, where, when you're like not technically fundraising yet.

Michael Royzen57:31

Or about to start fundraising.

Swyx57:33

Yeah, yeah, you're about to, yeah.

Michael Royzen57:33

Yeah, so we have like this demo, and like we showed him, and like there was still a lot of issues with the product. Uh, but I think like it like, it must have like still kind of like blown his mind in some way.

Um, and so, yeah, so like we're having fun, he's having fun. Um, we have this dinner planned, um, with this other friend that we had in SF, 'cause we were only there for that one day. So we thought, okay, you know, after an hour we'll be done, you know, we'll grab dinner with our friend, and we'll fly back to New York.

But PG was like, like, "I'm having so much fun," like, "Do you wanna-"

Swyx58:05

Have dinner?

Michael Royzen58:06

"... like, yeah, come to my house?" Or he's like, "I gotta, I gotta go have dinner with my wife-"

Swyx58:11

Uh-huh

Michael Royzen58:12

"... Jessica," um, who's also awesome by the way. Um-

Swyx58:14

She's like the heart of YC.

Michael Royzen58:16

Yeah.

Swyx58:16

Yeah.

Michael Royzen58:16

Yeah, like I... Jessica does not get enough credit, as an aside, for her role in-

Swyx58:21

He, he tries, he tries to-

Michael Royzen58:22

She, she tries

Swyx58:22

... yeah

Michael Royzen58:22

... but like, yeah, Jessica really deserves a lot of credit, 'cause she, like, she understands like the technical side and she understands people, and together they're just like a phenomenal team. But he's like, "Yeah, I gotta go see Jessica.

Um, but you guys are welcome to come with, do you wanna come with?" And we're like, "We have this friend who's like-

Swyx58:40

Blow him off

Michael Royzen58:41

... right now outside of, like literally outside the door- ... who like we also promised to get dinner with, so like we'd love to, but like I don't know if we can." He's like, "Oh, he's welcome to come too."

So like, yeah, so all of us just like hop in his car and we go to his house, and then we just like have this like, um, we have dinner, and we have this like just chat about the future of search.

Like I remember him telling Jessica distinctly, like, "Our kids, and our like kids', our kids' kids are like, are not gonna know what like a search result is. Like they're just gonna like have answers." So that was, that was really like a mind-blowing like inflection point moment, for sure.

Swyx59:21

Wow, that you changed your life.

Michael Royzen59:23

Absolutely.

Swyx59:23

And, and you also just spoiled the, um, the booking system for PG. So now everyone's just gonna go after the last slot.

Michael Royzen59:30

Oh, man. Yeah, but like- ... I don't know if he even does that anymore. What he, apparently-

Swyx59:34

He does, he does. Yeah, I've, I've met other founders that he did it this year.

Michael Royzen59:36

This year, gotcha. But, uh, when I, when we told him about how we did it, he was like, "I am like frankly shocked that like YC just did like a random like scheduling system. They didn't like do anything else."

But, um-

Swyx59:47

Okay, and then he introduces you to Ron Conway-

Michael Royzen59:48

Yes

Swyx59:49

... who is, uh, one of the most legendary angels in Silicon Valley.

Michael Royzen59:53

Yes. So after PG invested, um, the rest of our round came together pretty quickly.

Swyx59:58

Yeah.

Michael Royzen59:58

And so like-

Swyx59:59

I, I, I'm, by the way, I'm surprised, like it's, it might feel like playing favorites, right, within the current batch, to be like, "Yo, PG invested in this one."

Michael Royzen1:00:07

Right. Um, and like

yes. Um-

Swyx1:00:15

Too bad for the others.

Michael Royzen1:00:15

Too bad for the others, I guess. Like, so I think this is a, a bigger point about YC, and like these accelerators in general, is like YC gets like a lot of criticism from founders who feel like they didn't get value out of it.

But like in my view, YC is what you make of it. Like, and YC tells you this. They're like, "You really gotta grab this opportunity, like by the balls, and make the most of it. And if you do, then it could be the best thing in the world.

And if you don't, and if you're just kinda like a passive, even like an average founder in YC, you're still gonna fail." And they s- they, they tell you that. They're like, "If you're average in your batch, you're gonna fail Like you have to just be exceptional in every way.

And so yeah, after PG invested, the rest of our round came together pretty quickly, which I'm very fortunate for. And yeah, he introduced us to Ron. Um, and after he did, I get a call from Ron. Um, and Ron says like: "Hey, like, you know, PG tells me what you're working on.

I'd love to come meet you guys." Um, and we're like: "Wait, no way." Um, and we're just holed up in this like little house in, uh, San Mateo, which is a little small, but you know, it had a nice patio.

In fact, we had like our, our, our monitor set up outside on, on the, uh, deck out there. And so Ron Conway comes over. We go over to the patio where like our workstation is. Um, and Ron Conway, he's known for having like this notebook that he, he goes around with, where he like s-sits down with the notebook and he like takes very, very detailed notes, so he never like forgets anything.

Um, so he sits down with his notebook and he asks us like: "Hey guys, like what do you need?" And we're like: "Well, we need GPUs." Like the GP... Back then, the GPU shortage wasn't even nearly as bad as it is now, but like even then it was still challenging, um, to get like the quota that we needed.

And he's like: "Okay, no problem." Um, and then like he leaves. A couple hours later, um, we get an email, uh, and we're CC'd on an email that Ron wrote to Jensen, the CEO of NVIDIA, um, saying like: "Hey, like these guys need GPUs."

Alessio1:02:07

He didn't say how much? He was just like: "Just give them GPUs"?

Michael Royzen1:02:09

Basically, yeah. Ron is known for writing these like one-liner emails that are like very short but very to the point. And I think that's why like, like everyone responds to Ron. Everyone loves Ron. Um, and so Jensen responds.

He responds quickly, like tagging this VP of AI NVIDIA. Um, and we start work-working with NVIDIA, which is great. Um, and something that, that I love about NVIDIA, by the way, is that after that intro, um, we got matched with like a dedicated team.

And um, at NVIDIA, they know that they're gonna win regardless, so they don't care where you get the GPUs from. They're like, they're truly neutral, unlike various sales reps that you might encounter at, at various like clouds and, you know, hardware companies, et cetera.

Like they actually just wanna help you because they know-- they don't care. Like regardless, they know that if you're getting NVIDIA GPUs, they're still winning. So um, so I guess that's, that's a tip, um, is that like if you're looking for GPUs, like-

Alessio1:03:04

Yeah

Michael Royzen1:03:04

...NVIDIA, yeah, they're-- they'll, they'll help you do it.

Alessio1:03:06

So like, so okay. And, and then just to tie up this thing because it, it... So first of all, that's a fantastic story and like, you know, I just wanted to let you tell that because it's-

Michael Royzen1:03:14

Thank you

Alessio1:03:14

...special. Uh, that, that is a strategic shift, right? You're, uh, that you already decided made by the time you, you met Ron, which is, uh, we are going to have our own hardware. We're gonna rack them in a data center somewhere.

Michael Royzen1:03:25

Well, not even that we, uh, need our own hardware, because actually we don't, um-

Alessio1:03:28

Right

Michael Royzen1:03:28

...in fact say, but we just, we just need GPUs, period.

Alessio1:03:31

Okay.

Michael Royzen1:03:31

And like every cloud loves-- like they have their own sales tactics and like they wanna make you commit-

Alessio1:03:37

Commit

Michael Royzen1:03:37

...to long terms-

Alessio1:03:38

Fees, yeah

Michael Royzen1:03:38

...and like very non-flexible terms and like there's all these-- There's a web of different things that you kind of have to navigate. NVIDIA will kind of be to the point and be like: "Okay, you can get-- do this on this cloud, this on this cloud.

Um, if like this is your budget, maybe you want to consider buying as well." Like they'll, they'll help you walk through what the options are. Um, and in terms of software... And, and, and the reason why they're helpful is because like they, they look at the full picture.

So they, they'll help you with the hardware. And in terms of software, um, they actually implemented a custom feature for us in FasterTransformer, um, which is one of their libraries.

Alessio1:04:14

For you?

Michael Royzen1:04:14

For us, yeah. Uh, which is wild. Yeah, I don't think they would have done it otherwise.

Alessio1:04:18

Uh-huh.

Michael Royzen1:04:18

Um, they implemented streaming generation for T5-based models, which we were running at the time, um, up until we switched to GPT in, um, in February, March of this year. Um, so they, they implemented that just for us actually, um, in FasterTransformer.

Um, and so like they'll help you like look at the complete picture and then just help you, um, get done what you need to get done.

Alessio1:04:41

Cool. And I know one of your interests is also local models, open source models and hardware kind of goes hand in hand. Any fun projects, explorations in that space that you wanna share with local llamas and-

Local Models1:04:41

Michael Royzen1:04:54

Yeah. Um, so it's something that we're very interested in, um, because something that kind of we're hearing a lot about is like people want something like Phind, especially companies, but they wanna have it like within like their own sandbox.

They wanna have it like on hardware that they control. And so I'm super, super interested in how we can get, um, big models to run efficiently on, um, on local hardware. And so like Ollama is great. Llama CPP is great.

Um, very interested in like where the whole quanti-quantization thing is going, because like obviously there are all these like great quantization libraries now, um, that go to four-bit, eight-bit, um, but specifically int eight, um, and int four. Um-

Alessio1:05:44

Which is the lowest it can go, right?

Michael Royzen1:05:46

Right.

Alessio1:05:46

There's no-

Michael Royzen1:05:46

But with int eight, there's not necessarily a speed increase. It's just a storage optimization. Yeah. So we have these great quantization libraries, um, that, you know, for the most part are able to get the size down with not that much quality loss.

But there is some, like the quantized models currently are actually worse than the non-quantized ones. Um, and so I'm very curious if the future is something like what NVIDIA is doing with their implementation of FP8, um, which they're implementing in their Transformer engine library, where basically, um,

once FP8 support is kind of more widespread, um, and hardware can support it efficiently, you can kind of switch between the two different FP8 formats, uh, one with greater precision, one with greater range. Um, and then combine that with only, um, not doing F-FP8 on every layer and doing like a mixed precision with like FP32 on some layers.

Um, and like NVIDIA claims that this strategy, um, that they're kind of demoing with the H100 has no degradation. Um, and so it remains to be seen whether that is really true in practice, but that's something that we're excited about, and whether that can be applied to, like, Macs and other hardware once they get FPA support as well.

Alessio1:07:06

Oh, we should also talk about hiring. Uh, how do you get your info, right? Like you seem to know, um, you are-- you seem self-taught.

Autodidact AI1:07:06

Michael Royzen1:07:13

Yeah. So- ... um, I've always just-- Well, I'm fortunate to have like a decent systems background, um, from UT Austin, um, and somewhat of a research background. Um, even though like I didn't publish any papers, but like I went through all the motions.

Like I didn't publish the, the thesis that I wrote mainly out of time because I was doing both that and the startup at the same time-

Alessio1:07:34

Oh

Michael Royzen1:07:34

... and then I graduated, and then it was YC, and then everything was kind of one after another. Um, but like I'm very fortunate to kinda have like the systems and like a bit of like a research background.

But, um, yeah, for the most part, outside of that foundation, like I've always just whenever I've been interested in something, I just like, I, I go deep.

Alessio1:07:50

Like give, give people t-tips, right? Like-

Michael Royzen1:07:51

Yeah

Alessio1:07:51

... where do you-- what fire hose do you drink from?

Michael Royzen1:07:53

Yeah, exactly. So like whenever I see something that blows my mind the way that that initial Hugging Face demo did, that was like the start of everything, I'll just, yeah, I'll just like I'll start from the beginning. I'll like, if I don't know anything, then like I'll just, I'll start by just trying to get a mental model of what is happening.

Like first I need to understand what so I can understand like the why, the how and the why. And once I can understand that, then I can make my own hypothesis, um, about like, okay, here are the assumptions that the authors of this made, um, and here's why maybe they're correct, maybe they're wrong, and here's how like I can improve on it, um, and iterate on it.

And I guess that's the, the mindset that I, I approach it from is like once I understand something like how can it be better, how can it be faster, how can it be like more accurate? Um, and so I guess for anyone starting now like I would've, I would've used Phind-

Alessio1:08:46

Ah.

Michael Royzen1:08:46

... if, if I was starting now because like I would've loved to just have been able to say like, "Hey, like I have no idea what I'm doing. Can you just like be this like technical research assistant and kinda hold my hand and like ask me clarifying questions and like help me like formalize my assumptions like along the way?"

I would've loved that, but yeah, I just kind of did that myself and just-

Alessio1:09:03

Yeah. Recording moons of yourself using Phind actually would be pretty interesting-

Michael Royzen1:09:07

Yeah

Alessio1:09:07

... to it because I think you, you would use Phind differently than people would by themselves.

Michael Royzen1:09:11

I think so, yeah.

Alessio1:09:12

Unprompted.

Michael Royzen1:09:13

I, I generally use Phind, um, for everything- ... which is definitely, yeah, it's like no, no, even like non-technical questions as well 'cause that's just something I'm, I'm, I'm curious about. Um, but that's generally like n- that's less of a usage pattern nowadays.

Like most people generally for the most part do, um, technical questions on Phind, and that's an, and that is completely understandable because of very deliberate decisions that we've made in how we've optimized the product. Like we've optimized the product very much in a quality-first manner as opposed to a like speed-first or like some balance of the two matters.

So we're like we have to run GPT-4 or some GPT-4 equivalent by default and like, and it has to give like a good answer to like very, a very demanding technical audience or people will leave. So that's just, that's just the trade-off.

So like sometimes it's, it's slower for like simple questions, but like we did, we did that on purpose so.

Alessio1:10:07

Awesome. Before we do a lightning round, um, call for hiring any roles you're looking for, what should people know about working at Phind?

Michael Royzen1:10:16

Yeah. So we really straddle the line between, um, product and research at Phind. Like, um,

for the past little while, a lot of the work that we've done has been solely product. But we also do, especially now with the Phind model, um, a very particular kind of applied research in trying to, um, apply the very latest techniques and techniques that might not, that have not even been proven yet, um, to training the very, very best model for our vertical.

Um, and the two go hand in hand because, um, the product, the UI, the UX is kind of model agnostic. But when it has a better kind of kernel, as Andrej Karpathy put it, plugged into it, um, it gets so much better.

So we're doing really kinda both at the same time. And so someone who like enjoys seeing both of those sides, like doing something very tangible that affects the user, um, high quality, reliable code that runs in production, but also, you know, having that chance to experiment, um, with like building these models, um, yeah, we'd love to talk to you.

Alessio1:11:24

And the title is Applied AI Engineer.

Michael Royzen1:11:26

I don't know what the title is. Like- Like, ah, that, that, that is one title, but, um, I don't know if like this really exists 'cause I feel like we're, I feel like we're too rigid about like bucketing people into categories, you know, like-

Alessio1:11:38

Yeah, founding engineer is fine.

Michael Royzen1:11:39

Yeah. Well, yeah, we, we already have a founding engineer technically, but, uh-

Alessio1:11:42

Well, it's, uh-

Michael Royzen1:11:43

But-

Alessio1:11:43

... for what it's worth- ... uh, OpenAI is adopting Applied AI Engineer.

Michael Royzen1:11:46

Really?

Alessio1:11:47

So it's-

Michael Royzen1:11:47

Hmm

Alessio1:11:47

... becoming a thing.

Michael Royzen1:11:48

It's becoming a thing.

Alessio1:11:51

All right.

Michael Royzen1:11:51

We'll see.

Lightning Round1:11:51

Alessio1:11:52

We'll see. Uh, lightning round.

Michael Royzen1:11:54

Lightning round.

Alessio1:11:55

Yeah, we have three questions: acceleration, exploration, and then a takeaway. So the acceleration one is what's something that already happened in AI that you thought would take much longer?

Michael Royzen1:12:04

Yeah. The jump from these like models being, um, glorified summarization models to actual powerf-powerful reasoning engines happened much faster than we thought, um, 'cause like our product itself transitioned from being kind of, you know, this glorified summarization product to now like mostly a reasoning-heavy product.

And we had no idea that this would happen this fast. Like we thought that like there'd be a lot more time and like a lo- many more things that needed to happen before we could do some level of like

intelligent reasoning on a low level about people's code. Um, but it's already happened, and it happened much faster, um, than we, than we could've thought. Um, but I think that leads into your, uh, to your next point.

Alessio1:12:51

Uh, which is exploration.

Michael Royzen1:12:52

Exploration, yeah.

Alessio1:12:52

What do you think is the most interesting unsolved question in AI?

Michael Royzen1:12:54

Yes. I think- Solving hallucinations, um, determinant, like s- being able to guarantee that the answer will be correct is, I think, super interesting, and it's particularly re-relevant to us 'cause, like, we operate in a space where, like, everything needs to be correct.

Like- ... the code, like, not just the logic, but, like, the implementation, everything has to be completely correct. And there's a lot of very interesting work that's going on in this space. Um, some of it is approaching it from, uh, the angle of, uh, formal grammars.

There's a very interesting paper that came out recently. Um, I forget where it came out of, but it came, it-- The, the paper is basically, um, you can define a grammar, um, that restricts and modifies the model's, uh-

Swyx1:13:39

Log props

Michael Royzen1:13:40

... e-exactly, like decoding strategy to, to only conform to that grammar. Um, and that, that helps it-

Swyx1:13:49

Is this LMQL?

Because I, I feel like-

Michael Royzen1:13:53

I'm not sure

Swyx1:13:53

... LMQL is a little bit too structured for if the goal is avoiding hallucination. That's such a vague goal.

Michael Royzen1:13:59

Yeah.

Swyx1:14:00

I, I, I yeah, I haven't seen anybody-

Michael Royzen1:14:01

This is, yeah, the, uh, this is only something we've begun to take a look at. I haven't fully read the paper yet. Like, I've only kind of skimmed the abstract, but it's something that, like, we're definitely interested in exploring further.

Um, but something that we are, like, a bit further along on is also, like, exploring reinforcement learning, um, for correctness as opposed to only harmfulness the way it has typically been used in this area.

Swyx1:14:21

We just need to see your paper. We just need to see your paper on that.

Michael Royzen1:14:22

Yeah.

Swyx1:14:23

Uh, just a quick follow-up. Do you have e- uh, internal evals for what hallucination rate is on stock GPT-4 and then maybe what yours is after fine-tuning?

Michael Royzen1:14:33

We-- Yeah. So we don't measure hallucination directly in our internal benchmarks. We more measure, like, was the answer right or was it wrong? We, but we measure hallucination indirectly by, um, evaluating the context, like the RAG context fed into the model as well.

Swyx1:14:54

Yeah.

Michael Royzen1:14:54

So basically, if, if the context was bad and the answer was bad-

Swyx1:14:58

Yep

Michael Royzen1:14:58

... then chances are, like-

Swyx1:15:00

It's the context

Michael Royzen1:15:00

... the, it just, it's the context. But if the context was good and the, and it just, like, misinterpreted that or had the wrong conclusion, um, then, like, we can take different steps there. Um-

Swyx1:15:10

Yeah. Harrison from LangChain has been talking about this sort of two-by-two matrix with the RAGAs-

Michael Royzen1:15:14

Right

Swyx1:15:14

... people. Um, it's a pretty simple concept. Like-

Michael Royzen1:15:16

Yes

Swyx1:15:17

... what, what's the source of error?

Michael Royzen1:15:18

Exactly.

Swyx1:15:18

Yeah.

Michael Royzen1:15:18

And I was-- I've been talking to Harrison actually about, like, a more, like, structured way, perhaps within LangChain to, like, do evals, 'cause I think that's a massive problem. Like, every single eval is different, um, for these big large language models.

Um, and doing them in a quantitative way is really hard. Um, but it's possible with, with, like, a platform that I think harnesses GPT-4 in the right way. Um, that and also, um, perhaps a stricter prompting language, like a prompting markup language for prompting models is something I'm also very interested in.

Um, 'cause we've written some very, very complex prompts, particularly for a VS Code extension, um, to, like, do, like, very fancy things with people's code. And, like, I wish there was a way that you could have, like, a more formal way, like a Python for LLM prompting, that you could activate desired things within, like, the model's execution flow through some other abstraction above language that has been, like, tested to do that some of the time, perhaps, like, combined with, like, formal grammar limitations and stuff like that.

Swyx1:16:33

Interesting. I have no idea what that looks like.

Michael Royzen1:16:35

These, these are all things that have kind of emerged directly from the issues we're facing ourselves internally, so-

Swyx1:16:40

Yeah, yeah, yeah.

Michael Royzen1:16:41

But, but yeah, definitely, definitely very abstract still so far.

Swyx1:16:44

Awesome. And yeah, just to wrap, what's, uh, what message, idea you want people to, uh, remember and think about?

Michael Royzen1:16:51

Yeah. I think pay attention to those moments that, like, really jump out at you. Like, like when you see, like, a crazy demo that you can't, like, forget about or, like, like something that you just think is really, really cool.

Um, yeah, don't let that go. 'Cause I see a lot of people trying to start startups from the angle of like, "Hey, I just, like, I just wanna start a startup," or, "I'm just, like, bored at my job," um, or like, "I'm, like, generally interested in the space."

Um, and I personally disagree with that. My take is that, like, it's much easier, having been on both sides of that coin now, it's much easier to stay, like, obsessed every single day when the genesis of your startup is, like, something that really spoke to you in an incredibly meaningful way, um, beyond just being kind of some insight that you've noticed.

Um, and I guess that's-- I think, like, what we're discovering now is that, like, um, in the long, long term, like, what you're really building is, like, you're building a group of people that believe this thing, um, that believe that, like, the future of solving problems and, and making things will be just, like, focused more on, like, the human thought process as opposed to the, the implementation part.

Um, and, and it's like it's that belief that I think is what really gets you through the tough times, um, and hopefully gets you to the other side someday.

Swyx1:18:26

Awesome.

Michael Royzen1:18:27

Yeah.

Swyx1:18:27

I kinda wanna play, uh, "Lose Yourself" as the outro music. Then we'll get DMCA strike. That'd be great though. Thank you so much for coming on.

Michael Royzen1:18:36

Yeah. Thank you so much for having me. This was really fun.