LALatent SpaceApr 24, 2024· 1:02:59

High Agency Pydantic over VC Backed Frameworks — with Jason Liu of Instructor

Jason Liu, creator of the Instructor library, explains why structured outputs from LLMs are best handled by a simple requests-like wrapper rather than a VC-backed framework, arguing that Pydantic-defined schemas via function calling outperform JSON Mode for typed responses. He details his journey from being bearish on LLMs at StitchFix to building Instructor on a bullet train to Japan after GPT-3 proved him wrong. Liu advocates for workflow-based DAGs over reactive agent loops, recommends using rankers rather than cramming 60+ tools into an API call, and credits high agency—trying many experiments and documenting conditions for revisiting failures—as key to his success. He also critiques the MLE hiring hype, urging startups to empower motivated AI engineers instead.

  1. 0:00Introductions
  2. 2:50Stitch Fix AI
  3. 9:39Instructor Origin
  4. 13:17Function Calling
  5. 20:40Evals & Use Cases
  6. 26:41Workflow Design
  7. 31:58AI Stack
  8. 33:40Solo Consulting
  9. 37:08High Agency
  10. 42:59Prompts as Code
  11. 51:06AI Engineers

Powered by PodHood

Transcript

Introductions0:00

Alessio0:00

Hey everyone, welcome to the Latent Space Podcast. This is Alessio, partner and CTO in residence at Decibel Partners, and I'm joined by my co-host Swyx, founder of Small AI.

Swyx0:10

Hello, we're back in the remote studio with Jason Liu from Instructor. Welcome, Jason.

Jason Liu0:14

Hey there. Thanks for having me.

Swyx0:17

Jason, you are extremely famous, so I don't know what I'm gonna do introducing you, but you're one of the-

Jason Liu0:22

Oh my God

Swyx0:22

... Waterloo, uh, clan. Uh, there's like this small cadre of you-

Jason Liu0:25

Yes

Swyx0:26

... that's just completely dominating machine learning. Um, actually, can you list like Waterloo alums that you're like you, you, you know are just dominating and crushing it right now?

Jason Liu0:35

So like John from like, uh, RaisonA is-

Swyx0:40

Uh-huh

Jason Liu0:40

... is doing his like, uh, inversion models, right? I know like, uh-

Swyx0:43

Oh, we should talk to him

Jason Liu0:45

... Clive Chan. Clive Chan from Waterloo. He was like one of the kids where I-- when I started the data science club, he was one of the, the guys who were like joining in and just like hanging out in the room, and now he's like...

Well, he's at Tesla working with like Karpathy, now he's at OpenAI, you know, uh-

Swyx0:58

Yeah. He's in my, uh, climbing club, uh-

Jason Liu1:01

Oh, hell yeah.

Swyx1:02

So yeah.

Jason Liu1:02

Yeah. I was-- I haven't seen him in like six years now.

Swyx1:05

Uh, to, to get in the social scene in San Francisco, you have to climb. Um, so -

Jason Liu1:09

Yeah. For sure. I'm sure

Swyx1:10

... both in career and in rocks.

Jason Liu1:12

Yeah. I mean, a lot of good like problem-solving there, but oh man, I feel like now that you've put me on the spot, I don't know-

Swyx1:18

Yeah, yeah

Jason Liu1:18

... who in Waterloo is doing it, but-

Swyx1:18

It's okay. Uh, there was, there was a, yeah, there was, there was a riff. Okay, but anyway, so, uh, you started the data science club at Waterloo. We can talk about that. Uh, but then also spent five years at Stitch Fix as an MLDE.

Jason Liu1:27

Mm-hmm.

Swyx1:28

Um, you pioneered the use of OpenAI's LLMs to increase stylist efficiency. Um, so you must have been like a very, very early user. This is, this was like pretty early on.

Jason Liu1:38

Yeah. I mean, this was like, I mean, this was like GPT-3. Okay, so we actually were using transformers at Stitch Fix, like before the GPT-3 model. So we were just using transformers for recommendation systems. At that time, I was very skeptical of transformers.

I was like, "Why do we need all this infrastructure? We can just use like matrix factorization." When GPT-2 came out, I fine-tuned my own GPT-2 to write like rap lyrics, and I was like, "Okay, this is cute. Okay, I gotta go back to my real job," right?

Like who cares if I can rap or like write a rap lyric? When GPT, uh, Instruct came out, again I was very much like, "Why are we using like a POST request to review every comment a person leaves?

Like we can just use like classical models." So I was very against language models for like the longest time. And then when ChatGPT came out, I basically just like wrote a long apology letter to like everyone at the company.

I was like, "Hey guys, you know, I was very dismissive of some of this technology. I didn't think it would scale well, and I am wrong. This is incredible." And I immediately just transitioned to go from computer vision recommendation systems to LLMs.

But uh, funny enough, now that we have RAG, we're kind of going back to recommendation systems.

Swyx2:47

Yeah, speaking of that, I think Alessio's gonna bring up the next point.

Alessio2:49

Yeah, I was gonna say, we had Bryan Bischof from X on the podcast.

Stitch Fix AI2:50

Swyx2:52

Mm-hmm.

Alessio2:52

Did you work-- did you overlap at Stitch Fix?

Jason Liu2:54

Yeah, yeah. He u- he was like one of my main users of the, uh, recommendation framework that I had built out at Stitch Fix.

Swyx3:00

Yeah. We talked a lot about RecSys-

Jason Liu3:02

So the system's called-

Swyx3:02

So it makes sense.

Jason Liu3:03

Yeah. Yeah. W- so, uh, I actually like now I have adopted that line that RAG is r- uh, RecSys, and y- you know, if you're trying to reinvent new concepts, you, you should study RecSys first, uh, because you're gonna independently reinvent a lot of concepts.

So your system was called Flight. Yeah.

Swyx3:17

It's a recommendation framework with over eighty percent adoption servicing three hundred and fifty billion requests every day. Um, wasn't there something existing at Stitch Fix? Like why, why did you have to write one from scratch?

Jason Liu3:27

No, so I think because at Stitch Fix a lot of the machine learning engineers and data scientists were writing production code, sort of every team's systems were very bespoke. It's like this c- this team ha- only needs to do like real-time recommendations with small data, so they just have like a FastAPI app with some like pandas code.

This other team has to do a lot more data, so they have some kind of like Spark job that does some batch ETL that does a recommendation, right? And so what happens is each team writes their code differently, and I have to come in and like refactor their code, and I was like, "Oh man, I'm, I'm refactoring four different code bases four different times.

Wouldn't it be better if all of the code quality was my fault?" Right? Like, "Let me just write this framework, force everyone else to use it, and now one person can maintain five different systems rather than five teams having their own bespoke system."

And so it was really a, a need of just sort of standardizing everything. And then once you do that, you can do observability across the entire pipeline and make like large sweeping improvements in this infrastructure, right? If we notice that something is slow, we can detect it on the like operator layer and just say, "Hey, like this team, you guys are doing this operation.

It's lowering our s- latency by like thirty percent. If you just optimize your Python code here- ... we can probably make an extra million dollars. So let's like jump on a call and figure this out." And then a lot of it was just like doing all this observability work to figure out what the heck is going on and optimize this system from not only just a code perspective of- but just like sort of like harassing the org and saying like, "We need to add caching here.

We're doing duplicated work here. Let's go clean up the systems."

Swyx5:04

Yeah. Uh, one more system that I'm interested in finding out more about is your similarity search system using CLIP, uh, and GPT-3 embedding in FAISS. Um, where you-

Jason Liu5:14

Mm-hmm

Swyx5:14

... you said fifty... over fifty million dollars in an-annual revenue. So of course-

Jason Liu5:18

Oh yeah. Beautiful

Swyx5:18

... they all gave all that to you, right? Um-

Jason Liu5:20

No, no, no. I mean, the stock went up and down, but you know, I, I got a little bit, so I'm pretty happy about that. Um, but there, you know, that was when we were doing like fine-tuning like ResNets to do image classification.

And so a lot of it was given an image, if we could predict the different attributes we have in the merchandising, and we can predict like the em- text embeddings of the comments, then we can kind of build a image vector or image embedding that can capture both descriptions of the clothing and sales of the clothing.

And then we would use these additional vectors to augment our recommendation system. And so with this, like the, the recommendation system really was just around like- What are similar items? What are complementary items? What are items that you would wear in a single outfit?

And being able to say on a product page, "Let me show you like 15, 20 more things." And then what we found was like, hey, when you turn that on, you make a bunch of money.

Swyx6:14

Yeah. So okay, so you didn't actually use GPT-3 embeddings, you, you fine-tuned your own?

Jason Liu6:19

Yeah.

Swyx6:19

Uh, because I was surprised that GPT-3 worked off the shelf. Okay, okay.

Jason Liu6:22

Yeah. We, we-- Because I mean, at this point, we would have like, you know, 3 million pieces of inventory over like a billion interactions between users and clothes. Um, any kind of fine-tuning would, would out- definitely outperform, uh, like some off, off-the-shelf model.

Swyx6:38

Cool. Um, um, um, we're about to move on from Stitch Fix, but, uh, you know, any other like fun stories from the Stitch Fix days that you wanna cover?

Jason Liu6:46

No, I think that's basically it. I mean, the biggest one really was the fact that like, I think for just four years, I was so bearish on language models.

Swyx6:53

Mm.

Jason Liu6:53

And just NLP in general, I was just like, "Oh, like none of this really works. Like, why would I spend time focusing on this? I gotta go do the thing that makes money, recommendations, bounding boxes, image classification."

Swyx7:04

Yeah.

Jason Liu7:04

And, uh, you know, now I'm like prompting an image model. It's like, oh man, I was wrong.

Swyx7:10

Um, I think, okay, so, uh, you know, be-- uh, so my Stitch Fix question would be, you know, I think you have a bit of a drip and I don't. You know, my, my primary wardrobe is, uh, free startup, um, uh-

Jason Liu7:20

Oh

Swyx7:20

... conference T-shirts.

Jason Liu7:21

Yeah.

Swyx7:21

Um, should more technology brothers be using Stitch Fix? What's your fashion advice?

Jason Liu7:28

Oh man, I mean, I'm not a user of Stitch Fix, right? It's like I enjoy going out and like touching things and putting things on and trying them on, right? I think Stitch Fix is a place where you kind of go because you want the work offloaded, whereas like I really love the clothing I buy, where I have to like-- When I land in Japan, I'm doing like a 45-minute walk up-

Swyx7:50

Mm

Jason Liu7:50

... a giant hill to find this like weird denim shop. Like, that's the stuff that really excites me. But I think the bigger thing this really captures is this idea that like narrative matters a lot to human beings.

Swyx8:02

Okay.

Jason Liu8:03

And I think the recommendation systems is- that's really hard to capture, right? Like it's easy to sell s- like it's easy to use AI to sell like a $20 shirt, but it's really hard for AI to sell like a $500 shirt.

Swyx8:14

Mm.

Jason Liu8:14

But people are buying $500 shirts, you know what I mean? Like there's, there's definitely something that we can't really capture just yet that we probably will figure out how to in the future, right?

Swyx8:24

Well, it'll probably output in JSON, uh, which is what we're gonna turn to next. Uh, so you, you then you went on a sabbatical to South Park Commons in New York, which is unusual-

Jason Liu8:34

Yeah

Swyx8:34

... because it's usually the other way around.

Jason Liu8:35

Yeah, so basically in 2020, 2020 really, I was like, just like enjoying working a lot, and so I was just like building a lot of stuff. This is where we were making like the, you know, tens of millions of dollars doing stuff.

And then, uh, I had a hand injury, and so I really couldn't code anymore for like a, a year, two years. And so I kinda took sort of half of it as medical leave. The other half, I became more of like a tech lead, just like making sure the systems were like lights were on.

And then when I went to, uh, New York, I spent some time there and kinda just like wound down the tech work, you know, did some pottery, did some jujitsu. And, uh, after GPT came out, I was like, "Oh, like I clearly need to figure out what the, what is going on here 'cause something is feels very magical and I don't understand it."

So I spent basically like five months just prompting and playing around with stuff, and then afterwards it was just my startup friends going like, "Hey Jason, you know, my investors want us to have an AI strategy. Can you help us out?"

And it just kinda it, it just snowballed more and more and became till I was sort of like make, making, making this my full-time job.

Swyx9:39

And, um, you know, you, you had YouTube University and a journaling app, you, you know, a bunch of other, uh, explorations. Uh, but it seems like the, the, the most productive or the most, um, best known thing that came out of your time there was Instructor.

Instructor Origin9:39

Jason Liu9:52

Yeah. Written on the bullet train to Japan.

Swyx9:54

Maybe you should tell us. Uh, yeah. Well, well tell us the origin story.

Jason Liu9:57

Yeah, I mean, I think at, at some point, you know, tools like Guardrails and Marvin came out, right? Those are kind of tools that like use XML and Pydantic to get structured data out. But they really were doing things sort of in the prompt, and these were built with sort of the Instruct models in mind.

And I really, like I'd already done that in the past, right? At Stitch Fix, you know, one of the things we did was we would take a, a Crest note and s- and turn that into a JSON object that we would use to send to, uh, our search engine, right?

So if you said like, "I wanted, you know, skinny jeans that were this size," that would turn into a JSON that we would send to our internal search APIs. But it always felt kinda gross. A lot of it is just like you read the JSON, you like parse it, you make sure the names are strings and ages are numbers, and you do all this like messy stuff.

But when Function Calling came out, it was very much sort of a new way of doing things, right? Function Calling lets you define the schema separate from the data and the instructions. And what this meant was, uh, you can kind of have a lot more complex schemas and just map them in Pydantic, and then you can just keep those very separate.

And then once you add like methods, you can add validators and all that kinda stuff. The one thing I really had with a lot of these libraries, though, was it was doing a lot of the string formatting themselves, which was fine when it was the instruction tune models.

You just have a string. But when you have, uh, these new chat models, you have like these chat messages, and I just didn't really feel like, um, not being able to access that for the developer was sort of a good benefit that they would get.

And so I just said, "Let me write like the most simple SDK around, uh, the OpenAI SDK, or simple wrapper on the SDK, just handle the response model a bit, and kind of think of myself more like a requests than an actual framework that people can use."

And so the goal's like, hey, like this is something that you can use to build your own framework, but let me just do all the boring stuff that nobody really wants to do, right? People want to build their own frameworks, but people don't wanna build like JSON parsing.

Swyx12:00

Uh, and the retrying and all, and all that, uh, other stuff.

Jason Liu12:02

Yeah. Um-

Swyx12:03

Yeah, isn't it... We had this, a little bit of this discussion before the show, but like that design principle of going for being requests rather than being Django, um-

Jason Liu12:11

Yeah

Swyx12:12

... m- t- like what inspires you there? Like w- um, is there a just... Does this come from a lot of prior pain? Are there other

Alessio12:20

Open source projects that kind of inspired your philosophy here?

Jason Liu12:23

Yeah, I mean, I think it would be Requests, right? Like, I think it is just some... It is just the obvious thing you install. Like, if you were gonna go make, like, HTTP requests in Python, you would obviously import Requests.

Maybe if you wanna do more async work, there's, like, future tools, but, like, you don't really even think about installing it. And then when you do install it, you don't think of it as like, "Oh, this is a Requests app."

Right?

Alessio12:46

It's just an app.

Jason Liu12:46

Like, no, this is, this is just Python. Like, the, the bigger question is, like, a lot of people ask questions like, "Oh, why isn't Requests, like, in the standard library?"

Alessio12:54

Yeah.

Jason Liu12:55

Like, that's how I want my library to feel, right? It's like, oh, if you're gonna use the LLM SDKs, you're obviously gonna install Instructor. And then I think the second question would be like, "Oh," like, "how come Instructor doesn't just go into OpenAI, go into Anthropic?"

Like, if that's the conversation we're having, like, that's where I feel like I've succeeded.

Alessio13:11

Yeah.

Jason Liu13:12

It's like-

Alessio13:13

Yeah, yeah

Jason Liu13:13

... so standard, you may as well just have it in the base libraries.

Alessio13:16

Mm-hmm. And the shape of the request has stayed the same, but initially function calling was maybe equal structure outputs for a lot of people.

Function Calling13:17

Jason Liu13:24

Mm.

Alessio13:25

I think n- now the models also support, like, JSON mode and some of these things.

Jason Liu13:29

Mm.

Alessio13:29

And, um, you know, return JSON of, "My grandma's gonna die." All of that stuff is maybe-

Jason Liu13:34

Yeah

Alessio13:35

... maybe to decide. How, how have you seen that evolution? Like, maybe what's the, the meta game today? Like, should people just forget about function calling for structure outputs? Or where, when is structure output, like JSON Mode, the best versus not?

Uh, would love to get any thoughts given that you do this every day.

Jason Liu13:50

Yeah. I would almost say these are, like, different implementations of, like... The, the real thing we care about is the fact that now we have typed responses to language models. And because we have that typed response, my ID is a little bit happier, I get auto-complete.

If I'm using the response wrong, there's a little red squiggly line. Like, those are the things I care about. In terms of whether or not, like, JSON Mode is better, I usually s- think it's, it's almost worse unless you want to spend less money on, like, the prompt tokens that the function call represents.

Um, primarily because with JSON Mode, you don't actually specify the schema. So sure, like, JSON Load works-

Alessio14:25

Mm

Jason Liu14:25

... but really I, I care a lot more than just the fact that it is JSON, right? Uh, I think function calling gives you a tool to specify the fact that like, okay, this is a list of objects that I want, and each object has a name or an age, and I want the age to be above zero, and I wanna make sure it's parsed correctly.

Um, that's where kind of function calling really shines.

Alessio14:43

Mm-hmm. Any thoughts on, um, single versus parallel function calling? Uh, when I first started, so I did a presentation at our AI in Action Discord, um, channel, and obviously showcased Instructor. Uh, one of the big things with, that we had before with single function calling is, like, when you're trying to extract lists, you have to-

Jason Liu15:04

Mm

Alessio15:04

... make these funky, like, properties that are lists to then actually return-

Jason Liu15:07

Yeah

Alessio15:07

... all the objects. Um, how, how do you see the hack being put on the developer's plate versus, like, more of the stuff just getting better in the model? Um, and I know you, you tweeted recently about Anthropic, for example, you know, some lists that are not lists, they're strings, and there's, like, all of these discrepancies.

Jason Liu15:26

Yeah. I almost would prefer it if it was always a single function call, but obviously there is, like, the agents workflows that, you know, Instructor doesn't really support that well, but are things that, you know, ought to be done, right?

Like, you could define I think maybe, like, 50 or 60 different functions in a single API call, and, you know, if it was, like, get the weather or turn the lights on or do something else, it makes a lot of sense to have these parallel function calls.

But in terms of an extraction workflow, I definitely think it's probably more helpful to have everything be a single schema, right? Just because you can sort of specify relationships between these entities, right, that you can't do in function, like, parallel function calling.

You can have a single chain of thought before you generate a list of results. Like, there's, like, small, like, API differences, right, where, yeah, like, if it's for a parallel function calling, if you do one, like, again, really, I really care about the, how the SDK looks, and so it's, okay, do I always return a list of functions or, or do you just wanna have the actual object back out and you wanna have, like, auto-complete over that object?

Alessio16:28

What's kind of the, the cap for, like, how many function definitions you can put in where it still works well? Do you have any, any sense on, on that?

Jason Liu16:36

I mean, for the most part, I haven't really had a need to do anything that's more than, like, six or seven different functions. I think in the documentation they support way more.

Alessio16:44

Mm-hmm.

Jason Liu16:44

But, yeah, I don't even know if there's any good evals that have, like, you know, over, like, two dozen function calls. I think there, like, if, if you're running into issues where you have, like, 20 or 50 or 60 function calls, I think you're much better having those specifications saved in a vector database and then have them be retrieved, right?

So if there are 30 tools, like, you should basically be, like, ranking them and then using the top K to, uh, do selection a little bit better, rather than just, like, shoving, like, 60 functions into a single API.

Yeah.

Alessio17:15

Well, I mean, so I think this is relevant now because previously I think context limits prevented you from having more than, like, you know, a dozen tools anyway, and now that we have million token context windows, um, you know, Claude recently, with their new function calling release, said they can handle two, uh, over 250 tools.

Jason Liu17:34

Yeah.

Alessio17:34

Which is insane to me. That's, uh, that's a lot. Um, yeah, I would say, like, you know, you, you're saying, like, you know, you don't think there's many people doing that. Um, I think anyone with a sort of agent-like platform where you have a bunch of connectors, um, they wouldn't run into that problem.

Probably you're right that they should use a vector database and kind of rag their tools. Um, I know Zapier has, like, a few thousand, like 8,000, 9,000-

Jason Liu17:56

Mm-hmm

Alessio17:56

... connectors that, um, you know, obviously don't fit anywhere. So-

Jason Liu18:00

Yeah

Alessio18:00

... yeah, I mean, that, I think that would be it, unless you need some kind of intelligence that chains things together, which is, is I think what Alessio is coming back to, right? Like, there's this trend about parallel function calling.

I don't know what I think about that. An- Anthropic's version was, um, I think they, they use multiple tools in sequence, but they're not in parallel. I don't, I, I haven't explored this at all. I'm just, like, throwing this open to you as to, like, what do you think about the obvious things?

Jason Liu18:23

Yeah, it's like, you know, do we assume that all function calls could happen in any order? Or like, I think there's a lot of, like- And like, in which case, like, we either can assume that or we can assume that, like, things need to happen in some kind of sequence as a DAG, right?

But if it's a DAG, really that's just, like, one JSON object that is the entire DAG, rather than going like, "Okay, the order of the function that return don't matter." Like, that's just, that's definitely just not true in practice, right?

Like, if I have a thing that's like turn the lights on, like unplug the power, and then like turn the toaster on or something, like the order doesn't matter, right? Um, and it's unclear how well you can describe the importance of that reasoning to a, uh, language model yet.

I mean, I'm sure you can do it with, like, good enough prompting, but I just haven't any... I don't, haven't had any use cases where the function sequence really matters.

Alessio19:09

Yeah. To me, the most interesting thing is the models are better at picking than your ranking-

Jason Liu19:15

Yeah

Alessio19:15

... is usually. Like, I'm, I'm incubating a company around system integration and, for example, with one system there are, like, 780 endpoints, and if you actually try and do vector similarity, it's not that good because the people that wrote the specs didn't have in mind making them, like, semantically apart, you know?

They're kinda like, "Oh, create this, create this, create this," versus when you give it to a model and you put, like in, in Opus, you put them all, it's quite good at picking which ones you should actually run.

Um, and I'm curious to see if the model providers actually care about some of those workflows, or if the agent companies are actually gonna build very good rankers to-

Jason Liu19:52

Yeah

Alessio19:52

... kind of fill that gap.

Jason Liu19:54

Yeah. My money is on the rankers because you can do those so easily, right? You could just say, "Well, given the embeddings of my search query and the embeddings of the description, I can just train XGBoost and just make sure that I have very high, like, MRR," which is, like, mean reciprocal rank.

And so, like, the only objective is to make sure that the tools you use are in the top end, uh, filtered. Like, that feels super straightforward, and you don't have to actually figure out how to fine-tune a language model to do tool selection anymore.

Um, yeah, I definitely think that's the case 'cause I f- for the most part, I imagine you either have, like, less than three tools or more than 1,000.

Alessio20:32

Mm-hmm.

Jason Liu20:32

Like, I don't know how, like, what kind of company said, "Oh, thank God we only have, like, 185 tools and it, this works perfectly," right?

Evals & Use Cases20:40

Alessio20:40

That's right. Um, and before we maybe move on just from this, um, it was interesting to me, you retweeted this thing about entropic function calling, and it was-

Jason Liu20:49

Mm-hmm

Alessio20:50

... uh, Joshua Brown's, uh, retweeting some benchmark that is like, "Oh my God, entropic function calling, so good." And then, uh, you retweet it, and then you tweet it later, and it's like, "It's actually not that good." Uh- ...

what's your flow for, like, uh, how do you actually test these things? Because obviously the benchmarks are lying, right? Because the benchmark-

Jason Liu21:07

Yeah, yeah

Alessio21:08

... says good and you said it's bad, and I trust you more than the benchmark. Uh, how, how do you think about that, and then how do you evolve it over time?

Jason Liu21:15

Yeah, like it's, it's mostly just client data. Like, I think when... Like, I actually have been mostly busy with enough client work that I haven't been able to reproduce public benchmarks, and so I can't even share some of the results to Anthropic.

But I would just say, like, in production, we have some pretty interesting schemas where it's like, you know, iteratively building lists where we're doing, like, updates of lists. Like, we're doing in-place updates, so like upserts and inserts, and in those situations we're like, "Oh yeah, we have a bunch of different parsing errors."

Like, numbers are being returned as strings. We were expecting lists of objects, but we're getting strings that are, like, the strings of JSON, right? So we had to, like, call JSON parse on individual elements. Um, overall, I'm, like, super happy with the, um, Anthropic models compared to the OpenAI models.

Like, Sonnet is very cost effective. Haiku is, in function calling, it's actually better. Um, but I think they just had to sort of like file down the edges a little bit where, like, our tests pass, but then we actually deploy to production, we get, like, you know, half a percent of traffic, you know, having issues where, like, if you ask for JSON, it'll still try to, it'll try to talk to you, or if you use function calling, you know, we'll have, like, a parse error.

And so I think these are things are, de- definitely gonna be things that are, uh, fixed in, like, the upcoming weeks. Um, but in terms of, like, the reasoning capabilities, man, like, it's hard to beat, like, 70%, uh, cost, cost reduction, especially when you're building consumer applications, right?

Like, if you're building something for, like, consultants for private equity, like you're charging $400, it doesn't really matter if it's a dollar or two dollars. But for consumer apps, it, it makes products viable. Like, if you can go from Florida Sonnet, you are, you might actually be able to price it better.

Swyx22:56

Yeah, I, I had this chart about the ELO versus the cost of all the models and, uh-

Jason Liu23:01

Mm-hmm

Swyx23:01

... uh, you know, you could, so you could put trends, trend graphs on each of the, each of those things about, like, you know, higher ELO equals higher cost, except for Haiku. Haiku kind of just broke the lines-

Jason Liu23:11

Yeah

Swyx23:11

... or the ISO ELOs, if you want to kinda call it.

Jason Liu23:15

Yeah.

Swyx23:15

Um, cool. Uh, before we go too far into, like, uh, you know, your opinions on-

Jason Liu23:19

Sure

Swyx23:19

... just the, the overall ecosystem, uh, I wanna make sure that we map out the surface area of Instructor. I would say that most people would be familiar with Instructor from your talks and your tweets and all that.

Uh-

Jason Liu23:29

Mm-hmm

Swyx23:30

... you had a, you had the number one talk at, from the AI Engineer Summit, uh, at Lu-

Jason Liu23:35

Two Jasons. Jason Liu and Jerry Liu.

Swyx23:36

Yeah, yeah, yeah, yeah, yeah. You have to be, uh, named, start with J and then, and then, and then a Liu for, uh, to do well. Um, but yeah, but, uh, until I, until I actually went through your cookbook, um, I didn't realize, like, the, the surface area.

Like, how would you ca- ca- categorize, like, the, the, the use cases, right? You have, like, um, so LLM self-critique, you have knowledge graphs in here, you have, uh, PII da- data sanitation. Um, how do you characterize to people, like, what is the surface area of Instruct, Instructor?

Jason Liu24:02

Yeah. So I mean, this is the part that feels crazy because really the difference is LLMs give you strings and Instructor gives you data structures, and once you get data structures, again, you can do every, like, LeetCode problem you ever thought of, right?

Um, and so I think there's a couple of really common applications. The first one, obviously, is extracting structured data. This is just be, okay, well, like, I wanna put in an image of a receipt, I wanna give back out a list of checkout items with a price and a fee and a coupon code or whatever That's one application.

Another application really is around extracting graphs out. So one of the things we found out about these language models is that not only can you define nodes, it's really good at figuring out what are nodes and what are edges.

And so we have a bunch of examples where, you know, not only do I extract that, you know, this happens after that, but also like, okay, these two are dependencies of another task, and you can do, you know, extracting complex entities that have relationships.

Given a story, for example, you could extract relationships of, like, families across different characters. This is gonna all be done by defining a graph. Um, and then the la- the last really big application really is just around query understanding.

The idea is that, like, any, any API call has some schema, and if you can define that schema ahead of time, you can use a language model to resolve a request into a much more complex request, um, one that an embedding could not do.

So, for example, I have a really popular post called, like, RAG is more than embeddings, and effectively, you know, if I, if I have a question like this, "What was the latest thing that happened this week?" That embeds to nothing.

Right? But really, like, that query should just be, like, select all data where the date time is between today and today minus seven days, right? Um, what if I said, "How did my writing change between this month and last month?"

Again, embeddings would do nothing, right? But really, if you could do like a group by over the month and a summarize, then you could again, like, do something much more interesting. And so this really just calls out the fact that embeddings really is kind of like the lowest hanging fruit, and using something like Instructor can really help produce that data structure, and then you can just use your computer science and reason about this data structure.

Maybe you say, "Okay, well, I'm gonna produce a graph where I wanna group by each month and then summarize them jointly." You can do that if you know how to define this data structure.

Swyx26:24

That-- In that part, you kind of run up against, like, the LangChains of the world that used to have that. Or, I mean, they still, they still do have, like, the self-querying, I think it's, they used to call it on, on, uh, when we had Harrison on in our episode.

Um, how do you, how do you see yourself interacting with the other, I guess, LLM frameworks in the ecosystem?

Workflow Design26:41

Jason Liu26:43

Yeah, I mean, if they use Instructor, I think that's totally cool. I think because it just... Again, it's like, it's just Python, right? It's like, it's like asking like, "Oh, how does, like, Django interact with requests?" Well, you just might make a request.get in a Django app, right?

Um, but no one would say, "Oh, like, I, like, went off of Django 'cause I'm using requests now." Like, these are so-- Like, they should be ideally, like, sort of the wrong comparison. In terms of especially, like, the agent workflows, I think the real goal for me is to go down, like, the LLM compiler route, which is instead of doing like a react type reasoning loop, I think my belief is that we should be using, like, workflows, right?

Swyx27:24

Uh.

Jason Liu27:24

If we do this, then we always have a request and a complete workflow. We can fine-tune a model that has a better workflow. Whereas it's hard to think about, like, how do you fine-tune a better react loop?

Swyx27:33

Yeah.

Jason Liu27:34

Do you wanna f- always train it to have less looping? In which case, like, you wanted to get the right answer the first time, in which case it was a workflow to begin with, right? But then when you-

Swyx27:43

Can, sorry, can you, can you define workflow? Because I think obviously, I w- used to work at a workflow company, but I'm not sure-

Jason Liu27:48

Yeah

Swyx27:48

... this is a well-defined term for everybody.

Jason Liu27:49

Oh, yeah. Like, like, I'm thinking workflow in terms of, like, the Prefect to Zapier workflow. Like, I wanna build a DAG. I want you to tell me what the nodes and edges are, and then maybe, maybe the edges are also, like, put in with AI.

But the idea is that, like, I wanna be able to present you the entire plan and then s- ask you to fix things as I execute it, rather than going like, "Hey, I couldn't parse the JSON, so I'm gonna try again.

I couldn't parse the JSON, like, I'm gonna try again." And then next thing you know, you spent like two dollars on OpenAI credits, right?

Swyx28:19

Yeah.

Jason Liu28:19

Um, whereas with the plan, you can just say, "Oh, the edge between node, like, x and y does not run. Let me just iter- iteratively try to fix that component."

Swyx28:31

Fix the plan.

Jason Liu28:32

"Once that sticks, go on the next component." Right? And obviously, you can get into a world where if you have ex- enough examples of the nodes x and y, maybe you can use like a vector database to find a f- good few shot examples.

You can do a lot if you sort of break down the problem into, like, that workflow and executing that workflow, rather than looping and hoping the reasoning is good enough to, uh, generate the correct output.

Swyx28:55

Yeah. I, I would say, um, you know, I, I've been hammering on Devin a lot. I got, um, access a couple weeks ago, and, um, w- obviously, for a simple task, it does, it does well. For the complicated, like, multi...

like, more than ten, twenty-hour tasks, um, I can see it churning a lot of plans.

Jason Liu29:12

That's a crazy comparison. Like, we used to talk about, like, three, four loops. Like, you're saying like, uh, only once it gets to, like, hour tasks, it's hard?

Swyx29:20

Yeah. Less than an hour-

Jason Liu29:22

That-

Swyx29:22

... it's, it's nothing.

Jason Liu29:24

That's crazy.

Swyx29:26

I mean, I don't know. Yeah. I, okay, maybe, maybe my goalposts have shifted. I don't know. But like-

Jason Liu29:31

That's incredible. Yeah, no. Like, I'm like-

Swyx29:33

Yeah

Jason Liu29:33

... I'm like sub one-minute executions. Like, the fact that you're ta- you're, you're talking about ten hours is incredible.

Swyx29:38

I think it's a spectrum. Um, I, I actually, I don't... I really, really, I, I, I think I'm gonna say this every single time I bring up Devin. Like, let's not reward them for taking longer to do things.

Do you know what I mean? Like, that's, that's a metric that is easily abusable.

Jason Liu29:52

Sure. Yeah.

Swyx29:52

Like, you know what I mean?

Jason Liu29:53

You can game it.

Swyx29:54

Yeah. But all I'm saying is like-

Jason Liu29:55

But I think if you can monotonically increase the success probability over an hour-

Swyx30:01

Mm

Jason Liu30:01

... like, that's winning to me, right? Like, obviously-

Swyx30:03

Correct

Jason Liu30:03

... if you run an hour and it-- you've made no progress, like, I think when we were in, like, AutoGPT land, there was that one example where it's like, "I wanted to, I wanted it to, like, buy me a bicycle."

Swyx30:11

Purchase a network.

Jason Liu30:11

"And overnight, I spent seven dollars on credits, and I never found the bicycle."

Swyx30:15

Yeah, yeah.

Jason Liu30:15

Right?

Swyx30:16

I wonder if he... I wonder if he'll be able to purchase a bicycle. Um, because it actually can do things in the real world, um, it just needs to-

Jason Liu30:22

Yeah

Swyx30:22

... suspend to you for auth and stuff. Um, but, uh, well, the, the point I was trying to make was that-

Jason Liu30:27

Yeah

Swyx30:27

... uh, I can see it churning plans. Like, when it, when it gets on... I, I think one of the agent's, um, loopholes or one of the things that, that is a real barrier for agents is LLMs really like to get stuck into a lane.

And-

Jason Liu30:38

Yeah

Swyx30:39

... and, uh, you know, what you're talking about, what, what, what I've seen, uh, Devin do is it gets stucked in, stuck in a lane, and it will just kinda change plans based on-

Jason Liu30:46

Yeah

Swyx30:46

... the performance of the, the plan itself.

That's kinda cool.

Jason Liu30:51

Uh, yeah, I, I feel like we've gone too much in the looping route, and I think a lot of more plans and, like, DAGs and data structures are probably gonna come back to help fill in some holes.

Swyx31:00

Yeah.

Alessio31:00

What's, like, the interface to that? You know, do you see it's like an existing, like, stately, like, state machine kind of thing that, like, connects to the other lamps, um, the traditional, like, DAG player? So, like, do you think we need something new for, like, uh, AI DAGs?

Jason Liu31:15

Yeah, I mean, I think that the hard part is gonna be describing visually the fact that this DAG can also, like, change over time when it should el- still be allowed to be fuzzy, right? Um, I think in, like, mathematics, we have, like, plate diagrams and, like, Markov chain diagrams and, like, you know, recurrent states and all that.

Like, that-- some of that might come into this, like, workflow world, but to be honest, I, I'm not too sure. I think right now the, the first steps are just how do we take this DAG idea and break it down to modular components that we can, like, prompt better, have few shot examples for, and ultimately, like, fine-tune against?

Um, but in terms of even the UI, it's hard to say what will, uh, likely win. I think, you know, people like Prefect and Zapier have a pretty good shot at, uh, doing a good job.

Swyx31:58

Yeah. Uh, so you s- you seem to use Prefect a lot. I actually worked with a Prefect competitor, Temporal, uh, and I'm also very-

AI Stack31:58

Jason Liu32:03

Oh

Swyx32:03

... familiar with Daxter. Um-

Jason Liu32:05

Mm-hmm.

Swyx32:06

Uh, what else would you call out as, like, particularly interesting in the AI engineering stack?

Jason Liu32:12

Man, I almost use nothing.

Swyx32:15

You just write everything-

Jason Liu32:16

I just use Cursor and, like, pytest.

Swyx32:19

Oh, okay.

Jason Liu32:20

Like, I, I think that's basically it. You know, a lot of the observability companies have-

Swyx32:26

Next, I think you're on Next, right?

Jason Liu32:26

The more observability companies I've tried-

Swyx32:28

Yeah

Jason Liu32:28

... the more I just use Postgres.

Swyx32:31

Really? Okay. Postgres for observability?

Jason Liu32:35

But, like, but the issue really is the fact that these observability companies isn't actually doing observability for the system. It's just doing the LLM thing.

Swyx32:42

Yeah.

Jason Liu32:43

Like, I still end up using, like, Datadog, right? Or, like, you know, Sentry to do, like, latency. And so I just have those systems handle it, and then the, like, prompt in, prompt out latency token costs, I just put that in, like, a Postgres table now.

Swyx32:56

So you don't need, like, 20 funded startups, uh, building LLM ops?

Jason Liu33:02

Yeah, but I'm also, like, a, like, an old tire guy, you know what I mean? Like- ... I think, like, because of my background, it's like, yeah, like, the Python stuff I'll write myself, but, you know, I will also just use Vercel happily.

Swyx33:14

Yeah, yeah.

Jason Liu33:14

Right? Uh, because I'm just not familiar with that world of, uh, you know, tooling. Whereas, like, I think, you know, I spent, like, three good years building observability tools for recommendation systems, and I was like, "Oh, compared to that, like, Instructor is just one f- call."

I just had to put time start, time end, and then count-

Swyx33:32

Mm-hmm

Jason Liu33:32

... the prompt token, right? 'Cause I'm not doing a very complex looping behavior. I'm doing mostly, um, workflows and extraction.

Solo Consulting33:40

Swyx33:40

Yeah. I mean, wh- while we're on this topic, uh, we'll just kinda get this out of the way. Like, uh, you, you famously have choo- decided to not be a v- uh, not be a venture-backed company. You want to do the consulting route.

Like, obv- the obvious-

Jason Liu33:51

Oh, yeah

Swyx33:51

... the obvious route for, you know, someone as successful as Instructor is like, "Oh, here's hosted Instructor with, like, all tooling." Um, and you just said-

Jason Liu33:58

Yeah

Swyx33:58

... you just said you had a whole bunch of experience building observability, tooling. Like, you have the perfect background to do this, and you're not.

Jason Liu34:04

Yeah. Isn't that sick?

Swyx34:06

I think that's sick. I know-- I mean, I know why, because you wanna go free dive. Um, but

Jason Liu34:10

Yeah. Um, yeah, 'cause I think there's two things, right? Like, one, it's like if I tell myself I wanna build Requests, Requests is not a venture-backed startup, right? I mean, o- one could argue, like, whether or not, like, Postman is, but I, I think for the most part it's like, having worked so much, I'm kind of...

Like, I am more interested in looking at how systems are being applied and just being, having access to the most interesting data, and I think I can do that more through a consulting business where I can come in and go, "Oh, you wanna build perfect memory.

You wanna build an agent. You wanna build, like, automations over construction or, like, insurance into supply chain," or, like, "You wanna handle, like, writing, like, private equity, like, mergers and acquisitions reports based off of user interviews." Like, those things are super fun, whereas, like, maintaining the library, I think, is mostly just kind of like a utility that I try to keep up, keep up, especially because if it's not venture backed, I have no reason to sort of go down the route of, like, trying to get 1,000 integrations.

Like, in my mind, I just go like, "Okay. Okay, 98% of the people use OpenAI. I'll support that, and if someone contributes another, like, platform, that's great. I'll merge it in." But, um-

Swyx35:21

Yeah, I mean, you only added Anthropic support, like, this year. Uh-

Jason Liu35:24

Yeah, yeah. The thing, a lot of it was just like, you couldn't even get an API key until, like, this year, right?

Swyx35:29

That's true. That's true.

Jason Liu35:30

And so I was like, okay, if, if I add it, like, last year, I was kinda, I'm trying to, like, double the code base to service, you know, half a percent of all downloads.

Swyx35:38

Do you think the market share will shift a lot now that Anthropic has, like, a, you know, compe- very, very competitive offering?

Jason Liu35:44

I think it's still hard to get, uh, API access. I don't know if it's, it's fully GA now, and if it's GA, if you can-

Swyx35:51

It's GA

Jason Liu35:51

... if you can get, uh, commercial access really easily.

Alessio35:54

I don't know. I, I got commercial after, like, two weeks. I had to reach out to their sales team-

Jason Liu35:58

Okay, yeah, two weeks

Alessio35:59

... and get it, but not too bad, but-

Swyx35:59

Yeah, there's a, there's a call list here, and then, um, anytime you run into rate limits, just, like, ping one of the Anthropic, um, staff members. They're-

Jason Liu36:06

Sweet, sweet. Then we need to, like, cut that part out, so I don't need to, like-

Swyx36:08

No, it's cool. It's cool

Jason Liu36:09

... spread false news. But, um-

Swyx36:10

It's a common question

Jason Liu36:11

... surely just from the price perspective, it's gonna make a lot of sense. Like, if you are a business, you should totally consider, like, Sonnet, right? Like, the cost savings is just gonna justify it if you actually are doing things at volume.

And yeah, I think their SDK is, like, pretty good. Uh, but the, to back to the Instructor thing, I just don't think it's a billion-dollar company, and I think if I raise money, the first question is gonna be, like, "How are you making a billion-dollar company?"

And I would just go like, "Man, like, if I make a million dollars as a consultant, I'm super happy. I'm, like, more than ecstatic." I can have, like, a small staff of, like, three people. Like, it's fun. And I think a lot of my happiest founder friends are those who, like, raised a tiny seed round, became profitable.

They're making, like, 70, 60, 70, like, MRR, uh, 70,000 MRR. And they're like, "We don't even need to raise the seed round. Like, let's just keep it, like, between me and my co-founder. We'll go traveling, and it'll, it'll be a great time."

I think it's a lot of fun. Whereas-

High Agency37:08

Alessio37:08

I repeat to the seed investor in the company. I, I think that's, like, one of the things that people get wrong sometimes, and I see this a lot. Um, they have an insight into, like, some new tech, like say LLM, AI, and they build some open source stuff, and it's like, "I should just raise money and do this."

And I tell people a lot, it's like, look, you, you can make a lot more money doing something else than doing a startup. Like, most people that do a company could make a lot more money just working somewhere else-

Jason Liu37:33

Yeah

Alessio37:33

... than in the company itself. Do you have any advice for folks that are maybe in a similar situation? They're trying to decide, oh, should I stay in my, like, high-paid FAANG job and just tweet this on the side and, and do this on GitHub?

Should I go be a consultant? Like, being a consultant-

Jason Liu37:48

Oh, man

Alessio37:48

... seems like a lot of work. It's like you gotta talk to all these people, you know?

Jason Liu37:53

There's, there's a lot of... There's a lot to unpack, because I think the open source thing is just like, well, I'm just doing it for, like, purely for fun, and I'm doing it because I think I'm right. But part of being right is the fact that it's not a venture-backed startup.

Like, like, I think I'm right because this is all you need, right? Like, you know. Um, so I think a part of it is just, like, part of the philosophy is the fact that this... all you need is a very sharp blade to sort of do your work, and you don't actually need to b- build, like, a big enterprise.

So that's, that's one thing. I think the other thing too that I've kind of been thinking around, just because I have a lot of friends at Google that wanna leave right now, it's like, man, like, what we lack is not money or...

like, money or ti- like skill. Like, what we lack is courage. You should like... You just have to do this-

Alessio38:37

Mm-hmm

Jason Liu38:37

... the hard thing, and you have to do it scared anyways, right? In terms of, like, whether or not you do wanna do a founder, I think that's just a matter of, like, optionality, but I definitely recognize that the, like, expected value of being a founder is still quite low.

Alessio38:53

It is.

Jason Liu38:54

Right. Like, I know, I know as many founder breakups than as I know friends who raised a seed, raised a seed round this year, right? And like, that is, like, the reality and, like, you know. Even, even in...

from that perspective, it's been tough, where it's like, oh, man, like, a lot of incubators want you to have co-founders. Now you spend half the time, like, fundraising and then trying to, like, meet co-founders and find co-founders rather than building the thing.

And I was like, man, like, this is a lot of stuff sp- a lot, a lot of time spent out doing, uh, things I'm not really good at.

Alessio39:29

I, I think-

Jason Liu39:29

Right?

Alessio39:29

I do think there's a rising trend in, uh, solo founding. Um-

Jason Liu39:33

Yeah

Alessio39:33

... you know, I, I am a solo. Um, I think that, uh, something like 30% of, like, I think... I, I forget what the exact stat is. Something like 30% of startups that make it to, like, Series B or something actually are solo founder.

Jason Liu39:43

Mm-hmm.

Alessio39:44

Um, so I think... I feel like this must-have co-founder idea mostly comes from YC, and most everyone else copies it, and then, yeah, you, like, plenty of companies break up over co-founder breakups.

Jason Liu39:56

Yeah. And I bet it would be like, I wonder how much of it is the people who don't have that much, like... And I hope this is not a diss to anybody, but it's like you sort of, you go through the incubator route because you don't have, like, the social equity you, you, you would need to just sort of like send an email to Sequoia and be like, "Hey, I'm, I'm, I'm going on this ride.

Do you want a, do you want a ticket on the rocket ship," right? Like, that's very hard to sell. Like, if I was to raise money, like, that's kind of... Like, my message if I was to raise money is like, "You've seen my Twitter.

My life is sick." "I've decided to make it much worse by being a founder, because this is something I have to do."

Alessio40:29

Yeah.

Jason Liu40:30

"So do you wanna come along? Otherwise, I'm gonna fund it myself." Like, if I can't say that, like, I don't need the money. 'Cause, like, I can, I can, like, handle payroll and, like, hire an intern and get an assistant.

Like, that's all fine. But, like, what I don't wanna do, it's like, I really don't wanna go back to Meta. I gotta... I, I wanna, like, get two years to, like, try to find a problem worth solving. That feels like a bad time.

Alessio40:52

Yeah.

Jason Liu40:53

Right.

Alessio40:53

J- Jason is like, "I wore a YSL jacket on stage at AI Engineer Summit. I don't need your accelerator money." I got enough.

Jason Liu40:59

And boots. You don't forget the boots.

Alessio41:01

That's true. That's true. You have really good boots. Really good boots.

Jason Liu41:04

Um, but I, I think that is a part of it, right? I think it is just, like, optionality, and also just, like, I'm a lot older now. I think 22-year-old Jason would've been probably too scared, and now I'm, like, too wise.

But I think it's a matter of like, oh, if you raise money, you have to have a plan of spending it, and I'm just not that creative with spending that, that much money.

Alessio41:24

Yeah. I mean, to be clear, you just celebrated your 30th birthday. Happy birthday.

Jason Liu41:28

Yeah. It's, um, it's awesome.

Alessio41:29

So, you know-

Jason Liu41:29

Going to Mexico next weekend.

Alessio41:31

Uh, you know, a lot older is relative to some, some of the folks listening.

Staying on the, on the career tips, um, I think Swix had a great post about, are you too old to get into AI? Um, I saw one of your tweets in January '23, you applied to, like, Figma, Notion, Cohere, Anthropic, and all of them-

Jason Liu41:49

Oh, yeah

Alessio41:50

... rejected you because you didn't have enough LLM experience. Uh-

Jason Liu41:53

Yeah

Alessio41:53

... I think at that time it would be easy for a lot of people to say, "Oh, I kinda missed the boat, you know? I'm too late, not gonna make it," you know? Um, any advice for people that feel like that, you know?

Jason Liu42:05

Yeah, I mean, uh, like, the biggest learning here is actually from a lot of folks in jujitsu. They're like, "Oh, man, like, am I... Is it too late to start jujitsu?" Like, "Oh, I'll go to the... I'll join jujitsu once I get in more shape," right?

It's like there's a lot of, like, excuses. And then you say, "Oh, like, why should I start now? I'll be, like, 45 by the time I'm any good." And it's like, well, you'll be 45 anyways. Like-

Alessio42:28

Mm-hmm

Jason Liu42:28

... time is passing. Like, if you don't start now, you start tomorrow, you're just, like, one more day behind. And if you're like, if you're worried about being behind, like, today is, like, the soonest you can start, right?

And so you gotta recognize that, like, maybe you just don't want it, and that's fine too. But, like, if you wanted it, you would've started. Like, like, you know. I think a lot of these people, again, probably think things, think of things on a too, too short time horizon, but again, you know, you're, you're gonna be old anyways, you may as well just start now.

Alessio42:59

You know, one more thing on, I guess, the, um, uh-

Prompts as Code42:59

Swyx43:02

Career advice slash, slash, like, sort of, uh, blogging. Um, you always go viral for this post that you wrote on advice to young people and the lies you tell yourself.

Jason Liu43:10

Oh, yeah, yeah.

Swyx43:11

You said that you were writing it for your sister. What... Like, why is that?

Jason Liu43:14

Yeah. Yeah. She was, like, bummed out about, like, you know, going to college and, like, stressing about jobs, and I was like, "Oh, what I really want to hear. Okay." And I just kinda, like, text-to-speech the whole thing.

It's, it's crazy. It's got, like, fifty thousand views. Like, I'm, I'm mind-blowing.

Swyx43:31

Yeah.

Jason Liu43:31

My career as well.

Swyx43:31

I mean, your, your average tweet has more .

Jason Liu43:35

But that thing is, like, you know, a thirty-minute read now.

Swyx43:39

Yeah, yeah. Um, so there's lots of stuff here which I agree with. I, you know, I'm, I'm also of, occasionally indulge in the sort of life reflection phase. Um, there's the, the how to be lucky. There's the how to, how to have high agency.

Um, I feel like the, the agency thing is always making a, is always a trend in SF or just in tech circles.

Jason Liu43:57

Mm-hmm.

Swyx43:58

Um, how do you define having high agency?

Jason Liu44:00

Yeah, I mean, I'm almost, I'm almost, like, past the high agency phase now. Now, my biggest concern is, like, okay, the agency is just, like, the norm of the vector. What also matters is the direction, right? It's like, how pure is the shot?

Swyx44:14

Yeah, I mean, I think agency is just a matter of, like, having courage and doing the thing that's scary, right? Like, you know, if, if you wanna go rock climbing, it's like, do you decide you wanna go rock climbing, and then you show up to the gym, you rent some shoes, and you just fall forty times?

Or do you go like, "Oh, like, I'm actually more intelligent. Let me go research the kind of shoes that I want. Okay, like, there's, there's flatter shoes and more inclined shoes. Like, which one should I get? Okay, let me go order the shoes on Amazon.

It'll come back in three days. Like, oh, it's a little bit too tight. Maybe it's too aggressive. I'm only a beginner. Let me go ch-" No. I think the higher agent person just, like, goes and, like, falls down twenty times, right?

Jason Liu44:49

Yeah.

Swyx44:49

Um, I think the higher agency person is more focused on, like, process out, uh, metrics versus, um, outcome metrics, right? Like, from pottery, like, one thing I learned was if you wanna be good at pottery, you shouldn't count, like, the number of cups or bowls you make.

You should just weigh the amount of clay you use, right? Like, the successful person says, "Oh, I went through a thousand pounds of clay, a hundred pounds of clay," right? The less agent person's like, "Oh, I made six cups, and then after I made six cups," like, there's not really...

What are you, what do you do next? Like, no, just pounds of clay, pounds of clay. Same with the work here, right? It's like, oh, you just gotta write the tweets, like, make the commits, contribute open source, like, write the documentation.

There's no real outcome. It's just a process, and if you love that process, you just get really good at the thing you're doing.

Jason Liu45:34

Y- yeah.

Swyx45:35

So just to push back on this, 'cause I obviously, I, I, I mostly agree, um, how would you design performance review systems?

Because you're, you're effectively saying we can count lines of code for developers, right? Like, did you put out-

Jason Liu45:49

No, I don't think that would be the actual- Like, I think the... If you make that an outcome, like, I can just expand a for loop, right?

Swyx45:55

Yeah.

Jason Liu45:55

I think-

Swyx45:55

Okay

Jason Liu45:55

... so for, for performance review, this, this is interesting because I've mostly thought of it from the perspective of science and not engineering. Like, I've been running a lot of engineering stand-ups, uh, primarily because there's not really that many machine learning folks.

Like, the process outcome is, like, experiments and ideas, right? Like, if you think about outcomes, what you might wanna think about an outcome is, "Oh, I wanna improve the revenue or whatnot." But that's really hard. But if you're someone who is going out like, "Okay, like, this week, I wanna come up with, like, three or four experiments that might move the needle.

Okay, nothing worked." To them, they might think, "Oh, nothing worked," like, "I suck." But to me, it's like, wow, you've closed off all these other possible avenues for, like, research. Like, you're gonna get to the place that you're gonna figure out that direction really soon, right?

Like, there's no way you try thirty different things and none of them work. Usually, like, you know, ten of them work, five of them work really well, two of them work really, really well, and one thing was, like, you know, a g- the, the nail on the, nail on the head.

So agency lets you sort of capture the volume of experiments, and, like, experience lets you figure out, like, oh, that other half, it's not worth doing, right? Like, I think ex-experience is gonna go, "Half these prompting papers don't make any sense.

Just use chain of thought and just, you know, use a for loop." Um, but that's kinda-- That's b- basically it, right? It's like usually performance for me is around, like, how many experiments are you running? Like, how, how oftentimes are you trying?

Swyx47:18

Yeah.

Jason Liu47:19

Right.

Alessio47:20

When do you give up on an experiment? Because at Stitch Fix, you kinda give up on language models, I guess, in a way, and as, as a tool to use. And then-

Jason Liu47:28

Yeah, I mean, I never really-

Alessio47:28

... maybe the tools got better. Th- they got better-

Jason Liu47:30

Yeah

Alessio47:30

... before, you know, the... You were kinda like, you were right at the time, and then the tool improved. I think there are-

Jason Liu47:35

Yeah

Alessio47:35

... similar, uh, paths in my engineering career where I try one approach, and at the time it doesn't work, and then the thing changes, but then I kinda soured on that approach, and I don't go back to it soon enough.

Um, I-

Jason Liu47:46

I see.

Alessio47:46

Yeah. How do you think about that, that loop?

Jason Liu47:48

So usually, when I, when I, like, h- when I'm coaching folks, and I s- they say, like, "Oh, these things don't work. I'm not gonna pursue them in the future." Like, one of the big things, like, hey, the negative result is a result, and this is something worth documenting.

Like, this isn't academia. Like, if it's negative, you don't just, like, not publish.

Alessio48:01

Mm-hmm.

Jason Liu48:01

Right? But then, like, what do you actually write down? Like, what you should write down is, like, here are the conditions. This is the inputs and the outputs we tried the experiment on. And then one thing that's really valuable is basically writing down w- under what conditions would I revisit these experiments, right?

It's like th- these things don't work because of what we had at the time. If someone is reading this two years from now, under what conditions will we try again? That's really hard, but again, that's, like, another, that's, like, another skill you, you kind of learn, right?

It's like you, you do go back and you do experiments and you figure out why it works now. I think a lot of it here is just, like, scaling worked.

Alessio48:39

Yeah.

Jason Liu48:39

Right? Like, you could actually, like... Rap lyrics, you know, like, that was because I did not have high enough quality data. If we phase shift and say, "Okay, you don't even need training data." It's like, "Oh, great, then it might just work."

Alessio48:51

Mm-hmm. Yeah.

Jason Liu48:52

Different domain.

Alessio48:53

Do, do you have any- anything in your list that is like, it doesn't work now, but I wanna try it again later? Something that people should maybe keep in mind. You know, people always like AGI when, you know?

When are you gonna know the AGI is here? Maybe it's less than that, but, uh, any stuff that you tried recently that didn't work that you think, uh, will get there?

Jason Liu49:11

I mean, I think like the personal assistance and the writing I've shown to myself is just not good enough yet. So I hired a writer and I hired a personal assistant. So now I'm gonna basically like work with these people until I figure out like what I can actually like automate and what are like the reproducible steps.

Swyx49:30

Right.

Jason Liu49:30

But like I think the, the experiment for me is like, I'm gonna go like pay a person like $1,000 a month to like help me improve my life and then let me sort of get them to help me figure out like what are the components and how do I actually modularize something to get it to work.

'Cause it's not just like OAuth, Gmail, Calendar and, and like Notion. It's a little bit more complicated than that, but we just don't know what that is yet. But those are two sort of systems that I wish GPT-4 or Opus was actually good enough to just write me an essay, but most of the essays are still pretty bad.

Swyx49:57

Mm-hmm. Yeah. I, I would say, um, you know, on the personal assistant side, Lindy is probably the one I've seen the most. Um, he was-- uh, Flo was at, a speaker at the summit. I don't know if you've checked it out or, or any other sort of a-agent assistant startup.

Jason Liu50:11

Not recently. I haven't tried Lindy. It was-- they were like behi-- they were not GA-

Swyx50:14

Yeah

Jason Liu50:14

... uh, last time I was considering it.

Swyx50:16

Yeah, yeah. They're not GA yet.

Jason Liu50:17

A lot of it now, it's like, "Oh, like really what I want you to do is like take a look at all of my meetings and like write like a really good weekly summary email for my clients."

Swyx50:26

Yeah.

Jason Liu50:26

"Remind them that I'm like, you know, thinking of them and like working for them." Right? Or it's like, "I want you to notice that like my Mondays were way like, like way too packed and like block out more time.

And, and also like email the people to sche- do the reschedule and then try to opt in to move them around. And then I want you to say, 'Oh, Jason should have like a 15-minute prep break after four back-to-back-"

Swyx50:47

Mm.

Jason Liu50:47

"... meetings sometimes." Those are things that like now I know I can prompt them in, but can it do it well?

Swyx50:53

Yeah.

Jason Liu50:53

Like before I didn't even know that's what I wanted to prompt for, was like defragging a calendar and adding breaks so I can like eat lunch.

Swyx51:01

Right. Yeah, that's the AGI test.

Jason Liu51:04

Yeah, exactly. Compassion, right?

AI Engineers51:06

Alessio51:07

I think one thing that, yeah, we didn't touch on it before, but I think was interesting, uh, you, you had this tweet a while ago about prompts should be code, and then-

Jason Liu51:15

Mm-hmm

Alessio51:15

... there were a lot of companies trying to build prompt engineering tooling, kind of trying to turn the prompt into a more structured thing.

Jason Liu51:23

Mm-hmm.

Alessio51:24

What's your thought today? Like, you know, now you wanna turn the, the thinking into DAGs. Like do prompts should still be code? Like, uh, any, any updated ideas?

Jason Liu51:33

Nah, it's the same thing, right? I think like, you know, with Instructor, it is very much like the output model is defined as a code object. That code object is sent to the LLM and in return you get a data structure.

So the outputs of these models sh- I think should also be code, like code objects, and the inputs somewhat should be code objects. But I think the one thing that Instructor tries to do is separate instruction, data and the, and the types of the output.

Um, and beyond that, I really just think that, you know, most of it should be still like managed pretty closely to the developer. Like so much of it is changing that if you give control of these systems away too early, you end up ultimately wanting them back.

Like many companies I know that when I reach out are ones who are like, "Oh, we're going off of the frameworks because now that we know what the business outcomes we're trying to optimize for-

Swyx52:24

Mm-hmm

Jason Liu52:24

... these frameworks don't work."

Swyx52:26

Yeah.

Jason Liu52:26

'Cause like we do RAG, but we wanna do RAG to like sell you supplements or to have you like schedule the fitness appointment. And like the, the prompts are kinda too baked into the systems to really pull them back out and like start doing upselling or something.

It's really funny, but a lot of it ends up being like, once you understand the business outcomes, uh, you care way more about the prompt, right?

Swyx52:49

Actually, this is fun. So we were try-- uh, in our prep for this call, we were trying to say like, "What can you, what can you as an independent person say that maybe me and Alessio cannot say, or may- you know, someone who works- ...

at a company say?" What do you think is the market share of the frameworks? The LangChain, the LlamaIndex, the everything else.

Jason Liu53:06

Oh, massive. 'Cause not everyone wants to care about the code.

Swyx53:10

Yeah.

Jason Liu53:11

Right? It's like, I think that's a different question to like what is the business model and are they gonna be like massively profitable b-businesses, right? Like making hundreds of millions of dollars, that feels like so straightforward, right?

Swyx53:24

Nice.

Jason Liu53:24

'Cause not everyone is a prompt engineer. Like there's gonna be, there's so much productivity to be captured in like, like back, back office optom-automations.

Swyx53:34

Mm.

Jason Liu53:35

Right? It's not because they care about the prompts, or they care about managing these things. Um-

Swyx53:39

Yeah, but those are the sort of low-code experiences, you know?

Jason Liu53:43

Yeah. I, I think it, it... I think the, the bigger challenge is like, okay, $100 million, probably pretty, pretty easy. It's just time and effort, and they have both like the manpower and the, and the money to sort of solve those problems.

Um, I think just like, again, if you, if you go the VC route, then it's like you're talking about billions, and that's really the goal. That stuff, uh, for me, it's like pretty unclear.

Swyx54:08

Okay.

Jason Liu54:09

But again, that is to say that like I sort of am building things for developers who wanna use Instructor to build their own tooling, but in terms of the amount of developers there are in the world versus like downstream consumers of these things or even just like, you know, like think of how many companies will use like the Adobes and the IBMs, right?

Because they, they want something that's fully managed and they want something that they know will work. And if the incremental 10% requires you to hire another team of 20 people, you might not wanna do it. Um, and I think that kind of organization is really good for, uh, those like bigger companies.

Swyx54:41

And yeah, I just wanna capture your thoughts on one more thing, which is you, you said you wanted, uh, most of the prompts to stay close to the developer. Um-

Jason Liu54:47

Mm-hmm

Swyx54:48

... I would, uh, and Hamel Hussein wrote this like post which I really love called, uh, F- like "F You, Show Me the Prompt".

Jason Liu54:55

Yeah.

Swyx54:55

I think he cites you in, in one of those, uh, part of the blog post. And I think DSPy is kind of like the complete an-antithesis of that, uh, which is I think is interesting 'cause I, I, I also hold the strong view that AI is a better prompt engineer than you are.

Um, and I don't know how to square that. Um-

Jason Liu55:12

I think-

Swyx55:13

Wondering if you have thoughts

Jason Liu55:15

I think so something like DS5 can work because there are like very short-term, uh, metrics to measure success, right? It is like, did you find the PII? Or like, did you write the multi-hop question the correct way? But in these like workflows that I've been managing, a lot of it is like, are we minimizing, like minimizing churn and maximizing retention?

Swyx55:45

Yeah. That's a very long loop.

Jason Liu55:46

Like that's not, like it's not really like a, like Optuna, like training loop, right? Like those things are much more harder to capture. So we don't actually have those metrics for that, right? And obviously we can figure out like, okay, is the summary good?

But then like, how do you mea-measure the quality of the summary, right? It's like that, that feedback loop ends up being a lot longer. And then again, when something changes, it's really hard to make sure that it works across these like newer models or again, like changes to work for, uh, the current prompts.

Like when we migrate from like Anthropic to, to OpenAI, like there's just a, a ton of changes that are like infrastructure related and not necessarily around the prompt itself.

Swyx56:26

Any other AI engineering startups that you think should not exist before we, we wrap up?

Jason Liu56:31

No, I mean-- Oh my gosh. I mean, a lot of it again, is just like every time of investors, like, "What is... How does this make a billion dollars?" Like, it doesn't. I'm gonna go back to just like tweeting and holding my breath underwater.

Yeah, like I don't really pay attention too much to most of these. Like most of the stuff I'm doing is, is around like the consumer layer, right? Like it's not the consumer layer, but like the consumer-

Swyx56:54

Mm-hmm

Jason Liu56:54

... of like LLM calls. I think people just wanna move really fast and they're willing to pick these vendors, but it's like I don't really know if anything has really like blown me out the water. Like I only trust myself, but that's also a function of like, just of being an old man.

Like I think, you know, many companies are definitely very happy with using most of these tools anyways. Um, but I, I definitely think I, I occupy like a very small space in the AI engineering ecosystem.

Swyx57:23

Yeah. Uh, I would say, um, one of the challenges here, you know, you talk, you call about the dealing in the consumer of-

Jason Liu57:30

Yeah

Swyx57:30

... LLMs space. Um, I think that's what AI engineering differs from ML engineering. And I think-

Jason Liu57:35

Mm

Swyx57:35

... a constant disconnect or cognitive dissonance in, in this field, in, in the AI engineers that have sprung up, um, is that they're not as good as the ML engineers. They're not as qualified. Um, I think that, you know, you are someone who has credibility in the MLE space, and you are also, uh, you know, a very, a very authoritative figure in the AIE space.

And, um-

Jason Liu57:58

Authoritative?

Swyx57:59

I think so. And you know I, I think you've built the-

Jason Liu58:01

Oh gosh

Swyx58:02

... the de facto leading library. I think yours, I think Instructor should be part of the standard lib, even though I try to not use it. Like I also try to figure out, um, that I, I, I basically also end up rebuilding Instructor, right?

Like, um, that, that's, that's a lot of the, the back, back and forth that we had over the past, uh, two days. Um, but like, yeah, like, uh, I, I think that's a fundamental thing that we're trying to figure out.

Like there's, there's a very small supply of MLEs. They are-- They're not, not, like not everyone's gonna have that experience that you had. Um, but the, the global demand for AI is going to far outstrip the existing MLEs.

So what do we do? Do we force everyone to go through this, the standard MLE curriculum, or do we make a new one?

Jason Liu58:40

I got some takes.

Swyx58:41

Go.

Jason Liu58:42

I think a lot of these app layer startups should not be hiring MLEs 'cause they end up churning.

Swyx58:49

Chu- Yeah, they wanna work at OpenAI .

Jason Liu58:51

'Cause they're just like, "Hey, guys, I joined, and you have no data, and like all I did this week was like fix some TypeScript build errors and like figure out why we don't have any tests, and like what is this framework X and Y?

Like how come, like what am I, like what are, like how do you measure success? What are your business outcomes? Oh, no? Okay, let's not focus on that. Great, I'll focus on like these TypeScript build errors." And then you're just like, "What am I doing?"

And then you kind of sort of feel really frustrated. And I, I, I already recognize that because I've made offers to machine learning engineers, they've joined, and they've left in like two months. And, and the, the response is like, "Yeah, I think I'm gonna join a research lab."

So I think it's not even that. Like I don't even think you should be hiring these MLEs. On the other hand, what I also see a lot of is the really motivated engineer that's doing more AI engineering is not being allowed to actually like fully pursue the AI engineering.

So they're the guy who built a demo, it got traction, now it's working, but they're still being pulled back to figure out like why Google Calendar integrations are not working, or like how to make sure that like, you know, the button is loading on the page.

And so I, I'm sort of like in a very interesting position where the companies wanna hire an MLE they don't need to hire, but they won't let the excited people who've caught the AI engineering bug to go do that work more full-time.

Um, this is something I'm literally wrestling with like this week, as I just wrote something about it. This is one of the things that I'm probably gonna be recommending in the future, is really thinking about, like, where is the talent coming from?

How much of it is internal? And do you really need to hire someone who's like writing PyTorch code?

Swyx1:00:28

Yeah.

Jason Liu1:00:28

Right?

Swyx1:00:28

Exactly. Uh, you-- Most of the time you're not. You're gonna need someone to write Instructor code .

Jason Liu1:00:34

Right. And you're just like, yeah, you're making this like... And like I feel goofy all the time just like prompting. It's like, "Oh, man, like I wish I just had a target dataset that I could like train a model against."

Swyx1:00:43

Yes.

Jason Liu1:00:43

And I can just say it's right or wrong.

Swyx1:00:45

Yeah.

Jason Liu1:00:46

That's like-

Swyx1:00:46

So, you know, uh, I guess, uh, what Lane Space is, what the AI Engineering World's Fair is, is that we're trying to create and elevate this, this industry of AI engineers, where it's legitimate to actually take these motivated software engineers who wanna build more in AI and do creative things in AI, to actually say, "You have the blessing."

Like, and, and this is leg-legitimate subspecialty of software engineering.

Jason Liu1:01:07

Yeah. Yeah. I think there's gonna be a mix of that, product engineering. Like I think a lot more data science is gonna come in versus machine learning engineering, 'cause a lot of it now is just quantifying, like what does the business actually want as an outcome, right?

The outcome is not RAG app.

Swyx1:01:23

Yeah.

Jason Liu1:01:24

The outcome is like reduced churn or something like that, but we need to figure out what that actually is and how to measure it.

Swyx1:01:29

Yeah. Yeah. All the, the data engineering tools, uh, still apply, uh, BI layers, semantic layers, whatever.

Jason Liu1:01:36

Yeah.

Swyx1:01:37

Cool. Um-

Jason Liu1:01:37

We'll see.

Swyx1:01:38

We'll, we'll have you back again for the World's Fair. Um, we kn- we don't know what we're, what you're gonna talk about, uh, but I'm sure it's gonna be amazing. Um, you're a very-

Jason Liu1:01:47

The title I've written, it's just, uh, "Pydantic Is Still All You Need".

Swyx1:01:53

I, I'm worried about having too many all you need titles because that's obviously very trendy. Um, so, so yeah, you have one of them, but I n- I need to keep a lid on like, you know, everyone saying their thing is all you need.

Um, but yeah, we'll figure it out .

Jason Liu1:02:05

Pydantic is not my thing. It's someone else's thing.

Swyx1:02:06

Yeah, yeah, yeah.

Jason Liu1:02:07

I think that's why it works.

Swyx1:02:08

Yeah. It's true. Um, cool. Well, uh, it was a real pleasure to have you on, uh, awesome. Uh-

Jason Liu1:02:13

Of course

Swyx1:02:13

... everyone, everyone should go follow you on, on Twitter and, uh, check out Instructor.

Jason Liu1:02:17

Oh my goodness.

Swyx1:02:17

There's also Instructor JS, I think, which, um, which I'm very happy to see. And-

Jason Liu1:02:21

Yeah, yeah

Swyx1:02:21

... what else? Any other-

Jason Liu1:02:22

Use instructor.com.

Swyx1:02:24

Yeah. A- a- anything else to plug?

Jason Liu1:02:26

Use instructor.com. We got a domain name now.

Swyx1:02:28

Nice. Nice. Awesome. Cool.

Jason Liu1:02:31

Cool. Thanks, Tristan. Thanks.