LALatent SpaceMar 13, 2025· 27:34

[Lightning Pod] Evals: How to Improve AI Consistently — with Hamel Husain and Shreya Shankar

Hamel Husain and Shreya Shankar join the Latent Space podcast to argue that systematic evaluation (evals) is the critical missing piece for moving AI applications from demo to production. They estimate 75% of evals in the wild use LLM-as-judge, but 80% of those are not helpful without proper validation against domain experts. Husain advocates building custom annotation UIs (using tools like Lovable or Cursor) to speed error analysis, while Shankar emphasizes synthetic data generation that balances real-world data and LLM outputs, drawing on social science methods. They preview their free lightning lesson on March 21 and a four-week paid course covering the eval lifecycle: synthetic data creation, LLM-as-judge calibration, error analysis, and iterative improvement, with hands-on coding assignments. The episode also touches on trends like dedicated judge models (e.g., Haizelabs' Verdict) and the value of basic data literacy (e.g., pivot tables) for eval participation.

  1. 0:00Introduction
  2. 1:05The Course
  3. 7:06Course Evolution
  4. 12:31LLM as Judge
  5. 17:26Evals Trends
  6. 20:52Tools & Culture
  7. 25:23Lightning Lesson

Powered by PodHood

Transcript

Introduction0:00

Hamel Husain0:03

Hey everyone, welcome back to another Latent Space Lightning Pod. This is Alessio, partner and CTO at Decibel, and I'm joined by my co-host Swyx, founder of Small AI.

Swyx0:13

Hey, and today we have the epic return of Shreya Shankar. I think you were guest number, like, 10 or 12 on the pod, uh, but welcome back.

Shreya Shankar0:23

Yeah. I was early, early adopter of your podcast. Early-

Swyx0:27

Early adopter

Shreya Shankar0:27

...something.

Swyx0:29

Uh, I remember we were talking a lot about grounder research at the time, and then, you know, since then I've had the opportunity to visit you on campus and that was-

Shreya Shankar0:38

Sure

Swyx0:38

...that was one of my most, uh, favorite, uh, trips, I think, uh, last year. Um, and also-

Shreya Shankar0:45

Oh, I'm happy

Swyx0:45

...yeah, um, I mean, I learned about Doc ETL, but also Lotus, and then I was like, "Oh, like, this is actually, like, a, a, a very promising area of research." Anyway, we also have Hamel Husain. Uh, I don't know if you've actually been on the pod, but I feel like we, we did the, the Rasa Hub pod together.

Anyway, but welcome.

Hamel Husain1:03

Yeah. Glad to be here.

The Course1:05

Swyx1:05

Uh, you are both doing a course on evals. Tell us about it.

Hamel Husain1:09

Yeah, so I can go first. So, you know, the thing that we keep talking about and writing about all the time is evals. And it's not because we're somehow obsessed with evals or there's some kind of, you know, fairy that gives me money every time I talk about evals under my pillow.

It's more, okay, any time that I h- try to help someone build an AI application, uh, they always get stuck on how to move beyond a demo product, and they, they get stuck, like, how to systematically improve things and measure it, and that's a very central part to improving AI products.

And so, and so I've learned over the last few years that people really get stuck on this systematic measurement and evaluation of AI, which is, like, really important. And so, you know, I found that there's not that much material out there on how to do it.

Um, it can be a little bit mysterious for people. And so, you know, what's motivating me specifically is, okay, can, you know, we create materials that are accessible to everybody, um, you know, in, even beyond, let's say, writing about it and giving examples, like, can we do lots of hands-on exercises and things like that, 'cause sometimes people need more.

Um, I'll let Shreya talk as well or you can ask me questions. Oh, I think you might be muted.

Swyx2:40

Uh, yeah, I was just pulling up your evals chart because I, I, you know, my role here is to kind of illustrate while we talk. But, uh, everyone should have th- this in, this in mind. I think I have a slightly different formulation of this, but, uh, I think we all agree that evals are important for going to production.

Yeah, Shreya, go ahead.

Shreya Shankar2:58

Yeah. Echoing everything, and of course you, Swyx, said evals are important, but there's two motivations for me. One is that there's a bunch of people who are now AI engineers that didn't have the ML engineer and data engineering persona, but still need to have data literacy.

So how can we kind of help educate this new group of people so they don't get lost beyond the demo? The other thing is I, I think that AI engineering evaluation or evals here is actually different from MLOps or ML evaluation for traditional ML models.

We were in a much more, you know, data-rich setting in MLOps, so we were taught to come up with loss metrics or even just better metrics to quantify performance. But today we're in this more data-scarce setting and then we have to kind of train ourselves how do we look at data?

How do we quantify our vibes? How do we systematically go about doing this? And, and it's not all hand-wavy. Like, there's a lot of tools that we can draw on, for example, from the social science literature, from qualitative research on how to make sense of, you know, unstructured text data that comes out of the LLM.

So I'm really excited to be able to put together this course with Hamel to teach the stuff that people just don't know about.

Swyx4:13

Yeah. Um, so what, uh, I guess the... I think it's a really interesting point that the sort of type of evals that we used to do are maybe not the same kinds of evals. And I feel like I, we need a terminology or a taxonomy of, like, these kinds of evals.

We had in the New York conference, the Datadog guy, I, I forget his... Diamond Bishop, I think he presented, uh, some eval- some ways in which of he broke down evals, which I thought was pretty cool. Hamel, I think your talk didn't really touch on...

I think you touched on evals a little bit, but yours were more about pr- sort of product strategy, right?

Hamel Husain4:48

Yeah. I didn't... I tried to venture a little bit outside-

Swyx4:50

Yeah, yeah

Hamel Husain4:51

...the, the what I always talk about.

Swyx4:54

Yeah. It was a very brave talk. Uh, I was... I would definitely say, like, some people on YouTube didn't get the, like, you have to invert everything. They failed the-

Hamel Husain5:00

Yeah, it was a little risky. Yeah, for sure

Swyx5:02

...they failed the reasoning test. But, but yeah, like, so what, what is the syllabus?

Hamel Husain5:06

Yeah. It might be good to pull up the landing page-

Swyx5:08

Right

Hamel Husain5:08

...and we can talk through it. So let me get the link. Uh, and I can put it in the chat.

Swyx5:13

No, I have it. Right?

Hamel Husain5:14

You have it? Okay.

Swyx5:15

This thing.

Shreya Shankar5:15

That's not the syllabus. Okay.

Hamel Husain5:17

Oh, yeah. That's for the lightning lesson. Uh, maybe we can do, like, we can talk about lightning lesson too. We can talk about the course-

Swyx5:24

Okay

Hamel Husain5:24

...uh, either way. Um-

Swyx5:25

I mean, the course is here, and then the sy- I think the syllabus-

Shreya Shankar5:28

Yeah, there's a brief syllabus there.

Swyx5:30

Okay.

Hamel Husain5:32

Yeah.

Swyx5:32

Yeah, here.

Hamel Husain5:32

It's at the... It's a little bit if you scroll down a little bit.

Swyx5:35

There we go.

Hamel Husain5:35

No, scroll up maybe.

Yeah. No. Okay. Yeah, you're right. Scroll down. All right .

Swyx5:42

Sorry. Sorry for leading you astray.

Hamel Husain5:46

Yeah.

Shreya, you want to talk about this?

Shreya Shankar5:50

Sure.

Hamel Husain5:51

Um-

Shreya Shankar5:51

I, I... Okay, it's, it's all in flux, but the high level idea is to, to break down this evals life cycle from first how do you even get started? What does it mean to- Come up with evaluations for something that you've never reasoned about evaluating.

Um, so part of that is creating synthetic data, part of that is creating good evals, and kind of we slowly venture off into more complicated topics like LLM as a judge, um, error analysis, how do you kind of create this s- cycle of improving your AI application.

Um, and then throughout this whole thing, we will have at least one... or sorry, at least two, um, assignments, kind of like coding projects that you work on, um, to kind of like s- implement a lot of these lessons.

And we're hoping to add more, but, you know, we're... Yeah, it's gonna be a h- I, I would say that it's hands-on coding. We will do some live coding as well as assign homework, so I don't know. If people don't like homework, maybe the course is...

I mean, it's still kind of the course, but e- expect to kind of be in, in, feel like you're in a class for four weeks.

Course Evolution7:06

Alessio7:06

Yeah. Uh, I would say, so Hamel, this is like your second time kind of doing this. The first time was, I think-

Hamel Husain7:14

Mm

Alessio7:14

... the most successful Maven course ever. What have you... What are you doing again from last time, and then what are you doing differently?

Hamel Husain7:23

Yeah. It's mostly what am I doing differently because, like, first time, it was a course that started, it was about LLM fine-tuning, mostly like open models and stuff, and it kind of morphed into just it was a pile on-

Alessio7:37

Everything

Hamel Husain7:38

... of a lot, a lot of people, a lot of excitement, and, you know, I invited my friends and, uh, one thing led to another, and then it was like, "Let's..." And then it became like a conference.

Alessio7:47

Yeah.

Hamel Husain7:48

Which was not intentional, but it kind of, we just went with the flow of what people wanted. This time, we're gonna stay really targeted towards evals, 'cause we really wanna teach the subject. We don't want it to become a carnival, let's say.

You know, we want it to be, "Hey, we..." like we're really interested in teaching people evals, and we've been writing about it for so long that, you know, we want more information, you know, this like curated information.

I think like, yeah, what I learned from last time is, you know,

just, you know, the... it's mostly like the sort of logistics of doing a course which, you know, basically like how to, you know, how to... So I'm starting earlier. Um, last time was just sort of by the seat of my pants.

So.

Alessio8:41

Yeah, this one is a, there's a fair amount of, uh, lead time as well.

Yeah. Okay. Got it.

Shreya Shankar8:47

On the-

Hamel Husain8:48

You're being humble.

Shreya Shankar8:50

On the lead-

Hamel Husain8:50

Go for it.

Alessio8:51

No, no. Go ahead. Go ahead.

Hamel Husain8:53

Oh, I was just gonna say, Hamel's being very humble. I think that was a great course. It was like the first like massive LLM/AI course that I'd heard about, so I think that was very cool. Um, and there are things that we are taking from it, like we are figuring out interesting ways of integrating guest speakers or guest lectures.

Um, of course, I think the challenge is tying that to evals and then ensuring that that connection actually holds. But no, I, I think, I think Hamel did a great job, so we shouldn't be like, "Oh, that was...

We're doing everything differently."

Alessio9:31

Is the lead time a feature or a bug when you're trying to do things in, in AI? It almost feels like, you know, you're laying out the syllabus now, like you're probably starting to put together some content, and who knows what things are gonna change.

I think evals are maybe more time resistant in a way. They're less kind of like flavor of the week. So yeah, I'm curious how, how you think about that and kind of balancing tried and true things versus like a lot of people always just want to experiment with new stuff, but-

Hamel Husain9:58

Yeah

Alessio9:58

... sometimes it feels like a waste of time to me.

Hamel Husain10:00

Yeah. No, that's a really good question. So for the last year and a half or so, I've been actually... Like, I first started off being very hesitant about even talking about evals. I thought, "Okay, everyone knows it. Why do I even have to write about it?"

You know, surely people, you know... I don't, I don't know. I usually, uh, write for myself, so it's kind of strange for me to like write for other people. So I didn't do it for a while, and then it was only after a lot of encouragement, like people asking me the same question over and over again, that I started posting these.

And then I found that, like, the subject is pretty evergreen because we're not, you know, over the last year and a half, like the same principles apply and, you know, we're not really talking about like, you know, using specific tools and APIs.

This is more of a general process and like data literacy and how you go about analyzing data and how you think about that, that is missing, and that tends to... That's pretty stable. So I feel pretty good about it.

I feel that, hey, it's not going to change. The only thing that could change is if you truly get ASI, you know, or AGI, but I'm not, I don't, you know-

Alessio11:12

But then what do you need to eval?

Hamel Husain11:13

I don't, I don't do anything according to that. Yeah. This is like an existential thing. So, you know, barring that, I don't feel like things change too much.

Shreya Shankar11:21

Yeah. I also, I think techniques have stabilized. I think the kinds of failure modes of LLMs, I mean, they're still there, but it's not like changing every single day. We know that LLMs are bad at certain things. We, we know a little bit more about, say, limitations of the transformer architecture, and so we can kind of reason about what failure modes LLMs might have, especially, you know, for our data, for specific types of workloads.

Um, so I'm kind of seeing that now more and more, like the same pattern, the same advice that we're kind of telling people. Um, I tell Hamel things, and he knows everything that I tell him now, so I feel pretty comfortable that, you know, if I'm telling Hamel something that I think is new but it's not new, then we're probably ready to teach the course.

Hamel Husain12:04

And the thing that's really interesting is, like, Shreya's research is really good. Like, there's a lot of good UX patterns that make a lot of sense. You know, like, you know, you don't have to do every single step of evals manually, and a lot of those things have not been picked up by any commercial tooling yet, even though I use it in my workflow, like, you know, these concepts in my workflow.

So I feel like the, the field-

Swyx12:31

Something like this?

Hamel Husain12:31

... is very slow to adopt. Yeah, things, like in this paper, for example. You know, it's like one, one really huge thing about LLM as a judge, people love LLM as a judge, but you really have to make sure that you can trust the LLM as a judge.

LLM as Judge12:31

Hamel Husain12:47

And so how do you trust and how do you trust anything is that you have to check it, and you have to measure how good it is compared to some domain expert. There's no free lunch because, like, 'cause if a human has to trust it, you have to check it.

And so people don't have the... You know. So Shreya wrote this what I would say, like, a really long time ago in AI timelines.

Shreya Shankar13:13

That's true.

Hamel Husain13:13

Um, and you know, this is, this is not really appearing in any commercial tooling yet, even though we are using it, you know, with our clients and stuff like that. Um, and so it just goes to show, like, I mean, I think this is a very evergreen topic, like,

doesn't feel like... It feels like people need to know about these things and make, to make their life easier.

Shreya Shankar13:35

Yeah. Another thing to add to that, I mean, we don't talk about it in this paper, but, but people are very blocked in, you know, getting data to start testing their AI applications.

And it just feels like people don't know how to do anything other than plan A, which is to go out and try to collect as much real world data as possible and take months, because we're going to employ, like, human annotator teams to do this, or plan B, which is I'm going to ask an LLM and, like, never look.

I'm just going to prompt it for data,

and it's going to give me 100 samples, and it- I have no idea if that actually is what I would think is real. So really, like, you shouldn't do either one of them. You kind of want something in the middle, and what does that process look like?

Like, how do you kind of ground your own expectations or even domain experts' expectations into kind of synthetic generate, data generation? I feel like that stuff is really not talked about at all, but it's stuff that, you know, we think about on a pretty regular basis, and I'm excited to be able to teach that.

Swyx14:38

Yeah. Awesome. I mean, there, there's just a huge body of work. Uh, I think, uh, one thing I'm, I'm ki- I would like some baselines on is what percent of evals you see in the wild are LLM as judge.

Are we talking like, are people doing, like, 90% LLM as judge or, or like, you know, 30%?

Hamel Husain14:58

I feel like it's 75% LLM as a judge.

Swyx15:01

We should-

Hamel Husain15:01

Because it's low effort. It's kind of easy. But I would say out of the 75%, 80% is not helpful.

Swyx15:08

Yeah. I mean-

Shreya Shankar15:08

I think that's, that's a better metric to think about. It's like the cost of doing LLM as a judge now is quite low. So people, it's almost like a no-op to do it. Doing it right can give you a lot of value.

Doing it wrong could be negative if you're looking at it and trusting it all the time. But, uh-

Swyx15:24

And, and are you-

Shreya Shankar15:25

Yeah

Swyx15:25

... are you positive on, like, dedicated LLM as judge LLMs? So, like, uh, there's this, uh, we actually talked to Mahesh. I think his last name is, like, Satyamurty or something. Bespoke Labs is trying to do, like, sp- dedicated ones.

I, I'm sure there's, like, a couple others as well out there. Uh, you know, do, do you, like, is, is it, like, the large models or dedicated models?

Hamel Husain15:47

Yeah, and I also know, like, Haizelabs has something like that. Uh-

Swyx15:51

Oh, yeah. Okay

Hamel Husain15:51

... forgot what they called it. Um, starts with an R.

Swyx15:54

Yeah, they just launched this, like, at the conference.

Shreya Shankar15:56

Yeah.

Hamel Husain15:56

Starts with an R. Forgot the name. It's, uh...

Swyx16:00

I'll, I'll look it up.

Hamel Husain16:01

Anyways. Yeah. So th- those can be interesting. Like, the main thing to focus on is, like, I put very little value in benchmarks, like general benchmarks. It's, has some value, but, you know, what you really need to do is, like, measure it in your domain and see if, if that, uh, LLM as a judge is more aligned than, like, an off-the-shelf LLM.

And what I've found is, like, it's kind of, there's n- there's, yeah, it doesn't seem like there's a free lunch that I can see yet, at least on the clients that I'm working with. But I think that's the part that people are missing, right?

They're just saying, "Okay, like, so and such and such put out these benchmarks for LLM as a judge." They just use those LLM as a judge, but they don't go about measuring how good the LLM as a judge is.

And so, yeah.

Swyx16:52

Yeah. Uh, and that's something that obviously is, is, is work that needs to be done. Uh, it's called Verdict.

Hamel Husain16:57

Oh, Verdict. So it doesn't start with R.

Swyx16:59

Library for scaling-

Hamel Husain17:00

I'm sorry.

Swyx17:00

There is an R in there. Uh, scaling judge time compute. Yeah. I think the, the, the thing that throws me off is, like, inventing your own scaling for, uh, for compute, um, seems a little strange. Yeah, I d- I d- I didn't, I didn't read much into this.

I don't know.

Hamel Husain17:17

I didn't get a tr- I didn't get a chance to try this one specifically yet because it's so new.

Swyx17:23

Any other trends you, you, you see in evals that we should cover?

Hamel Husain17:26

Trends in evals. Okay. Well, I think that people-

Evals Trends17:26

Swyx17:29

Maybe just patterns that you think that is-

Hamel Husain17:30

... also thinking this, this is the, this is a really big one. One counterintuitive thing that has extreme value that people kind of discover maybe accidentally or, you know, if they're working with us, they discover very fast, is, is this, there's a really...

So you really want to look at your data a lot, and there's a really high payoff to building your own application, like your own little web application that lets you annotate your data. And like, why is that? Well, it turns out that a lot of data is, has domain-specific elements to it, meaning you have like various metadata about your traces.

You know, it doesn't live in a silo. And when you want to see that data You wanna render it in a way that allows you to check it as fast as possible. So it's not just you might wanna render the trace in a spec- very specific way, maybe it has like markdown in that trace, maybe it has like rich elements, uh, maybe it has widgets.

Every, every application is different, and what you wanna do is render it in a way that allows you to see all the context you need, both the trace itself, but around the trace, things even external to the trace, and really dial it in for your domain and, you know, allow you to annotate data and make notes and do your error analysis.

And, you know, AI is really, really good. Use Lovable, use Cursor, whatever, and it's, it's a really high payoff. And sp- I think people are starting to discover that.

Swyx18:51

Something that, uh, you just reminded me is, uh, Y- Eugene's, um, Align Eval. Would this count as one of those dedicated things? I guess this is still, still too generic, right?

Hamel Husain19:02

I think this is the definitely the right spirit, and I love Eugene. He's gonna... We're gonna have him as a guest speaker, by the way, in the course.

Swyx19:10

Nice.

Hamel Husain19:11

Eugene is basically one of my favorite people in AI. And, you know, he's one of the people that also realized this importance of aligning the judge to a human, and that's what this product, this, uh, project is about.

It's... And it's like helps people grok the idea by h- like giving them an exercise to do that is kind of fun and it ga- gamifies it. So yeah, I mean, it's like this is the right idea. It's the right idea in a lot of ways.

So like not only is this like, you know, a nice web application, but it's like makes it fun. And these are the kinds- ... of things you should try to do to like, you know, like make it, make it as painless as possible to look at data.

So that's like one of the things that Eugene is, like doing here.

Swyx19:55

It's like the Duolingo-ification of, uh...

Shreya Shankar19:59

It's very interesting, and I think he did a good job figuring out the right way to render these like paragraph long summaries. I think going back to what Hamel was saying before, people's LLM applications, like outputs might not be just paragraphs long.

It might be words, it might be numbers, it might be long essays. Um, it like, it might be code, it might be charts. So, you know, just having the right way to render that so you can really quickly scan many of them is super helpful.

And, and there's just no kind of one size fits all. It's, yeah, like there's no like Excel like equivalent, you know, that everybody's gonna find is the perfect way to look at their data. Um, I love Excel. I think it's great.

You sh- You do Excel over nothing, but at some point people kind of get to that next level where they need something more, uh, custom for their use case.

Tools & Culture20:52

Swyx20:52

I would say like, yeah, spreadsheets actually is a surprisingly minimal viable UI. The, the ones that, the one that I reach for, you know, you know, fair disclosure, I'm a, I'm an investor, is Quadratic. I don't know if you guys have seen.

Uh-

Shreya Shankar21:05

No

Swyx21:05

... you know, in the way, this- in the same way that I think, uh, Hamel likes Merimo, which is like a sort of Python reinvention of the notebook. This is a Python reinvention of the s- the spreadsheet. Um, and, uh, yeah, you can, you can, uh, you know, do all, do all sorts of-

Shreya Shankar21:20

Oh, nice.

Swyx21:23

This is a spreadsheet, right?

Hamel Husain21:24

Yeah, definitely.

Swyx21:24

That just-

Hamel Husain21:26

Yeah

Swyx21:26

... that just has some custom cells and then some, some custom code behind it, and Quadratic lets you write Python or, uh, yeah, it, it, it seems like I can't really load it right now here.

Hamel Husain21:36

Yeah, it's, it's, uh, it's really interesting. I think we talked about it before Swix, like, so one of the, the foundation of Evals is error analysis. So like looking at your data and doing data analysis on-

Swyx21:49

Yes

Hamel Husain21:50

... your, your traces.

Swyx21:51

Yes.

Hamel Husain21:51

So a lot of people, when we say data literacy, that, that can mean, that can sound scary, but it can come from a lot of different places. It can be, you can... You know, it's not just data scientists and stuff like that.

It's like people that have worked in finance, people that have worked as business analysts. A lot of people have some basic data literacy, you know, and like, you know, if you know how to use a pivot table, you can go a really long ways to do some basic data analysis.

And there's a lot of like latent capabilities for people to get involved with like being part of Evals. It's not like there's some, you know, the bar doesn't have to be so high as being a data scientist, let's say.

So I think that's interesting, you know. In a lot, a lot, in a lot of situations, you can use a spreadsheet. Not always, but you can. You can start there, and you can muck around, and that's important.

Swyx22:40

Yeah. Yeah. My, my, my first, uh, e- uh, pi- uh, spreadsheet pilling was, uh, Clouden Sheets, which I, I always, I always love telling the story of how I got access to this while interviewing at Anthropic, and that actually prevented me from building so much 'cause I was like, "Why would you do anything?

You just use Clouden Sheets." Uh, and obviously that was wrong. Okay. Yeah, so cool. Uh, you know, like I, I think this is topic, you know, it's, it's, it's weird because it's like not everyone knows it's important, and I think you get a lot of points every time you say Evals is important, FYI.

And then, you know, how do you make it like actually stick? Um, it occurred to me that I think the similar thing that happened was kind of like the Six Sigma movement. You need like some cool name like Six Sigma Black Belt, right?

Hamel Husain23:24

Mm.

Swyx23:25

So you can be like some Align Evals or, um, whatever, EvalGen Black Belt. And to be a Black Belt, you, you have to do like this, this number of Evals. You have to like, you know, get your organization to commit to, to, uh, doing s- achieving some kind of standard.

But just like slap some cool title on it. People will do it.

Hamel Husain23:44

So do you think it's like the name and the title would help? 'Cause like one thing that I wrestle with is that Evals is a solution, but the problem is, okay, your AI doesn't work or it doesn't work as well as you want it to, and people don't associate the solution with the problem cleanly enough.

'Cause they don't know. They don't know what they don't know.

Swyx24:06

Well, you know, I think you're gonna educate people on that and, and at some point, like you, you, you'll run it, you'll independently reinvent it no matter, um, how, where, where you're starting from. You just, you have to.

So I, I have some f- like I think I have some faith that like- Anyone who's, like, actually serious enough will, will find the problem, and then, you know, y- here you are sitting with the solution. So I don't...

I find it's, I find it's okay. I, I just think, like, people just need, like, a, like, an agile manifesto, like a, like a really stamped way to do things that is, that, that has enough sort of, uh, broad appeal.

Um, and you know, like, you know, the, the way that I often do it is I just, like, I put a, put a cool name to the thing, and I just really give it a lot of wood, and, and then it just starts taking off by itself.

Like, you don't really need to do anything else beyond that. So, like, yeah. D- uh, I don't know. That, that's something, that's something I think about.

Hamel Husain24:53

Maybe we could rain on that.

Swyx24:54

Yeah.

Shreya Shankar24:54

Yeah. That's good advice. Here we are just like, "Oh, let's have some cool assignments." I need to think bigger.

Swyx24:59

I think this is, this is part of the process. You're figuring it out, and like y- y- I, I actually, I tell people not to force these things, right? Like, it'll come when it comes, and, like, it'll be obvious.

You... Like, you... It'll be the something that you naturally describe in one of your, uh, your, your... Then you're like, "Oh, you know, I really like the way that rolled off my tongue." Like, people start using it organically, repeating it back to me.

That's the sign of the traction. You don't, you can't force it from, you know, top down.

Shreya Shankar25:23

Yeah.

Swyx25:23

You know? Anyway, cool. Yeah. C- uh, congrats on, on putting this together, and I think, uh, we wanna remind people that there's some kind of lightning lesson. That's, uh, it's like a free sampler course or something.

Lightning Lesson25:23

Hamel Husain25:32

Yes. Yeah. And then that's, we're gonna talk a, we're gonna talk a bit about error analysis, that first, the very first stage. We're gonna go over that, so it's a good teaser.

Swyx25:44

Awesome.

Shreya Shankar25:44

We've announced the date, right? Or we haven't announced the date-

Swyx25:48

Yeah

Shreya Shankar25:49

... of Lightning Lesson.

Swyx25:49

It is March 21st, so we got 10 days.

Shreya Shankar25:52

Perfect.

Swyx25:53

Yeah. Uh, good luck. Uh, I think you're doing, fighting the good fight. Ha- always happy to spread the word.

Shreya Shankar25:58

Thank you.

Swyx25:58

For me, the challenge is always with my conferences, how to present, uh, new trends on evals, which is why I, I asked you about trends. So it's, so want, I want us to innovate as a field while obviously s- reminding people of the fundamentals, and it's always a tricky balance.

Hamel Husain26:15

Yeah. I think for new trends, it's really, like, Shreya has a few different papers of good... Uh, she does r- like, really good research into human-computer interaction and UX, and also evals in the workflow, and those things are kind of...

Like, if I were to fast-forward one or two years, I would expect to see those in all the tools. It just hasn't arrived yet. I think people are slow, but, you know, I've definitely-

Shreya Shankar26:42

Well, it might not be perfect, so I think it's gonna take adaptation.

Hamel Husain26:46

Yeah. I mean, it certainly works really well for me, like, on many clients. I mean, I adapt a little bit, but...

Shreya Shankar26:52

Yeah. So.

Swyx26:54

You know, I think there's, there's some interesting work right now on verifiers, which is not quite the same thing, but I think at some point this starts to coincide, collapse into the s- into that field, and maybe you can make it interesting from the point of view of, you know, like, you also need evals for chains of thought.

Um-

Shreya Shankar27:11

Yeah

Swyx27:11

... and that, that could be a cool thing.

Shreya Shankar27:13

You would need evals to train your reasoning models.

Swyx27:15

All right. Awesome. Well, thank you so much for joining us, and this has, this has been a nice, uh, quick catch-up.

Hamel Husain27:19

Yeah. Thank you.