Intro0:00
Hey everyone, welcome to the Latent Space Podcast. This is Alessio, partner and CTO in residence at Decibel Partners, and I'm joined by my co-host, Swyx, founder of Smol.ai.
Hey, and today we're very happy and excited to welcome Anastasios and Weilin from LMSys. Welcome, guys.
Hey. How are you?
Hey.
Nice to see you. Thanks for having us.
Anastasios, I actually saw you, I think, at last year's NeurIPS. You were presenting a paper which I don't really, uh, super understand, but it was some theory paper about how your method was very dominating over other sort of search methods.
I don't remember what it was, but, uh, I remember that you were a very confident speaker.
Oh, I totally remember you. Didn't ever connect that, but yes, that's definitely true. Yeah, nice to see you again.
Yeah, I was frantically looking for the name of your paper and I couldn't find it. Basically, I had to cut it 'cause I didn't understand it.
Was this conformal PID control-
Conformal, yes
... or was this the all control? Blast from the past, man.
Blast from the past. Uh, it's always interesting how NeurIPS and, you know, all these academic conferences are sort of six months behind where what, you know, people are actually doing. But conformal risk control, I would recommend people check it out.
We-- I have the recording, I just never published it just 'cause I was like, "I don't understand this enough to explain it," so
Yeah. I mean, you're like, "People won't be interested." It's all good.
But ELO scores, ELO scores are very easy to understand. You guys are responsible for the biggest revolution in language model benchmarking in the last few years. Maybe you guys wanna introduce yourselves and maybe tell a little bit of the brief history of LMSys.
Origins1:17
Hey, I'm Weilin. I'm a fifth-year PhD student at UC Berkeley, uh, working on ChatBot Arena these days, doing crowdsourcing AI benchmarking.
I'm Anastasios. I'm a sixth-year PhD student here at Berkeley. I did most of my PhD on like theoretical statistics and sort of foundations of model evaluation and, and testing, and now I'm working 150% on this ChatBot Arena stuff.
It's great.
And what was the origin of it? How did you come up with the idea? Uh, how did you get people to buy in? And then maybe what were one or two of the pivotal moments early on that kind of made it the standard for, for these things?
Yeah, yeah. ChatBot Arena project was started last year in April, May, around that. Before that, we were basically experimenting in the lab how to fine-tune a chatbot open source based on the Llama 1 model Meta released. At that time, Llama 1 was like a base model, and people didn't really know how to fine-tune it, so we were doing some explorations.
We were inspired by Stanford's Alpaca project. So we basically, yeah, crawl a dataset from the internet, which is called ShareGPT dataset, which is like the dialogue dataset between user and, and ChatGPT conversation. It turns out to be like pretty high quality data, dialogue data.
So we fine-tune on it, and then we train it and release the model called Vicuna. And people were very excited about it because it kinda like demonstrate open weight model can reach this conversation capability similar to ChatGPT. And then we basically release the model with and also build a demo website for, for the model that people were very excited about it.
But during the, the open, the biggest challenge to us at the time was like how do we even evaluate it? How do we even argue this model we trained is better than others? And like what's the gap between this open source model and other proprietary offering?
At that time it was like GPT-4 was just announced, and it's like Claude 1, right? What's the difference between them? And then after that, like every weeks, there's a new model being fine-tuned, released. So even until still now, right?
And then we have that demo website for Vicuna, and then we thought like, "Okay, maybe we can add a few more open model as well, like API model as well." And then we quickly realized that people need a tool to compare between different models.
So we have like a side-by-side UI implemented on the website too, that people choose, you know, to compare. And we quickly realized that maybe we can do something like, like a battle on top of these LLMs, like just anonymize it, anonymize the identity, and that people vote which one is better.
So the community decide which one is better, not us, not us arguing, you know, our model is better or what. And that turns out to be like people were very excited about this idea. And then we tweet, we launch in last, yeah, last April, May, and then it was like first two, three weeks, like just a few hundred thousand views, tweet on our launch tweet.
And then we have re- regularly double update weekly at the beginning at the time, adding new model GPT-4s as well. So that was like, that was the, the, you know, the like initial-
Another pivotal moment just to jump in would be private model testing, like the GPT, I'm a little, I'm a little chat bot.
Yeah, yeah. That was this year. Yeah.
That was this year.
Early this year.
Huge. That was also huge.
In the beginning, I saw the initial release was May 3rd of the leaderboard. On April 6th, we did a Benchmarks 101 episode for our podcast, just kind of talking about, you know, how so much of the data is like in the pre-training corpus and blah, blah, blah, and like the benchmarks are really not what we need to evaluate whether or not a model is good.
Why did you not make a benchmark? Maybe at the time, you know, it was just like, "Hey, let's just put together a whole bunch of data again, run a mega score." That seems much easier than coming up with a whole website where like users need to vote.
Any thoughts behind that?
Arena vs Benchmark5:41
Thing is more like fundamentally we don't know how to automate this kind of benchmarks when it's more like, you know, conversational multi-turn and more open-ended task that may not come with a ground truth in. So let's say if you ask a model to help you write an email for you for whatever purpose, there's no ground truth.
How do you score that, right? Or write a story or a creative story, or many other things. Like how we use ChatGPT these days, it's oftentimes more like open-ended. You know, we need human in a loop to give us feedback which one is better.
And I think the nuance here is like sometimes it's also hard for human to give the absolute rating, so that's why we have this kind of pairwise comparison easier for people to choose which one is better. So from that we, you know, use these pairwise comparison votes to, to calculate the leaderboard.
Mm-hmm.
Yeah. You can ask us, you can know more about this methodology behind it.
Yeah. I mean, I think the, the point is that, and you guys probably also talked about this at some point, but, uh, static benchmarks are in- intrinsically to s- some extent unable to measure generative model performance. And the reason is because you cannot pre-annotate all the outputs of a generative model.
You change the model, it's like the distribution of your data is changing. New labels to deal with that. New labels are great automated labeling, right? Which is wha- which is why people are pursuing both. And yeah, static benchmarks, they allow you to zoom in to particular types of information like factuality, historical facts.
We can build the best benchmark of historical facts, and we will then know that the model's great at historical facts. Okay. But ultimately, that's not the only axis, right? And we can build 50 of them, and we can evaluate 50 axes, but it's just so...
The problem of generative model evaluation is just so expansive and it's so subjective that it's just maybe not intrinsically impossible, but at least we, we don't see a way, you know? We didn't see a way of encoding that into a fixed benchmark.
But on the other hand, I think there is a challenge where this kind of like online dynamic benchmark is slow, is more expensive than static benchmark or offline benchmark where-
Yeah
... people still need it, like when they build models.
Yeah, of course. Yeah.
They need, they need static benchmark to track where they are.
Yeah. It's not like our benchmark is like uniformly better than all other benchmarks, right? It just measures a different kind of performance that has proved to be useful.
You guys also published MT-Bench as well, which is a static version, let's say, of ChatBot Arena-
Yeah
... right? That, that people can actually use in their development of models.
Right. That's I think one of the reason we still do this static benchmark, we still wanted to explore, experiment whether we can automate this because people eventually model developer need it-
Mm-hmm
... to fast iterate their model.
Yeah.
So that's why we explored like LMSys Judge and Arena Hard try to filter, you know, select the high quality data we collected from ChatBot Arena, the high quality subset, and use that as the questions and then automate the judge pipeline so that people can quickly get high quality signal, benchmark signals using this offline benchmark.
Community Building9:03
You know, as a community builder, I'm, I'm curious about just the initial early days. Obviously, when you offer effectively free A/B testing inference for people, people will come and use your arena. What do you think were like the, the key unlocks for you?
Was it like funding for this arena? Was it like marketing? When people came in, do you see a no- noticeable skew in, in the data? Which obviously now you have enough data sets you can separate things out like coding and hard problems, but in the early days it was just all sorts of things.
Yeah. Maybe one thing to establish at first is that our philosophy has always been to maximize organic use. I think that really does speak to, to your point, which is, yeah, you know, why do people come? They came to use free LLM inference, right?
And also a lot of users just come to the website to use direct chat, because you can chat with a model for free. And then you could think about it like, "Hey, let's just be kind of like more on the selfish or conservative or protectionist side," and say, "No, we're only giving credits for people that battle," or, you know, so on and so forth.
The strategy wouldn't have worked, right? Because what we're trying to build is like a big funnel, a big funnel that can direct people. And some people are passionate and interested, and they battle. And yes, the, the distribution of the people that do that is different.
It's like as you're pointing out, it's like that's not necessarily-
Enthusiastics
... they're enthusiastic. They're-
Yeah. Early adopter of the, this technology.
Or they like games.
Yeah.
They like ga- you know, people like this. And we've run a couple of surveys that indicate this as well of like our user base. Yeah.
We do see a lot of like developers come to the site asking coding questions, 20, 30%.
Yeah, 20, 30%. That's obviously not reflective of the general population.
Yeah, for sure. No.
But it's like reflective of some corner of the world of people that really care. And to some extent, maybe that's all right because those are like the power users and, you know, we're not trying to claim that we represent the world, right?
We represent the people that come and vote.
Did you have to do anything marketing wise? Was anything effective? Did you struggle at all or was it success from day one?
At some point, almost like-
Okay
... because as you can imagine, this leaderboard depends on part user, like part, like community engagement, participation. If no one come to vote tomorrow, then no leaderboard. So-
We had some period of time when the number of users was just after the initial launch, it went, it just went lower.
Yeah.
And it, you know, at some point it didn't, it did not look promising.
Yeah. Right.
Actually, you know, I joined the project a, a couple months in to do the statistical aspects, right? As you can imagine, that's how it kind of hooked into my previous work. At that time, it wasn't like it, you know, it definitely wasn't clear that this was like gonna be the eval or something.
It was just like, "Oh, this is a cool project. Like, Weiliang seems awesome," you know, and that's it.
Definitely there's, there's in the beginning, because people don't know us, people don't know what this is for, so we had hard time But I think we were lucky enough that we have some initial momentum and as well as the competition between-
Model providers
... model, model providers just becoming, you know-
Yeah, it became very intense
... intense.
Very intense.
And then that makes the above fun to watch, right? Because always number one is number one.
There's also an element of trust. Our main priority in everything we do is trust.
Yeah.
We wanna make sure we're doing everything, like all the I's are dotted and the T's are crossed, and nobody gets unfair treatment. And people can see from our profiles and from our previous work and from whatever that, you know, we're trustworthy people.
We're not, like, trying to make a buck, and we're not trying to become famous off of this or something. It's just we're trying to provide a great public leaderboard.
Community venture project.
Community venture.
Yeah. Yes.
I mean, you, you are kind of famous now. I don't... You know, nobody recognizes me, but that's fine.
Just to dive in more into biases and, you know, of some of this is like statistical control. The classic one for human preference evaluation is humans demonstrably prefer longer contexts or longer outputs, which is actually something that we don't necessarily want.
You guys, I think maybe two months ago, put out, uh, some length control studies. Apart from that, there, there are just other documented biases. Like I, I'd just be interested in your review of the what you've learned about biases and, uh, maybe a little bit about how you've controlled for them.
Style Control13:16
At a very high level, yeah, humans are biased. Totally agree, like in various ways. It's not clear whether that's good or bad. You know, we try not to make value judgments about these things. We just try to describe them as they are.
And our approach is always as follows: We collect organic data, and then we take that data and we mine it to get whatever insights we can get. And, you know, we have m- many millions of data points that we can now use to extract insights from.
Now, one of those insights is to ask the question: What is the effect of style, right? You have a bunch of data. You have votes. People are voting either which way. We have all the conversations. We can say, "What components of style contribute to human preference, and how do they contribute?"
Now, that's an important question. Why is that an important question? It's important because some people want to see which model would be better if the lengths of the responses were the same, were to be the same, right? People wanna see the causal effect of the model's identity controlled for length or controlled for markdown, number of headers, bulleted lists.
Is the text bold? Some people don't, they just don't care about that. The idea is not to impose the judgment that this is not important, but rather to say, ex post facto, can we analyze our data in a way that decouples all the different factors that go into human preference?
Now, the way we do this is via statistical regression. That is to say, the arena score that we show on our leaderboard is a particular type of linear model, right? It's a linear model that takes, it's a logistic regression that takes model identities and fits them against human preference, right?
So it regresses human preference against model identity. What you get at the end of that logistic regression is a parameter vector of coefficients. And when the coefficient is large, it tells you that GPT-4.0 or whatever, very large coefficient, that means it's strong, and that's exactly what we report in the table.
It's just the predictive effect of the model identity on the vote. Another thing that you can do is you can take that vector. Let's say we have M models. That is an M-dimensional vector of coefficients. What you can do is you say, "Hey, I also wanna understand what the effect of length is."
So add, I'll add another entry to that vector, which is trying to predict the vote, right? That tells me the difference in length between two model responses. So we have that for all of our data. We can compute it ex post facto.
We add it in as the regression, and we look at that predictive effect. And then the idea, and this is formally true under certain conditions, not, not always verifiable ones, but the idea is that adding that extra coefficient to this vector will s- kind of suck out the predictive power of length and put it into that M plus first coefficient and, quote-unquote, "de-bias" the rest so that the effect of length is not included, and that's what we do in style control.
Now, we don't just do it for M plus one. We have, you know, five, six different style components that have to do with markdown headers and bullet- bulleted lists and so on that we add here. Now, where is this going?
Y- you guys see the idea. It's a general methodology. If you have something that's sort of like a nuisance parameter, something that exists and provides predictive value, but you really don't want to estimate that, you wanna remove its effect.
It's, uh, in causal inference, these things are called, like, confounders often. What you can do is you can model the effect. You can put them into your model and try to adjust for them. So another one of those things might be cost.
You know, what if I wanna look at the cost-adjusted performance of my model? Which models are punching above their weight, parameter count? Which models are punching above their weight in terms of parameter count? We can ex post facto measure that.
We can do it without introducing anything that compromises the organic nature of the data that we collect. Hopefully, that answers the question.
It does. For someone with a background in econometrics, this is super familiar.
You're probably better at this than me, for sure.
Well, I mean, uh, so I used to be, you know, a, a quantitative trader, and so, you know, controlling for-
Oh, okay
... multiple effects on, on stock price is effectively the job. Um, so it's interesting. Obviously, the problem is proving causation, which is hard, uh, but you don't have to do that.
Yes. Yes, that's right, and causal inference is a hard problem, and it goes beyond statistics, right? It's like you have to build the right causal model and so on and so forth. But w- we think that this is a good first step, and we're sort of looking forward to learning from more people.
You know, there's some good people at Berkeley that work on causal inference. Forward to learning from them on, like, what are the really most contemporary techniques that we can use in order to estimate true causal effects if possible.
Maybe we could take a, a step through the other categories. So Style Control is a category, it is not a default. I have thought that when you wrote that blog post, actually, I thought it would, it would be the new default because seems like the, the most obvious thing to control for.
But you also have other categories. You have coding, you have hard prompts.
We consider that we're still actively considering it.
Category Leaderboards18:27
Okay.
It's just, you know, once you make that step, once you take that step, you're introducing your opinion. And I'm not-- you know, why should our opinion be the one? That's kind of a community choice. We could put it to a vote.
We could ask-
Yeah, maybe we do a poll.
Maybe do a poll. I don't know.
No opinion is an opinion, you know what I mean?
Yeah.
There's no neutral choice here.
That's fair.
Yeah, you, you have all these others. You have instruction following too. Just, like, pick a few of your favorite categories that you'd like to talk about. Maybe tell a little bit of the stories. Te-tell a little bit of, like, the hard choices that you had to make.
Yeah, yeah, yeah. I think the, uh, initially, the reason why we want to add these new categories essentially to answer comm- some of the questions from our community, which is you won't have a single leaderboard for everything. So these models behave very differently in different domains.
Let's say this model is trained for coding, this model trained for more technical questions, and so on. On the other hand, to answer people's question about like, okay, what if all these low quality... You know, because we crowdsource data from the internet, there will be noise.
So how do we de-noise? How do we filter all these low quality data effectively? So that was like, you know, some questions we want to answer. So basically, we were-- we spent a few months, like, really diving into these questions to understand how do we filter all these stuff, because these data are like millions of data points.
And then if you want to label it yourself, it's possible. But we need to kind of like to automate this kind of data classification pipeline for us to effectively categorize them into different categories, say coding, math, instruction following, and also harder prompts.
So that was like the hope is when we slice the data into these meaningful categories to get... give people more, like better signals, more direct signals. And that's also to clarify what we are actually, actually measuring for, because I think that's the core part of the benchmark.
That was the initial motivation. Does that make sense?
Yeah. Also, I'll, I'll just say this does, like, get back to the point that the philosophy is to like mine organic, to, to take organic data and then mine it ex post facto.
Is the data cage-free too, or just organic?
It's cage-free. It's grass, grass-fed. No GMO. Yeah, and all of these efforts are like open source. Like we open source all the data cleaning pipeline, filtering pipeline. Um-
Yeah, I love the notebooks you guys publish. Actually really good-
Yeah
... just, just for learning statistics.
Yeah, yeah. I will share this, uh, insights to everyone.
I agree on the initial premise of, hey, writing an email, writing a story, there's like no ground truth. But I think as you move into like coding and like red teaming, some of these things, there's like kinda like skill levels.
So I, I'm curious how you think about the distribution of skill of the users. Like maybe the top one percent of red teamers is just not participating in the arena. So how, how do you guys think about adjusting for it in like fields like this where there's kinda like big differences between the average and the, the top?
Yeah. Red teaming, of course, red teaming is quite challenging. So, okay, moving back, there's definitely like some tasks that are not as subjective that like pairwise human preference feedback is not the only signal that you would want to measure.
And to some extent, maybe it's useful, but it may be more useful if you give people better tools. For example, it'd be great if we could execute code within Arena. It'd be fantastic. We wanna do it. There's also this idea of constructing a user leaderboard.
What does that mean? That means some users are better than others, and how do we measure that? How do we quantify that? Hard in chatbot arena, but where it is easier is in red teaming. 'Cause in red teaming, there's an explicit game.
You're trying to break the model, you either win or you lose. So what you can do is you can say, "Hey, what's really happening here is that the models and humans are playing a game against one another." And then you can use the same sort of Bradley-Terry methodology with some, some extensions that we came up with in one of-- You can read one of our recent blog posts for, for the sort of theoretical extensions.
You can attribute like strength back to individual players and jointly attribute strength to like the models that are in this jailbreaking game, along with the target tasks, like what types of jailbreaks you want. So yeah, and I think, uh, th-this, this is a hugely important and interesting avenue that we wanna continue researching.
We have some initial ideas, but you know, all thoughts are welcome.
Yeah. Well, first of all, on the code execution, the E2B guys, I'm sure they'll be happy to, to help you, help you set that up. They're big fans. We're investors in a company called Dreadnought, which, uh, we do a lot in AI red teaming.
I think to me the most interesting thing has been how do you do... Sure, like the model jailbreak is one side. We also had Nicola Scarlini from DeepMind on the podcast, and he was talking about, for example, like, you know, context stealing and like, uh, weight stealing.
So there's kinda like a lot more that goes around it. I'm curious just how you think about the model and then maybe like the broader system. Even with Red Team Arena, you're just focused on like jailbreaking of the model, right?
You're not doing kinda like any testing on the more system level thing of the model where like maybe you can get the training data back, you can exfiltrate some of the layers and the weights and, and things like that.
So right now, as you can see, the Red Team Arena is at very early stage, and we are still exploring what could be the potential new games we can introduce to the platform. So the idea is still the same, right?
Can we build a community-driven project platform also for people, they can have fun with this website for sure. That's one thing. And be used and then help everyone to test these models. So one of the aspect you mentioned is stealing secrets, right?
Stealing training sets. That could be one, you know, could be designed as a game. Say, can you still use their credential? You know, we hide-- maybe we can hide the credential into system prompts and so on. So there are, like, a few potential ideas that we want to explore-
Yeah.
For sure. You want to add more?
I think that this is great. This idea is a great one. There's a lot of great ideas in the red teaming space. You know, I'm not personally like a red teamer. I don't like go around in red team models.
But there are people that do that, and they're awesome. They're super skilled. And when I think about the Red Team Arena, I think those are the really the people that we're building it for. Like, we wanna make them excited and happy, build tools that they like.
And just like ChatBot Arena, we'll trust that this will end up being useful for the world. And all these people are, you know, I won't say all these people in this community are actually good-hearted, right? They're not doing it because they wanna, like, see the world burn.
They're doing it because they, like, think it's fun and cool. And yeah. Okay, maybe they wanna see, maybe they want a little-
Not, not all. Some. Majority.
Yeah. You know what I'm saying. So, you know, trying to figure out how to serve them best, I think. I don't know where that fits. I just... I'm not expressing-
And give them credits, right?
And give them credit, yeah.
Yeah.
So I'm not trying to express any particular value judgment here as to whether that's the right next step. It's just... That's, that's sort of the way that I think we would think about it.
Yeah. Then we also talked to Sander Schulhoff of the HackerPrompt competition, and he's pretty interested in red teaming at scale. Let's just call it that. Um, you guys maybe wanna talk with him.
Oh, nice.
We wanted to cover a little, a few topical things and then go into the other stuff that your group is doing. You know, you're not just running ChatBot Arena. We can also talk about the new website and your, your future plans.
But I just wanted to briefly focus on o1. Uh, it is the, the hottest, latest model. Obviously, you guys already have it on the leaderboard. What is the impact of o1 on your evals?
o1 Impact26:06
Made our interface slower.
Exactly.
You made it slower.
Yeah.
Yeah.
Because it needs like 30, 60 seconds, sometimes even more to, uh... The latency is, like, higher. So that's one for sure. But I think we observe very interesting, uh, things from this model as well. Like, we observe, like, significant improvement in certain categories, like more technical one, math categories.
Yeah. I think actually, like, one takeaway that was encouraging is that I think a lot of people before the o1 release were thinking, "Oh, like, this benchmark has saturated." And why were they thinking that? They were thinking that because there was a bunch of models that were kind of at the same level.
They were just kind of like incrementally competing, and it sort of wasn't immediately obvious that any of them were any better. Nobody, including any individual person, it's hard to tell. But what o1 did is it was-- it's clearly a better model for certain tasks.
I mean, I used it for like proving some theorems and, you know, there's some theorems that, like, only I know because I still do a little bit of theory, right? So it's like I can go in there and ask like, "Oh, how would you prove this exact thing?"
Which I, I can tell you has never been in the public domain. It'll do it. It's like, what? Okay, so there's this model, and it crushed the benchmark. You know, it's just like, really like a big gap. And what that's telling us is that it's not saturated yet.
And so it's still measuring some signal. That was encouraging.
Yeah.
That point, the takeaway is that the benchmark is comparative. It's not-- There's no absolute number. There's no maximum ELO. It's just like, if you're better than the rest, then you, then you win. I think that was actually quite helpful to us.
I think people were criticizing... I saw some, some of the academics criticizing it as not apples to apples, right? Like, because it can take more time to reason, it's basically-
Yeah
... doing some search, doing some chain of thought that if you actually let the other models do that same thing, they might do better.
Absolutely. But I mean, to be clear, none of the leaderboard currently is apples to apples.
Yeah.
Because you have like Gemini Flash, you have, you know, all sorts of tiny models like Llama 8B. Like 8B and 405b are not apples to apples. Totally agree.
Yeah, they have different latencies.
Different latencies. So when we-
Control for latency.
Yeah, latency control. That's another thing. We can do Style Control. We'll latency control. You know, things like this are important if you wanna understand the trade-offs involved-
Yeah
... in using AI.
o1 is a developing story. We still haven't seen the, the full model yet, but, um, you know, it's, uh, it's definitely-
Yeah
... a very exciting new, new paradigm. I think one community controversy I just wanted to give you guys space to address is the collaboration between you and the large model labs. People have been suspicious, let's just say, about how they choose to A/B test on you.
Lab Collaboration28:40
I'll state the argument and let you respond, which is basically they like have... They run like five anonymous models and basically argmax their ELO on LMSys or, or ChatBot Arena, and they'll, they release the best one, right? Like, what has been your end of the, the controversy?
How, how have you decided to clarify your policy going forward?
On a high level, I think our goal here is to build the best eval for everyone, and including everyone in the community can see the leaderboard, can understand, compare the models. More importantly, I think we want to build best eval also for, for model builders.
Like all these frontier labs building model, they also internally facing a challenge, which is, you know, how do they eval the model. So the- that's the reason why we want to partner with all the frontier lab people, not just...
And then to help them testing. So that's one of the... We want to solve this technical challenge, which is eval.
Yeah. I mean, ideally it benefits everyone, right?
Yeah. And ideally, if model for other people-
And people also are interested in like seeing the leading edge of the models. People in the community seem to like that. You know, "Oh, there's a new model out," it's, you know, "Is this strawberry?" People are excited. People are interested.
Yeah.
And then there's this question that you bring up of, is it actually causing harm?
Mm-hmm.
Right? Is it causing harm to the benchmark that we are allowing this private testing to happen? Maybe like stepping back, why do you have that instinct? The reason why you and others in the community have that, have that instinct is because when you look at something like a benchmark, like an ImageNet, a static benchmark, what happens is that if I give you a million different models w- that are all slightly different and I pick the best one, there's something called selection bias that plays in, which is that the performance of the winning model is overstated.
This is also sometimes called the winner's curse, and that's because statistical fluctuations in the evaluation, they're driving which model gets selected as the top. So this can... s- selection bias can be a problem. Now, there's a couple things that make this benchmark slightly different.
So first of all, the selection bias that you inc- when you're only testing five models is normally empirically small. Um-
And that's why we have these kind of like confidence-
Yeah
... interval constructed.
That's right. Yeah, our confidence intervals are actually not multiplicity adjusted, but one thing that we could do, like immediately tomorrow in order to like address this concern is if a model provider is testing five models and they want to release one, and we're constructing the models at level like one minus alpha, we can just construct the intervals instead at level one minus alpha divided by five.
That's called a Bonferroni correction. What that'll tell you is that like the final performance of the model, like the interval that gets constructed, is actually formally correct. We don't do that right now, partially because we kind of have-- know from simulations that the amount of selection bias you incur with these five things is just not huge.
It's not huge in comparison to the variability that you get from the, from just regular human voters. So that's one thing. But then the second thing is the benchmark is live, right? So what ends up happening is it'll be a small magnitude, but even if you suffer from the winner's curse after testing these five models, what will happen is that over time, because we're getting new data, it'll get adjusted down.
So if there's any bias that gets introduced at that stage, in the long run, it actually doesn't matter because asymptotically, basically like in the long run, there's way more fresh data than there is data that was used to compare these five models against-
Got it
... these like private models.
The announcement effect is only just the first phase and the, it has a long tail.
Yeah, that's right, and it, it sort of like automatically corrects itself for this selection adjustment.
Every month I do a little chart of LMSys ELO versus cost just to track the price per dollar, uh, the, the amount of like how much money do I have to pay for one incremental point in ELO. And so I, I actually s- observe an interesting stability in most of the ELO numbers except for some of them.
So for example, GPT-4o August has fallen from twelve ninety to twelve sixty over the past few months, and it's surprising.
You're saying like a new version of GPT-4 versus the, the version in May?
There was May. May is twelve eighty-five. I could have made some data entry error, but it'd be interesting to, to track the, track these things over time anyway. I observe like numbers go up, numbers go down. Like it's, it's remarkably stable.
Gotcha. So you're, these are two different checkpoints and the ELO has fallen.
Yes. And sometimes ELOs rise as well. RekaCore rose from twelve hundred to twelve thirty. That's one of the things, by the way, the community is always suspicious about, like, "Hey, did, did this same endpoint get dumber after release?"
Right? It's, it's such a meme.
That's funny. But those are different endpoints, right?
Yeah, those are different API endpoints, I think.
Yeah. Yeah, yeah.
Or in fact, it's GPT-4o August and May. But if it's for s- like, you know, if endpoint versions we fixed, usually we observe small variation after release.
I mean, you can quantify the variations that you would expect in an ELO. That's a closed-form number that you can calculate. So if the variations are like larger than we would expect, then that indicates that we should look into that.
For sure.
That's an important for us to, to know. So maybe you should, you should send us your plot.
Yeah, please.
I'll send you some data. Yeah.
And I know we only got a few minutes before we wrap, but there are two things I would definitely love to talk about. One is RouteLLM. So talking about models maybe getting dumber over time, blah, blah, blah. Are routers actually helpful in your experience?
RouteLLM34:27
And, uh, Sean pointed out that MoEs are technically routers too. So how do you kind of think about the router being part of the model versus routing different models? And yeah, overall learnings from, from building it.
Yeah. So RouteLLM is a project we released a few months ago, I think. And our goal was to basically understand can we use the preference data we collect to route model based on the question, conditional on the questions, because we may make assumption that some model are good at math, some model are good at coding things like that.
So we found it somewhat useful. For sure, this is like ongoing effort. Our first phase for this project will, you know, is pretty much like open source, the framework that we develop. So for anyone, if they are interested in this problem, they can use the framework and then they can train their own like router model and then to, to do evaluation, to benchmark.
So that, that's like our goal, the reason why we released this framework. And I think there are a couple of future stuff we are thinking. Like one is like, you know, can we just scale this, do even more data, even more preference data, and then train a reward model, train a like, like a router model, better router model.
Another thing is really is a benchmark for this, because right now currently there seems to be... The, one of the endpoint when we developed this project was like there's just no good benchmark for, for router. So that would be another things we think would, could be a useful contribution to community.
And there's still for sure methodology, new methodology we think.
Yeah. I think my, my fundamental philosophical doubt is Does the router model have to be at least as smart as the smartest model? What's the minimum required intelligence of a router model, right? Like if it's too dumb, it's not gonna route properly.
Well, I, I think that you can build a very, very simple router that is very effective. So let me give you an example. You can build a great router with one parameter, and the parameter is just like, I'm gonna check if my question is hard, and if it's hard, then I'm gonna go to the, the big model.
If it's easy, I'm gonna go to the little model. You know, there's various ways of measuring hard that are like pretty trivial, right? Like does it have code? Does it have math? Is it long? That's already a great first step, right?
Because ultimately, at the end of the day, you're competing with a weak baseline, which is any individual model, and you're trying to ask the question, how do I improve cost? And that's like a one-dimensional trade-off. It's like performance cost, and it's great.
Now, w- we can also get into the extension, which is what models are good at what particular types of queries. And then, you know, I think your concern starts taking into effect is can we actually do that? Can we estimate which models are good in which parts of the space in a way that doesn't introduce more variability and more variation and error into our final pipeline than just using the best of them?
That's kinda how I see it.
Your approach is really interesting compared to the commercial approaches where you use information from the Chat Arena to inform your model, which is, uh, I mean, smart, and it's the foundation of everything you do.
Yep.
As we wrap, can we just talk about LMSys and what that's gonna be going forward, like LMArena becoming its own thing. I saw you announced yesterday you're graduating. I think maybe that was confusing since you're PhD students, but this is a different type of, of graduation.
Just for context, LMSys started as like a student club and-
Future38:11
Student-driven.
Yeah.
Yeah.
Student-driven, like research projects of ver- you know, many different research projects are part of LMSys. Sort of ChatBot Arena has, of course, like kind of become its own thing, and Lianmin and Ying, who are, you know-
Created
... sort of created LMSys, have kind of like moved on to working on SGLang, and now they're doing other projects that are sort of originating from LMSys. And for that reason, it w- we thought it made sense to kind of decouple the two, just so, A, the LMSys thing, it's not like when someone says LMSys, they think of ChatBot Arena.
That's not fair, so to speak. Um-
And we wanna support new projects.
And we wanna support new projects-
Yeah
... so on and so forth. But of course, these are all like, you know, our friends, so that's why we call it a graduation.
I agree. That's like one thing that people were maybe a little confused by, where LMSys kinda starts and ends and where Arena starts and ends. So-
Yeah. Yeah
... um, I think you reach escape velocity now that you're kinda like your own, your own thing, so.
I have one parting question. Like what, what do you want more of? Like what do you want people to approach you with?
Oh my God, we need so much help. One thing would be like we're, we're obviously expanding into like other kinds of arenas, right? We definitely need like active help on red teaming. We definitely need active help on our-
Different modalities
... different modalities, vision.
Yeah.
Co-pilot. It's all-
Yeah. Coding.
Coding. You know, if somebody could like help us implement this like REPL in, REPL in ChatBot Arena, massive, right? That would be a massive delta. And I know that there's people out there who are passionate and capable of doing it.
It's just we don't have enough hands on deck. We're just like an academic research lab, right? We're not equipped to support this kind of project. So yeah, we need help with that. Um, we also need just like general backend dev.
Yes, for sure.
Um, and new ideas, new conceptual ideas. I mean, honestly, the work that we do spans everything from like foundational statistics, like new proofs, to full stack dev. And like anybody who's like wants to contribute something to that pipeline is should definitely reach out-
Yeah
... 'cause we need, we need it.
And it's an open source project anyways.
Yeah.
Anyone can make a PR.
And we're happy to, you know, whoever wants to contribute, we'll give them credit, you know?
Yeah.
We're not trying to keep all the credit for ourselves. We want it to be a community project.
That's great and fits the spirit of everything you've been doing over there. So awesome, guys. Well, thank you so much for, for taking the time, and we'll put all the links in the show notes so that people can find you and, and reach out if they need it.
Thank you so much.
It's very nice to talk to you, and thank you for the wonderful questions.
Thank you so much.






