Intro0:00
All right. All right. Take it away. Cool. Um, yeah, thanks for, for handing me over. Um, I'm Luca. I'm a research scientist at the Allen Institute for AI. Um, I threw together a few slides on sort of like a recap of, like, interesting themes in open models for, for twenty twenty-four.
Um, have about maybe twenty, twenty-five minutes of slides, and then we can chat if there are any questions. If I can advance to the next slide. Okay, cool. Um, so, um, I did the quick check of like to sort of get a sense of like how much twenty twenty-four was different from twenty twenty-three.
Um, so I went on Hugging Face and sort of get-- tried to get a picture of what kind of models were released in twenty twenty-three and, like, what do we get in twenty twenty-four. Um, twenty twenty-three we get-- we got things like, uh, both Llama one and two.
We got Mistral, we got MPT, Falcon models. I think the Yi model came at the tail end of the year. It was a pretty good year. Um, but then I did the same for twenty twenty-four, um, and it's actually quite stark difference.
Um, you have, uh, models that are, you know, reveling frontier-level performance of what you can get from closed models from, like, Qwen, from Deepseek. We got Llama 3, we got all sorts of different models. Um, I added our own, uh, OLMo at the bottom.
2024 Surge1:15
Uh, there's this, uh, growing group of like fully open models that I'm gonna touch on a little bit later. Um, but, you know, just looking at the slides, it feels like twenty twenty-four was just smooth sailing, happiness, much better than previous year.
Um, and you know, you can plot, um, you can pick your favorite benchmark, um, or least favorite, I don't know, depending on what point you're trying to make, um, and plot, you know, your closed model, your open model, um, and sort of spin it in ways that show that, oh, you know, um, open models are much closer to where closed models are today versus to-- versus last year, where the gap was fairly significant.
Um, so one thing that, um, I think, I don't know if I have to convince people in this room, but usually when I give these talks about like open models, there is always like this background question in, in, in people's mind of like, why should we use open models?
Why Open?2:21
Um, is it just use model APIs argument? You know, it's, it's just an HTTP request to get output from a-- from one of the best model out there. Why do I have to set up infra and use local models?
Um, and there are really like two answer. Um, there is the more research-y answer for this, which is where my background lays, which is, um, just research. If you wanna do research on language models, research thrives on, on open models.
There is a large swath of research on modeling, on how these models behave, on evaluation, on inference, on, uh, mechanistic interpretability that could not happen at all if you didn't have open models. Um, there are also, um, for AI builders, there are also like good use cases for using, um, local models.
Um, you know, you have some... This is like a very not, uh, comprehensive slides, but you have things like, uh, there are some applications where local models just blow closed models out of the water. Um, so like retrieval is a very clear example.
Um, you might have like constraints like edge AI applications where it makes sense. Uh, but even just like in terms of like stability, being able to say this model is not changing under the hood, um, it's-- there's plenty of good cases for, for, um, open models.
Um, and the community's just not models. Um, is I stole this slide from, uh, one of the Qwen two announcement blog post. Uh, but it's super cool to see like how much, um, tech exists around, um, open models, on serving them, on making them efficient and hosting them.
It's pretty cool. Um, and um, it's, um, if you think about like where the term opens come from, it comes from like the open source. Um, really open models meet the core tenets of, of, um, open, of open source, uh, specifically when it comes around collaboration.
There is truly a spirit like through these open models, you can build on top of other people innovation. Um, we see a lot of this even in our own work of like, you know, as we iterate in the various version of OLMo, um, it's not just like every time we collect from scratch all the data.
No, the, the first step is like, okay, what are the cool data sources, uh, and datasets people have put together for language model pre-training? Um, or when it comes to like, um, our post-training pipeline, um, we s-- uh, one of, uh, the steps is, um, you wanna do some DPO, and you use a lot of, uh, outputs of other models, uh, to improve your, your preference model.
OSI License5:32
So it's really, um, having like an open sort of ecosystem benefits and accelerates the development of open models. Um, one thing that, um, we got in twenty twenty-four, which is not a specific model, but I thought it was really significant, is we first got, uh, we got our first open source AI definition.
Um, so this is from the Open Source Initiative. Um, they've been generally the steward of a lot of the open source licenses when it comes to software. Um, and so they embarked on this journey in trying to figure out, okay- How does a license, an open source license for a model look like?
Um, majority of the work is very dry because licenses are dry, so I'm not gonna walk through the license step-by-step. But, um, I'm just gonna pick out, uh, one aspect that is very good, uh, and then one aspect that personally feels like it needs improvement.
On the good side, um, this, um, this open source AI license, actually this is very intuitive. If you ever build open source software and you have some expectation around like what open source, uh, looks like for software, uh, for, for AI sort of matches your intuition.
So the weights need to be freely available, uh, the code must be released with an open source license, uh, and there shouldn't be like license clauses that, uh, block specific use cases. So under this definition, for example, Llama or some of the Qwen models are not open source because the license says you can't, you can't use this, this model for this, or it says if you use this model you have to name the output this way or derivative needs to be, uh, named that way.
Those clauses don't meet open source definition, um, and so they will not be cover... The, the Llama license will not be cover under the open source definition.
Um, it's not perfect. Um, one of the thing that, um, um, internally, you know, in discussion with, uh, with OSI, we were sort of disappointed is around, um, the language for data. Um, so you might imagine that an open source AI model means a model where the data is freely available.
Uh, there were discussion around that, but at the end of the day they decided to go with a softened stance where, um, they say, um, a model is open source if you provide sufficient detailed information on how to sort of replicate the data pipeline so you have an equivalent system.
Sufficient, sufficiently detailed, uh, it's very, it's very fuzzy. Don't like that. Uh, an equivalent system is also very fuzzy. Um, and this doesn't take into account the accessibility of the process, right? It might be that you provide enough information, but this process costs, I don't know, ten million dollars to do.
Um, now the open source definition, like any open source license, has never been about accessibility, so that's never factor in open source software how accessible software is. Um, I can make a piece of software open source, put it on my hard drive and never access it.
That software is still open source. The fact that it's not widely distributed doesn't change the license. But practically there are expectation of like what we want good open sources to be, so it's kind of sad to see that, um, the, the data component, uh, in this license is not as, as open as some of us would like, uh, would like it to be.
And a link to blog post that Nathan wrote on the topic that it's less rambly, uh, and easier to follow through. Um, one thing that in general I think it's fair to say about the state of open models in twenty twenty-four is that we know a lot more than what we knew in, in twenty twenty-three.
Um, like, um, both on the training data, like the pre-training data, how you curate, um, on like how to do like all the post-training, especially like on the RL side. Um, you know, twenty twenty-three was a lot of like throwing random darts at the board.
Uh, I think twenty twenty-four we have clear recipes that, okay, don't get the same results as a closed lab because there is a cost in, in actually matching what they do. Um, but at least we have a good sense of like, okay, this is, this is the path to get state-of-the-art language model.
Um, I think that a... one thing that is a downside of twenty twenty-uh, four is that I think we are more resource constrained than twenty twenty-three. It feels that, you know, the barrier for compute that you need to, to move innovation along is just being ri- uh, rising and rising.
Um, so like if you go back to this slide, there is now this, this cluster of models that are sort of released by the compute rich club. Um, membership is hotly debated. Um, you know, some people don't wanna be called rich because it comes with expectation.
Some people want to be called rich, but, hmm, I don't know. There is debate. Like, these are players that have, you know, ten thousand, fifty thousand GPUs at minimum. Um, and so they can do a lot of work, um, and a lot of exploration in improving models that is not very accessible.
Compute Gap11:11
Um, to give you a sense of like how I personally think about research budget, um, for each part of the, of the language model pipeline is like on the pre-training side you can maybe do something with a thousand GPUs.
Really you want ten thousand. And like if you want real estate of the art, you know, your Deepseek at minimum is like fifty thousand. Um, and you can scale to infinity. The more you have, the better it gets.
Um, everyone on that side still complains that they don't have enough GPUs. Uh, post-training is a super wide, um, sorta, uh, spectrum. You can do as little with like eight GPUs, um, as long as you're able to, um, run, you know, a, uh, a good version of, say, a Llama model.
You can do a lot of work there. Um, you can scale a lot of the methodology just like scales with compute, right? If you're interested in, um, you know, your Open replication of what, uh, OpenAI's o1 is, um, you're gonna be on the 10K spectrum of OGpus.
Um, inference, you can do a lot with very few resources. Evaluation, you can do a lot with, well, I should say at least one GPUs if you wanna evaluate, um, uh, open models. Uh, but, um, in general, like if you are-- if you care a lot about intervention to do on this model, which is my, uh, preferred area of, of research, then, you know, the resources that you need, um, are quite, quite significant.
Fully Open12:48
Um, one of the trends, um, that has emerged in twenty-twenty-four is this cluster of, um, fully open models. Um, so OLMo, uh, the model that we built at AI2 being one of them. Um, and, you know, it's nice that it's not just us.
There's like a cluster of other mostly research, um, efforts who are working on this. Um, and so, uh, it's good to, um, to give you a primer of what like fully open means. Um, so fully open, the easy way to think about it is instead of just releasing a model checkpoint that you run, you release a full recipe so that, um, other people working on it, uh, working on that space can pick and choose whatever they want from your recipe and create their own model or improve on top of your model.
Um, you're giving out the full pipeline and all the details there, um, instead of just like the end output. Um, so I pulled up this screenshot from our recent, um, MoE model. Um, and like for this model, for example, we released the model itself, the data that was trained on, the code f-both for training and inference, um, all the logs that we got through, um, the training run, as well as, um, every intermediate checkpoint.
Um, and like despite the fact that you release different part of the pipeline, uh, allows others to do really cool things. Um, so for example, this tweet from early this year from, uh, folks at News Research, um, they use our pre-training data, uh, to do a replication of the BitNet paper in the open.
Um, so they took just a really, like the initial part of a pipeline, um, and then did their thing on top of it. Um, it goes both ways. So for example, for the OLMo 2 model, um, a lot of our pre-trained data for the first stage of pre-training, um, was from this DCLM, uh, initiative, uh, that was led by folks, uh, ooh, a variety of inst- a variety of institution.
It was a really nice group effort. But, um, you know, for when it was nice to be able to say, "Okay, you know, the state-of-the-art in terms of like what is done in the open has improved. We don't have to like do all this work from scratch to catch up the state-of-the-art.
We can just take it directly and integrate it and do our own improvements on top of that." Um, I'm gonna spend a few minutes, uh, talk-- doing like a shameless plug for, uh, some of our fully open recipes.
Um, so indulge me in this. Um, so a few things that we released this year was, as I was mentioning, this OLMoE model, um, which is, I think still is state-of-the-art, um, MoE model in its size class. Um, and it's also fully open, so every component of, of this model are variable.
AI2 Models15:33
Um, we released a multimodal model called Molmo. Um, Molmo is not just a model, but it's a full recipe of how you go from a text-only model to a multimodal model, and we apply this recipe on top of Qwen checkpoints, on top of OLMo checkpoints, as well as on top of OLMoE.
Um, and I think there've been replication doing that on top of Mistral as well. Um, um, on, on the post-training side, uh, we recently released Tulu 3. Um, same story. This is, uh, a recipe on how you go from a base model to, um, a state-of-the-art post-training model.
Uh, we use the Tulu recipe on top of OLMo, on top of Llama, and then there's been, um, open replication effort to do that on top of Qwen as well. Uh, it's really nice to see like, you know, when your recipe sorta it's kinda turnkey.
You can apply it to different models, and it kinda just works. Um, and finally, the last thing we released this year was OLMo 2, um, which so far is the best state-of-the-art, uh, fully open language model. Um, it sorta combines aspect from all three of these previous models, um, what we learn on the data side from OLMoE and what we learn on like making models that are easy to adapt from the Molmo project and the Tulu project.
Um, I will close with a little bit of reflection, like ways this, this ecosystem of open models, um, like it is not all roses. It's not all happy. Uh, it feels like day to day it's always in peril.
Threats17:14
Um, and, you know, I talked a little bit about like the compute issues that come with it, uh, but it's really not just compute. Um, one thing that is on top of my mind is due to like the environment and how, um, you know, growing feelings about like how AI is, is treated, it's actually harder to get access to a lot of the data that was used to train a lot of the models up to last year.
So this is a screenshot from really fabulous work from Shane Longpre, um, who's I think is at NeurIPS, um, about, um, just access of, uh, like diminishing access to data for language model pre-training. So what they did is they, um, went through every snapshot of Common Crawl Common Crawl is this publicly available scrape of the, of a subset of the internet, and they looked at how, um, for any given website, uh, whether a website that was accessible in, say, twenty seventeen, whe-whether it was accessible or not in twenty twenty-four.
And what they found is, as a reaction to, like, the closed, uh, like, of the existence of closed models like OpenAI or Claude, um, GPT or Claude, uh, a lot of content owners have blanket blocked any type of crawling to their website.
Um, and this is something that we see also internally at AI2. Um, like, one project that we started this year is, um, we wanted to, we want to understand, like, if you're a good, uh, citizen on the internet and you crawl, uh, following sort of norms and policy that have been established in the last twenty-five years, what can you crawl?
And we found that there's a lot of website where, um, the norms of how you express preference of whether to crawl your data or not are broken. A lot of people would block a lot of crawling but do not advertise that in robots.txt.
You can only tell that they're crawling, that they're blocking you in crawling when you try doing it. Sometimes you can't even crawl the robots.txt to, to check whether you're allowed or not. And then a lot of, um, websites, um, there's sort of this, like, all these technologies that historically have been, have existed to make web servers serving easier, um, such as, um, Cloudflare or a DNS.
Uh, they're now being repurposed for, um, blocking AI or any type of crawling in a way that is very opaque to the content owners themselves. Um, so, you know, you go to these websites, you try to access them, and they're not available.
Uh, you get a feeling it's like, oh, someone changed, something changed on the, on the DNS side that it's blocking this, and likely the content owner has no idea. They're just using, uh, Cloudflare for better, you know, load balancing, and this is something that was sort of sprung on them, uh, with very little notice.
Um, and I think the problem is this, this, um, blocking or ideas really, it, it impacts people in different ways. Um, it disproportionately helps, um, companies that have a head start, which are, um, usually the closed labs, and it hurts, uh, incoming, uh, newcomer players, um, where you either have now to do things in a sketchy way, um, or you're never gonna get that content, uh, that, uh, the closed lab might have.
So there's a lot, there's a lot of coverage. I'm gonna plug Nathan's blog post again, uh, that is, uh, that, um, uh, I think the title of this one is very succinct, uh, which is, like, we're actually not...
You know, before thinking about running out of training data, we're actually at, running out of open training data. And so if we want better open models, um, this should be on top of our mind. Um, the other thing that, uh, has emerged is that there's strong lobbying efforts on trying to define any kind of, uh, open source AI as, like, a new, um, extremely risky danger.
Um, and I wanna be precise here. Like, the problem is not, um, um, like, the problem is not not considering the risk of this technology. Every technology has risks that, that should always be considered. The thing that it's like to me is, um, sorry, disingenuous is, like, just putting this AI on a pedestal, um, and calling it, like, an unknown alien technology that has, like, new and undiscovered potentials to destroy, um, humanity, when in reality, all the dangers, I think, are rooted in dangers that we know from existing software inf- industry, um, or existing issues that come when, when using software on, um, on a lot of sensitive domains like medical, um, areas.
Um, and I also noticed a lot of efforts that have actually been going on on trying to make these mo- open models safe. Um, I pasted one here, uh, from AI2, but there's actually, like, a lot of work that has been going on on, like, okay, how do you make...
If you're distributing this model openly, how do you make it safe? Um, how-- What's the right balance between accessibility around open models and safety? Um, and then also there's annoying, uh, brushing of, um, sort of unf- concerns that are then proved to be unfounded under the rug.
You know, if you remember at the beginning of this year, it was all about bio risk of these open models. Uh, the whole thing fizzled out because there's been, finally, there's been, like, rigorous research, not just this paper from Cohere folks, but there's been rigorous resea-research showing that this is really not a concern that you, we should be worried about.
Again, there is a lot of dangerous use of AI application, but this one was just, like, a lobbying ploy to just make things sound scarier, uh, than they actually are. So I gotta preface this part and say this, this is my personal opinion.
It's not my employer. But I look at things like, uh, the SB ten forty-seven from, from California, and I think we kinda dodged a bullet, bullet on, on this leg-legislation. We, you know, the open source community, a lot of the community came together at the last, sort of the last minute, um, and did a g- very good effort trying to explain all the negative impact of this bill.
Um, but, um, there's like, I feel like there's a lot of excitement on building these open models, uh, or, like, researching on these open models, and lobbying is not sexy. Uh, it's kind of boring, uh, but, um, it's sort of necessary to make sure that this ecosystem can, can really thrive.
Um, this end of presentation, I have some links, emails, sort of standard thing in case anybody wants to reach out. And if folks have questions, um, or anything they wanted to discuss, it's our open floor.
Incentives24:50
I'm very curious how we should build incentives to build open models, things like Francois Chollet's, uh, ARC Prize and other initiatives like that. What is your opinion on how we should better align incentives in the community so that open models stay open?
The incentive bit is, like, really hard. Um, like, even... It's something that I, actually, even we think a lot about it internally, um, because, like, building open models is risky. It's very expensive. Um, and so people don't wanna take risky bets.
Um, I think the, um, definitely, like, the challenges, um, like our challenge, I think those are, like, very valid approaches for it. Um, and then I think in general promoting building, so, um, any kind of effort to participate in this challenge, in those challenges, if we can promoting doing that on top of open models, um, and sort of really lean into, like, this multiplier effect, um, I think that is a good way to go.
Um, if there were more money for, um, efforts, um, like research efforts around open models, there's a lot of... I think there's a lot of investments in companies that, um, at the moment are releasing their model in the open, which is really cool.
Um, but, um, it's usually more because of commercial interest and not wanting to support, um, this, this, like, open models in the long term. It's a really hard problem because I think everyone is operating sort of in what...
Everyone is at their local maximum, right? In ways that really optimize their position on the market. Um, the global maximum is harder to achieve.
Mistral Recap26:38
Um, yeah, I'm super excited to be here to talk to you guys, uh, about Mistral. Uh, a really short and quick recap of what we have done, what kind of models and products we have released in the past a year and a half.
So, um, most of y- you have already known that we are a small startup, uh, founded about a year and a half ago in Paris, um, in May 2023. It was founded by three of our co-founders, and in September 2023, we released our first open source mod- model, Mistral 7B.
Um, yeah, how, how many of you have used or heard about Mistral 7B? Yay. Pretty much everyone. Thank you. Uh, yeah, it's our, uh, pretty popular and, uh, community re- our community really love this model. And in December 2023, we, we released another popular, uh, model with the MoE architecture, um, Mistral 8x7B.
And ooh, going into this year, you can see we have released a lot of things this year. Um, first of all, in February 2024, we released, uh, Mistral Small, Mistral Large, uh, Le Chat, which is our chat interface.
I will show you in a little bit. Uh, we, we released a embedding model, uh, for, you know, converting your i- text into embedding vectors. And all of our models are available, um, the, the big cloud resources, so you can use our model on Google Cloud, AWS, Azure, uh, Snowflake, IBM.
So very useful for enterprise who wants to use our model through cloud. And in April and May this year, we released another powerful open source, um, MoE model, AX22B, and we also released our first code model, CodeStral, which is amazing at 80-plus languages.
Um, and then we provided another fine-tuning service for customization. Uh, so because we know the community love to fine-tune our models, so we provide you a very nice and easy option for you to fine-tune our model, uh, on our platform.
And also we released our, uh, fine-tuning code base called Mistral Fine-tune. It's open source, so feel free to take it a... uh, take a look. And more models. From July to November this year, we r- we released, um, many, many other models.
Uh, first of all is the two new small, um, best small models. We have Mistral 3B, uh, great for deploying on edge devices. Um, we have Mistral 8B. Uh, if you used to use Mistral 7B, Mist- Mistral 8B is a great re- replacement with much stronger performance than Mistral 7B.
Uh, we also collaborated with NVIDIA and open sourced another model, Nemo 12B, uh, another great model. And just a few weeks ago, we updated Mistral Large with the version two, with the updat- updated, uh, state-of-our features, um, and really great, uh, function calling capabilities.
It's, uh, supporting function calling natively. And we released two multimodal models. Um, Pixtral 12B, it says open source, and Pixtral Large. Uh, just amazing model f- models for not understanding, uh, images, but also great at text understanding. So yeah, a lot of the image models are not so good at text understanding, but Pix- Pixtral Large and Pixtral 12B are good at both image understanding and text understanding.
And of course, we have models for research. Uh, CodeStral Mamba is built on, uh, Mamba architecture, and Mathtrall, great with working with math, math problems. So yeah, that's a lot of models. Uh...
Uh, here's another view of our model offerings. Uh, we have several premier models, which means, uh, these models are mostly, um, available through our API. I mean, all of the, the models are available throughout our API, uh, except for Mistral 7...
uh, 3B. Um, but the, for the premium model, they have a special license, um, Mistral, uh, research license. You can use it for free for exploration, but if you want to use it for enterprise, um, for production use, you will need to purchase a license from us.
Uh, so on the top row here, we have Mistral 3B and 8B as our premier model. Um, Mistral Small for, uh, best, best low latency use cases. Mistral Large is great for your most sophisticated use cases. Um, Pixel Large is the frontier class multimodal model.
And, and we have CodeStral for... great for coding. And then again, Mistral Embedding model. And, um, the bot- the bottom of the slides here, we have several Apache 2.0 licensed open weight models, um, free for the community to use.
And also if you want to fine-tune it, use it for cus- uh, customization, production, feel free to do so. Uh, the latest we have PixStral 3... uh, 12B. Um, we also have, um, Mistral Nemo, um, Mum- CodeStral Mamba, and Mistral, as I real- as, as I, uh, mentioned.
And we have three legacy models that we don't update anymore, so we recommend you to, uh, move to our newer models, um, if you are, uh, still using them. Uh, and then just a few m- weeks ago, we did a lot of, um, uh, improvements to our, uh, code interface, Le Chat.
How many of you have used Le Chat? Oh, no. Only a few. Okay. I highly re- recommend Le Chat. It's chat.mistral.ai. Uh, it's free to use. It has all the amazing capabilities I'm gonna show you right now. Uh, but before that, Le Chat in French means, uh, cat, uh, so this is actually a cat logo.
Yeah, uh, if you can tell, this is cat eyes. Uh, yeah. So first of all, I want to show you something, uh... maybe let's ta-- let's take a look at, uh, image understanding.
Le Chat Demo33:01
So here I have a receipts, and I want to ask... Sorry, just going to get the prompts.
Going back. What is going
on? Yeah, I had an issue with Wi-Fi here, so hopefully it would work. Cool. So basically, uh, I have a receipt, and I said I ordered a, I don't know, coffee and a sausage. How much do I owe at a 18% tip?
Uh, so hopefully it was able to get the cost of the coffee and the sausage and ignore the other things. And, uh, yeah, I don't real-ly understand this, but I think this is coffee. Uh, it's, yeah, nine. Yeah.
And then cost of the sausage, we have 22 here. Yep. And then it was able to add the cost, calculate the tip, and all that. Uh, great. So it's great at image understanding. It's great at, uh, OCR tasks.
So if you have OCR tasks, please use it. It's free on Le Chat, uh, it's also available through our API. Um, and also want to show you a canvas example. Uh, a lot of you may have used canvas with other tools before, but, uh, with Le Chat it's completely free.
Again, here I'm asking it to create a canvas that's used PyScript to execute Python in my browser.
So... Ooh, what's going on?
Okay. Let's see if it works. Import this. Oh.
Yep. Okay. So yeah, so basically it's executing Python, uh, here, exactly what we wanted. Uh, and the other day, I was trying to ask Le Chat to create a game for me. Let's see if we can make it work.
Yeah, the Tetris game. Uh, yep.
Let's just get one row, maybe. Um...
Ah! Oh, no.
Okay, never mind. You get the idea. I failed my mission. Um...
Okay, here we go. Hey. Yay. Uh, cool. Yeah. So, uh, as you can see, Le Chat can write, like, a code about a simple game pretty easily, and you can ask Le Chat to explain the code, make updates, however you like.
Um, another example. Uh, there is a bar here I want to move. Okay. Great. Okay. And, uh, let's go back.
Another one. Uh, yeah, we also have web search capabilities. Like, you can ask what's the latest AI news. Uh, image generation is pretty cool. Generate an image about researchers in Vancouver.
Uh, yeah, it's Black Forest Labs, uh, Flux Pro. Uh, again, this is free, so...
Oh, cool. I guess researchers here are mostly from University of British Columbia. Uh, that's smart. Uh, yeah, so this is Le Chat. I re- Please feel free to use it, uh, and let me know if you have any feedback.
We're always looking for improvement, and we're gonna release a lot more powerful features in the coming years. Thank you.





