Intro0:00
When customers use this, uh, off-the-shelf closed model, what's very sad is that they are not leveraging, you know, this data that they have been collecting for, for years or sometimes for decades. Uh, so much data, sometimes it's trillions of tokens of data, uh, in a very specific domain, their domain, which is data that you will not find on, in the, in the public, uh, on the public internet.
So data on which like the closed model will actually not have access. So if they are using like closed source models, they are basically not benefiting from all these insights, all this data they have collected through years.
Voxtral TTS0:22
Okay. Welcome to Lean Space. We're here in the studio with our trusty co-host, uh, Vibhu. Welcome.
Very excited for this one.
As well as Guillaume and Pavan from Mistral. Welcome.
Excited to be here. Thank you for having us.
Uh, Pavan, you are leading audio research at, uh, Mistral, and, uh, Guillaume, you're our chief scientist. Uh, what are we announcing today? We're, we're sort of coordinating this, this release, uh, with you guys.
Yeah. So we are releasing, uh, Voxtral TTS. So it's our first, uh, audio model that generates speech. Uh, it's not our first audio model. We had a, a couple of releases before. We had one, uh, in the summer that was Voxtral, our first audio model, but it's-- it was like a, a transcription model, ASR.
Uh, like a few months later, we released some update on top of this, supporting more languages. Also a lot of table stake features for our customers like, uh, context biasing, uh, derization, uh, timestamping on the, on the transcription.
We also have some real-time model that can transcribe not just at the end of the... You just don't, don't need to fill your entire audio file, but that can transcribe in real-time. Uh, on here, this is kind of the natural extension, uh, in the audio, so basically speech generation.
So, so yeah. So we support, uh, nine languages, uh, and this is a pretty small model, a 3D model, so very fast. Very fast and also status, yeah, really performance is the same level of the, of the base model, but it's, uh, much more efficient in terms of cost and also, uh, much in terms of cost, it's also much, uh, only a fraction of the cost of our competitors.
And we are also releasing the weight of this model that are on inference. Yeah.
Yeah. May-May linked?
Not this time.
Well, yeah. What's the decision factor?
Um. It's a good question. There'll be more, Vibhu.
Ooh. Yeah, Pavan, any other sort of research notes to add on what-
No, we, uh, maybe we'll dive into it later in the forecast too, but it's a novel architecture that we developed in-house. Uh, we iterated it, uh, on several internal architectures and ended up with a autoregressive flow matching architecture, uh, and also have a, a new, uh, in-house, uh, neural audio codec, uh, which, uh, converts this, uh, audio into 1.5 hertz, uh, latent tokens, semantic and acoustic tokens.
And yeah, that's, uh, that's the new part about this model, and we're pretty excited that it, uh, uh, it came out, uh, with such good quality. And like Guillaume was mentioning, yeah, it's a 3B model. Uh, it's based off of the Mistral model that we actually released just a few months back and in trunk.
And it mainly meant for like the TTS stuff, but the need text capabilities are also there, uh, in the model.
Yeah.
Yeah.
So there's a lot to cover. I always-- I, I love any- anything to do with novel encodings and all, all those things, uh, because I think that's obviously increases a lot of efficiency, but also maybe bugs, uh, sometimes happen.
You were previously at Gemini, and you worked on post-training for language models, and maybe like a lot of people will have less experience with audio models just in general, uh, compared to, uh, pure language. Uh, what did you find that you have to sort of revisit from scratch as you joined Mistral and started doing this?
At least when it comes to for-- I think there are, there are two buckets, I guess, the audio understanding and audio generation.
Mm-hmm.
The audio understanding, like the Voxtral models that Guillaume was mentioning that we released earlier, the Voxtral chat that we released, uh, uh, I think July last year, and the follow-up transcription-only, uh, models family that we released in January.
That would be one bucket, I guess, and the generation is another bucket. I think, uh, you can also treat them as a unified, uh, set of models, but currently the approaches are a little different between these two to your question on, um, how audio is like fed to the model.
In the understanding model, it's very similar to actually Pixtral model that we also released, I guess, uh-
Yes, that was amazing
... a couple of years back. Yeah. It was pretty-- I-- That was the first project I worked on after joining Mistral. It was pretty, pretty nice. And Voxtral was, uh, very similar in spirit, I guess. So we feed, uh, audio through an audio encoder similar to, uh, images through a vision encoder, and it produces continuous, uh, embeddings, uh, and which are fed as tokens to the main transformer, decoder transformer model.
Yeah. And the model output is just text. So on the output side, there is nothing, uh, that needs to be done in these kinds of models. I guess the interesting part about the generation step is the output now has to produce audio, and the approach that we have is this, uh, neural audio codec, which, uh, converts audio into these, uh, latent tokens.
There is, uh, a lot of, uh, existing literature and a lot of models, uh, which, uh, are based off of this kind of approach. And we took a slightly, uh, different, I guess, design decisions around this. But at the end of the day, the neural audio product converts audio into a 12.5 hertz, uh, set of latents, and each latent is-- has a semantic token and, uh, a set of acoustic tokens.
And the idea is that you take these discrete tokens and then feed it. On the input side, there are several ways to fuse this at each frame, but we just sum the embeddings. So it's, it's kind of like having key different vocabularies and kind of combine all of them because they all correspond to one audio frame on the input side.
The output side is the interesting part. On, on the output side, the-- It's not the-- I don't know if it's the most popular, but one popular technique is to have a depth transformer because you have key tokens at each time step.
Like with our text, you just have one token at each time step, so you just do Like, uh, predict the token from the vocabulary with-- yeah, with just you get probabilities-
This is a very straightforward-
Very straightforward.
Yeah.
But if you have K tokens, then the main thing would be to predict all of them in parallel, but that doesn't work. Uh, at least that doesn't work that well because audio has more entropy. And, uh, the... one of the techniques people use is this depth transformer, where you, uh, you almost have a small transformer or it can be LSTM-RN as well, but people use transformers, and you predict the K tokens in autoregressive fashion in that.
So you have two autoregressive things going on. So the thing we did differently is instead of having this autoregressive K step prediction, we have a flow matching model. Instead of modeling this as a discrete token set, we, we train the codec to be both discrete and continuous to have this flexibility.
So we did try the discrete stuff too, and which it works well, but the continuous stuff works just better. So yeah, we took this flow matching. So the-- it's a flow matching head, which takes the latent from the main transformer and kind of like k- in diffusion it's denoising, but in this flow matching it's a velocity estimate.
So you kind of, uh, uh, go from this noised, uh, latent all the way to the audio latent, which corresponds to the eighty-millisecond audio and then which is sent, uh, through the vocoder to get back the eighty-millisecond audio frame.
Real-Time Voice7:25
Yeah. Is this the first application of flow matching in audio? Because-
Uh-
... usually I come, I come across this in the image.
Yeah. Actually, in some sense, there are models, uh, flow matching models in audio, but I think this, this specific combination-
Mm.
I, I, I could be wrong. There could be some work.
No, no, no.
I haven't seen, I haven't seen much work in this, so I, I think it's, it's novel. And a lot of, uh, it's just a way bigger community. It's they-- I think they pioneer a lot of these diffusion flow matching work, and it's interesting to adopt some of the ideas there into audio and-
Yeah.
Yeah. And personally, that's the exciting part, which is trying, trying out. One-- and more meta point is unlike text, even in vision, I think this is true, but in audio it's definitely true, is that there is no, uh, winner model yet.
There is no, "Okay, this is the way you do things." It's, uh, it's still evolving. I think people are still iterating and figuring out, like, what's the, uh, best overall recipe, I guess. The, the idea-- I-I mean, pretty sure, uh, there are models which are also completely end-to-end, like NATO audio and NATO audio, but it's still not like, uh, come to a convergence point where this is the right way to think.
That, that also makes the space pretty exciting to explore.
What are some of the ways to look at it? Like, there are ways where you can do diffusion for audio generation, but if you want, like, real-time generation, that's a big thing with the approach I'm assuming that you took.
Yeah.
Uh, and also, like, how do you go about evaluating different axes of what you care about, right? So yeah.
Good point. I think we, uh-- So you can do just flow matching diffusion for the whole audio. We didn't even go down that path because one of the main applications is, uh, uh, voice agents, and we want real-time streaming, and that's the use case.
That's not the only use case, but that's one of the primary use cases we want to get to. So we picked the autoregressive, uh, approach for that. And within the autoregressive space, again, you can do chunk by chunk or you can do...
Uh, so we picked the-- I think at least personally, I prefer the approaches which are the simplest, I guess. And so we try to see, can we just add audio as just another head to our regular de-- uh, transformer decoder model?
Because that kind of makes it easier for eventual end-to-end modeling of audio text native modeling. Yeah, and it works pretty well. So I guess we kind of, uh, went with that. And we tried it a little bit with the flow matching head itself.
Like, we had a discrete diffusion kind of approach, which also works well, but the flow matching worked better.
I was just curious about how you also think about, uh, this overall direction of research. Like, do you-- Basically, when you work with the audio team, do you set, uh, some high-level parameters and then let them explore whatever?
Or how does it work between you guys?
No, I, I think the, the way it works is that we have a-- we are prioritizing together, I think, what are the most important features. There are many, many things we can do in audio.
Yeah.
I think we try to decide, like, how we should do things. For instance, ultimately, what we want to do is to build this full duplex model, but we are not going to start this-- start there directly. I think it's some of my project people are doing, but-
Just to confirm, full duplex means it can speak while I'm speaking or-
Yeah.
Okay.
Yeah, audio in, audio out.
Yeah, yeah.
So ultimately, we are going to get there. But for us, it was-- I mean, we decided to take it like step-by-step. So we start with whatever is the most important, I think also for our customers, which is a transcription, is the most popular use case.
Then the speech generation, the real-time just a bit before that. And then the next thing is going to be, like, more like try combine everything all together. But, uh, but yeah, we thought it was also important to, like, separate things and, like, optimize each capability one by one before we-
Mm
... merge all of that together.
Then the, the, the super omni model, um-
But very interesting because as Pavan said, it's, uh, uh, when you work on some other domains of, uh, just LM and other things, there are many areas where I think it's not as interesting. For instance, many places it's essentially just around data or, like, creating new environments and a lot of kind of easy things.
But, um, things where I think the research is maybe not as interesting where in audio, there are so many ways to, to actually build these models, so many ways to go a-around this. That's-- This sense I think is really interesting.
Uh, and what we also tried for speech generation is that we tried multiple approaches. What was interesting is that even though they were extremely different, they ended up big at the end of the day, like, particulars. Uh, but the flow matching turned out to be quite more natural, so we are happy with this, uh-
Is there an intuition why? Uh, maybe like flow matching is just models speech better in some natural, fundamental latent dimension?
No, I think the main, main thing is, like, uh, e-even at a particular time step, there is a, a distribution of things-
Yes
... uh, to be predicted, like the way you inflect So you already know the word that you're speaking, and you have the in text space, let's say the word maps to just a single token for simplicity. In most cases, it does.
So there is, uh, not a lot of like... So you just pick the word. But within, within audio, even the same word could, even with your own voice, could be inflected in so many different ways. And I think, uh, a-any approach which like models this distribution and well, and flow matching is one, one of the tech- it, it's not the only one at all, but it's, it's, uh, one which works pretty reasonably well, I think that's better.
So you, you have to pick across several different like, uh... The, the intuition I have is it's, it's there are some several different clusters, each corresponding to some specific way you would inflect, pronounce that thing, and you can't predict the mean of it because that corresponds to some blurred out speech or something like that.
But you have to, uh, pick one and then like sharp-
Conditional inference.
Yeah, exactly.
Is that all covered under disfluencies? Which is I think the, the normal term of art.
Uh-
Disfluencies, pauses, intonations. By the way, I, I we have to thank Sophia for setting all this up, including like some of these really good notes because-
Yeah
... I'm less familiar with the audio domain.
No, no, no. I think disfluencies are definitely one such phenomenon. Uh, disfluencies is more like, uh-
Which is ums, uhs.
Yeah. Ums, uhs, and also repeat like you feel like, like, uh-
Yeah, yeah
... you do these filler words. You're thinking, so you repeat the word.
Okay.
And-
Whereas intonation is like a diff- it's up, uh, up speak and all this.
Yeah.
Okay.
And yeah, so I think there is a lot of like entropy and modeling it as a distribution and, uh, a-any technique which helps with it. And the depth transformer is a conditional way of modeling this, and transformers are actually really good at it, even though that's a mini transformer.
So I think that worked pretty well too for us too. It's just that the main consideration is when you have a depth transformer, if you have K tokens, you need to do K autoregressive steps. Even though it's a small thing, it's like K steps, which is very latency heavy.
Mm-hmm.
With flow matching, we were able to cut it down significantly, so we are able to do the inference in quad steps or 16 steps, and it works pretty well. And there are more normal techniques to bring it down even further to like, in the extreme case, one step.
Like we're not doing it yet, but it at least the framework lends itself to, uh, more efficient techniques.
Yeah. And the image guys have done-
Yeah
... incredible work guys.
Yeah. It, uh, now you just send a prompt, and you get an image.
Yeah. Surprisingly, not enough, I think, image model labs use those techniques in production. I think it's, it's-- I feel like it's a lot of research demos, but nothing-
Yeah
... nothing I can use on my phone today.
The thing is, there's a thing that would be interesting here is that, uh, since, I mean, indeed, there have been so much work that has been done in the vision community compared to audio in this domain. I think there are so many learning flows here.
Yeah.
And there are so many things we can do to actually improve this by even further. So like when we start first version, but we have like so many ways to make this much better and much more efficient, cost efficient.
Voice Agents14:53
So-
Yeah
... so really it's not a new field at all, of course, but there are still so many things that can be done with it.
I should also mention, uh, for those who are newer to flow matching, I think, uh, the creator is-- this guy's name is Alex. He's done, uh, I think a NeurIPS, uh, maybe two NeurIPS ago. There was, there was a, there was a very good workshop that's like one hour on like this, what flow matching is.
I th- I would recommend people look that up. That's the other thing, right? The efficiency-wise, like I, I imagine like the reason it's open weights, uh, the reason you pick three point six B backbone, uh, it, or three point four B, uh, you are trying to fit to some kind of hardware constraints.
You're trying to, to fit some kind of like basic constraints. What, what are they?
Not necessarily.
Yeah.
I think, uh, something we care about in our model is they are efficient. So we have like a lot of separate model, for instance. Um, so we have this audio model that is very small, very efficient. Uh, we also have like a small OCR model that is really very good, highly efficient as well.
And, uh, I think on that project, maybe other, I think companies are going to take is to have like a very general model that will do a bit of everything, uh, but that is also going to be expensive.
Uh, on here what you want to say is if you care about this specific use case, if you can actually use this model, it just does that. It's extremely good at it, but also very efficient. That's why we can actually add these models, audio models like OCR that are like really, really good at that, and that will be much more cost effective than those on the, on the general models that will contain a lot of capabilities you don't really need at that.
So, so yeah. So we are doing like general model, but also like more customized model like this.
How does it compare to other TTS models? It's, you know, we are going full open wave. We are just dropping it like-
I think it's pretty good.
Yeah, I think it's pretty good. Like, it, it's definitely one of the best for su- for sure. It's probably-- I, I would say it's the best open source model wise.
People can decide for themselves, right? Yeah. Why now? How does it fit into broader Mistral vision? How do you see voice agents? How do you see voice? Like I, I think every year I've heard, "Okay, you're a voice.
You're a voice." There's a lot of architectural stuff. There's a lot of end-to-end latency that you're solving, but where do you see voice heading?
We had so many customers asking for, for voice. That's also why we wanted to, to build it. Uh, what's kind of interesting in this domain is that in a sense, if you take something simple like transcription, it, it doesn't seem like something that should be very hard to do for a model.
It's es- it's essentially it's pattern recognition, it's classification, and these models are very good at classifying, right? Uh, and nonetheless, when you talk to them, I mean, it's, it's not there yet, right? It's not... You don't talk to them the same way you talk to a person or something.
Maybe people don't realize it. It's, uh, in English, it's still much better than in any other language. Even compared to French, for instance, if you talk to this model in French, I mean, when you see people talking to this model, they, they will talk very slow.
They will articulate as much as they can. So it's not natural, right? We are not yet to this, uh... And I think, yeah, maybe the next generation will not know this, but, um, yeah. I think people that are maybe our age will actually always keep this bias of speaking very slowly when they talk to this model, even if maybe probably in a couple of years or maybe next year, it will not be necessary anymore.
Uh, but yeah. But what's interesting is to see that, um, yeah, even for like languages like, yeah, French and Spanish, German that are not low res- low resource languages, you have a lot of audio with this, uh, there, and still it's, it's not as good.
Enterprise17:56
So and I think a conseque-- I mean, the reason for this, I, I suppose, is just there is not as much energy, uh, as much effort that has been put than in some other modalities like, for instance, vision or like coding.
But, uh, but yeah, there is still a lot of progress to be done. I think it's just a question of doing the work on this, uh, you know? So it's like a clear path, I think, to, to get there.
It's a little fascinating because I worked on Google Assistant a, I think, a while back at this point , but it's-- I think, um, it's, it's kind of like when you take a step back, it's fascinating. It's not that long ago.
It was like four years ago or five years ago.
Mm-hmm.
And it's like now it's completely audio in, audio out, and the function calling, and the whole thing happens completely end-to-end and in a very natural-
Yeah
... uh, natural way. And still ways to go, like you were saying. Even despite all the progress, it's not like you're speaking to a person, uh, when you talk to any of these, uh, uh, uh, agents, bots or voice mode kind of situation.
It's still, still like a gap, I think. That, that's the great part. And I, I feel like with even the existing stack, we should be able to get to this, uh, uh, very natural, uh, speech, uh, conversational abilities, uh, soon enough, I guess.
And, uh, we'll also hope, hope to get there.
On schedule the next step, right? Because when you talk to these agents, like usually people are just writing to them or sometimes there will be this very clear, for instance, you are-- you want to write code, but you are-- you have a very clear idea of how you want the model to implement what you have in mind.
But, uh, so here you are going to spend like a lot of time writing. So it's not very efficient on-
Mm.
Like audio is really like a natural interface that is just not there yet, but I think it's just going to be there very soon.
How is it like building, serving, inferencing? Like we see a lot about it's very easy to take LLMs off the shelf, serve them.
Mm-hmm.
Uh, fine-tuning, deploying. Uh, I know you guys have a whole, like you have Forge, you have a whole stack of customizing, deploying.
Mm-hmm.
Is there a lag in getting that like distribution channel? Are you, are you helping there? Is there... Is... So like prompting LLMs, you can have them be concise, verbose, all that. Uh, they're built on LLM backbones, these models-
Yeah
... so, but, you know, how do you see all that?
Yeah, I mean, I think this is a lot of what we are doing with our own customers. I mean, very often they come to us, so it's for different reasons. I think one reason is sometimes they have this lot of privacy concerns.
They have this data that, uh, it is very s-sensitive. They don't want the data to leave the company. They want it to stay, uh, inside the company. So we help them deploy model in-house, so it's on, uh, it's on-premise or on private cloud, so they are not worried that, uh, you know, it's given to a third party and that there is some leakage.
Sometimes they have this kind of, um, many, many companies have this different, you know, sensitivity of data. They have like sometimes tier one, tier two, tier three data. Tier three can send it to the cloud. Tier one it has to stay there.
So then it creates some kind of heterogeneous workflows where it's kind of annoying. I mean, you cannot send some data to the cloud. This one you can. So, so here when we actually deploy the model for them is they don't have this consideration.
They are like not worried that, you know, this is going to leak. Every-everything is much easier. So we help them basically do this on the... So it's one of the value proposition. But the, but the other is very often, I mean, when customers use this, uh, off-the-shelf closed model, what's very sad is that they are not leveraging, you know, this data that they have been collecting for, for years or sometimes for decades.
Uh, so much data, sometimes it's trillions of tokens of data, uh, in a very specific domain, their domain, which is data that you will not find on, in the, in the public, uh, on the public internet. So data on which like the closed model will actually not have access to, one which they're going to be really good.
So if they're using like closed source models, they are basically not benefiting from all these insights, all this data they have collected for years. They can always give it into the context that it finds, but it's, it's still not as good as if you actually train the model on this.
So, so yes, that's basically what we help them to do. We actually, uh, provide them some, some Mistral projects, basically what we announced at, uh, GTC this week. So we provide them with this. It's basically like a platform with a lot of tools, uh, to actually help them process data, train on that.
Mm-hmm.
Uh, yeah, it's actually the same thing that we are using in the science team. So it's actually very battle-tested, uh, infrastructure, like a lot of efficient training code base for like a continued pre-training, uh, like a fine-tuning, even doing SFT, IL.
Uh, so we, we help them do this using the same tools as what our science team is building, uh, is using. So since it's tools that we have been using for like two years now, it's really battle-tested. It's really like, uh, sophisticated.
So it's the same thing we are giving to them. We are giving the company the same thing that what our science team using internally to actually build their own AI. And it makes a very big difference. I think sometimes customers, um, many people in general don't realize, uh, how much better the model becomes when you fine-tune it on your own data.
And you can have like a, your model is here, and you start from there. You have like a closed source model, which is sort of here. But if you actually fine-tune, it can actually really go much further than this.
And then you have like a very big advantage. The model is trained on your entire company knowledge, so it knows everything. Uh, you don't have to feed like, uh, ten K tokens of context at every query. So it's, uh, it's much easier.
It's a bit... I think using a closed source model is really sad because it basically puts-- you are not leveraging all this data, and you are going to be using the same model as all your old competitors when you could actually using, I mean, everything you have been collecting for years, which is really valuable.
So, so yeah. So we help, uh, basically customers do this. We, we have a lot of solution, um, I mean, deployed for the engineer that go in the company that basically look at the problem customers are facing. Uh, they look at what they're struggling to do, uh, what we should do to solve it.
So we help them solve them together. So it's, uh, I think our approach is a bit different here than some other companies and competitors. It's, we don't just release an endpoint and say, "Do some stuff on top of that," or, "We don't just give a checkpoint."
We, we really look very closely with customers. We look at the issues they have. We help them solve them. We really make some tailored solution, uh, for, for the problems they are facing. Some example are also going to be...
Sometime where some customers, they really wanted to have a really good model, really performant on some like Asian, rare languages. On the-- If you take some of the shelf models, they, they can speak it. They can write in this language, but it's, it's not amazing.
This language will be like maybe zero or one percent of the mixture. So it, it has, it has been included during training, but very, very little. So what we did here is of course retrain a new model for them.
But, uh, this language was fifty percent of the mix, so it's much, much stronger. It knows of the dialects, it knows the like, and the slang. So it's, uh, yeah, so it's some example of things we can do.
And it's, it's really arbitrary custom. I think some of our customers, for instance, they wanted some, um, they wanted some 3D model that can do audio with like a very good at function key. So something you wanted to put in the car.
In particular, they wanted this to be offline because in a car you don't necessarily have access to internet. So, so yeah. So here we can actually build these solutions. There is no like model out of the box on this in the internet.
You have this very We have this very general model, very generalist, like reasoning on strong model. But for things like this, there, there always want like specific solutions on, on it. On some other reasons, sometimes they come to us is because, um, you know, like they, they experiment with some closed source model, or they get some prototype.
Uh, they are happy with what they build. They- it, it works well. They're happy with the performance, and then they want to go to production, and then they realize, oh, but it's extremely expensive. We cannot push this. It's, uh...
So then they come back to us and they say, "Can you help, help us build the same thing as this, but, uh, using something much cheaper on here?" On here, we can sometimes build something ten X cheaper by just fine-tuning a model, and it would be better on-prem, uh, on their own server, and also much cheaper as well.
So yeah.
That's the Mistral pitch right there. Take all the money.
And, and I mean- ... outside of that, you do, we do put open-weight models so people can do this themselves. I feel like not enough people go out of their way to-
They're not going to. They're gonna ask them to do it. They're...
Ask the expert.
Initially, we didn't know, I mean, we were not completely sure at the beginning of the company because, um, I think our strategy was not exactly the same as what it is today. But what we underestimated initially is the complexity of deploying this model and like connecting them to everything to be sure it has access to the company knowledge.
And, and it was, yeah, on-- We were seeing customers struggling with this, but it was even-- that was two years ago, and now things are much more complicated because now you don't, you don't just have, you know, text on safety, on like a simple instruction following.
Now you have reasoning, like agents. You have like, uh, tools. You have like a multimodal audio. So it's much more complicated than before. And even back then it was hard for customers, so they really need some support, and this is why we're actually providing like always some, uh, for deposition as well as in the processes.
I'm curious, is, is there also, uh, voice fine-tuning that people do?
So in this, in this Forge, we, we also have the-
Oh, really?
... a un-unified, uh, framework. Uh, and the, the hope is like the Voxers, uh, speech-to-text, uh, that we released earlier this year, and, uh, even the Voxtral chat that we released last year, and I think, uh, a big-- People-- I think there's a big, uh, a rich ecosystem of, uh, people fine-tuning Whisper, and people want the same thing with Voxer.
It's much stronger-
Mm-hmm
... than, than Whisper. And yeah, the, uh, the platform offers that kind of fine-tuning.
Yeah.
Which could be any kind of fine-tuning. Like for instance, even sometimes people want to support new languages to this, which are, are tail languages, which we hope to cover, um, ourself natively. But if there is a language where you have data and you want to fine-tune, I think this is a good use case.
Or the other use case is you-- it's the same language, like even English, uh, but it's in a very domain-specific way where-
Yeah, terminology, jargon, medical stuff.
Exactly.
Yeah.
And also the specific, uh, acoustic conditions, like, uh, there's a lot of noise-
Yes
... or other. And the model will do, uh, decently in most conditions, but you can always make it better. And that, those are some of the use cases where you can, uh, improve it e-even further. And, um, that's one good use case for this.
And for, uh, our text-to-speech, we're just releasing it, so we'll, we'll have support for that soon, too. I think it's similar use case. It's little different, the kind of things that you want to extend a text-to-speech model to, which could be like voice personalization, voice adaptation for enterprises.
And many enterprises need like very specific kind of tone, very specific kind of like personality for this kind of voice, and all of those are like good, good use cases-
Sure
... for fine-tuning.
This is what I was gonna ask you. You know, we never talked about cloning, voice cloning here. How, how important is it, right? Like I just clone a famous person's voice, okay. But...
The main use case would be like for enterprise personalization. Like enterprises need like lot of customization. Uh-
Yeah
... you don't want the same voice for like all the enterprises. Each enterprise, uh, want a customized, specialized, uh, something which is representative both their brand and also their, uh, I guess, safety considerations. And the use case, like I think the kind of thing that you would deploy as a empathetic assistant in the context of a healthcare domain would be very different from the kind of thing that would be in a customer support bot, uh, and would be different from, uh, like more conversational aspects.
I think th-those are the customizations you would expect from enterprise, and that's the main use case, at least from, from our side.
Mistral Small28:53
My, my basic example is you don't wanna call two customer services and have the same exact voice. You know, it's just, it's gonna be weird. But also, um, on the technical side of this, so there's like a few things in Voxtral that I thought were pretty interesting.
Um-
He's a big fan of this paper.
Oh.
Very good paper.
I think he said this is the best ASR paper he's ever read.
Yeah, I've, I've hyped up this voice paper enough. We covered it somewhere. But, uh, a big thing, so Whisper is known for thirty-second generation, uh, thirty-second processing. You extended this to forty minutes. There was a lot of good detail in the paper about how this was done.
Even little niches of how the padding is so... Like i-it's very much needed. You need to have that padding in there. The synthetic data generation around this, I'm wondering if you can share the same about the new speech-to-text, right?
Uh, text-to-speech. So how do you, how do you generate long-form coherent? How do you generate, you know, how do you do that? And then any gems, is there gonna be a paper?
Yeah, yeah. There, there would be a technical report.
Okay.
Yeah, I think it will have a lot of details. Uh, but m- uh, I think the summary of it, actually, some of the considerations in this paper were because we started with the Whisper encoder as the starting point, and now we have in-house encoders.
Like the real-time model, for instance, which we released in January, we also released a technical report, uh, for that real-time model as well, which is this dual stream architecture. It's an in- an interesting architecture. Uh, you should check it out.
And there we have a causal encoder, and I don't think there's any strong multilingual causal encoder, uh, out in the community, so we thought it's a good contribution. So that's, that's one nice encoder. If the other people want to adapt, that's, that's a good encoder.
And we train it from scratch. I think our full stack is now mature enough that we're able to train super strong encoders. And some of these considerations, like this padding and stuff, is a function of the Whisper encoder, and n-now that we, uh, train, uh, encoders in-house, uh, the design considerations are different.
And for the question on the, uh, text-to-speech, I think that also like leans on to the original, uh, autoregressive decoder, uh, backbone. I think, uh- It says we're a-almost identical considerations. I think the long context in, uh... It's not even long con- So the model processes audio at 12.5 hertz, so one second maps to, like, 12.5 tokens.
So I think one minute is like 720 tokens. So you can get like up to 10 minutes in like 8K context window.
Mm.
And you can get half an hour in 30K context window. So that's, uh, and 32K context is something that we are very comfortable training on. We can extend it to even much longer, 128K. So, again, you can naturally see how it can extend to even hour-long generations.
Yeah, we need the, like, data recipe and the whole algorithm to, uh, work coherently enough through such long, long context. But the techniques are some, some way, uh, very, um, similar to the text long context modeling. Uh, and the key difference is it's just doing flow matching autoregressively instead of like, uh, a text open prediction.
Okay. I think that was most, most of the, the, the sort of voice questions that we had, but, um-
I have a, I have a big question on-
Mistral Small?
Mistral Small.
Let's go.
Um, so what is Small? How do we define Small? What is this? What is this? I remember the days of Mistral 7B on my laptop. It's not fitting on my laptop. I, I could, I could run it on the big laptop, but-
It's just a different question of terminology. Like here what we mean by Small is the number of active parameters. But it's true, uh, we did give it a, another name. But, uh, yeah, we could have called it Medium, but then I think there's just a jump with the large models of this, uh, I suppose.
But yeah, it's a model that we released, a mixture of experts. Uh, it's a model that combines different, uh, model. Before, uh, what we were doing is simply that we had, uh, one model, general model for Mistral doing instruction following, where like a separate model that was Devtral, so really good at coding, specified specific to code.
Uh, we had another model for reasoning, Magistral. Uh, so these were separate artifacts built by different team at Mistral. And now what we are doing is basically merging all of this-
Mm
... into one. It was even Pixtral was the first vision model we had was like a separate model on the... The way we kind of do things internally is that we have like one team focus on one capability, build one model.
Open Source33:05
Uh, and then when it's mature enough, we decide to merge this into the main mixture. So, so here it was the first time we basically merged all of this into, into one. There are some other things we didn't have time to merge at the time.
For instance, like more capabilities or like, uh, function coding I think would be, uh, it's going to be mu-much better in this trial, smaller proper one. But, uh, but yeah, so it's, uh, our latest model on that we're working on.
Also the larger version of this.
And yeah, I mean, key things, it's, it's very sparse, 6B actives, so you know, pretty efficient to serve, uh, 256K context. Um, yeah.
I think what, what's interesting is just this general theory of like developing the individual capabilities in different teams and then merging them. Where, where is this going, gonna end up? Um-
Like we've seen the five things put together in this.
Yeah.
What are the next five team?
I think actually OpenAI has like kinda gone away from the original 4.0 vision of the omni model. That's what, what they were selling, right? All modalities in, all modalities out. Uh, but I feel like you might do it.
I mean, I think there are some, uh, modalities where it's not completely obvious. For instance, for, for audio. For audio here, if you want to, to do transcription, I think it makes no sense to use a model this large.
If you just want to tr-transcribe the... It's, it would be very inefficient. Like, uh, if you want to do audio, you probably just want to build one big 3D model. Performance would be essentially the same. It's going to be incredibly cheaper.
Uh, so here that's why we want to have like a, a separate model that just does this. Yeah, I think the question is just, yeah, if you are talking to your model by speech, and you're asking like a very, very complex questions and how you do this, uh, on the other approaches to cascade things, uh, do you want to put audio in a model that has like a one key or that's like a-
Mm
... not a competitive discussion, I think on our if you are going into the direction, but that's the possible approach. Uh, but yeah, but I think for, for us, the next capabilities we want to try to, to integrate into these models right now are going to be, yes, like more coding, more, more reasoning, but also, uh, I think more capabilities that people don't talk too much about, but that are important, I think for our customers in, uh, on different industries.
For instance, things are around like legal, finance, computer-aided design, a lot of these things that, uh, is this model out of the box are harder to put at that because people really don't prioritize this. There is no like two new benchmark on that.
Um, but it's not hard to make this build a model good on this. It's just that to do the work like sourcing some data, processing it, in-increasing it into the next version. So, so yeah. But, uh, we have other things we will merge into this.
Uh, yeah.
I, I, I think for, for voice, yeah, the key thing, I think over maybe like the last year or so with VO and Grok Imagine and all these things, is, uh, joining voice with video, right? Which people don't understand spatial audio because like most TTS is just like, "Oh, I'm speaking to a microphone in perfect studio quality."
Leanstral35:51
But when you have video, like the voice moves around.
That's true. The configuration was a little different in the sense that, uh, uh, there it's like a, uh, a standalone artifact where you get the whole, whole thing and you consume it. But in the conversational setting, it's, uh, you need the extreme low latency-
Yeah
... streaming-
Oh, yeah
... would be one of the primary concentrations. And-
You can build a giant company just doing that. So you don't need to do the voice there. But I, I was just, you know, on the theme of merging modalities, that is something I am like, wow, like I didn't...
Everyone up till, let's say, mid last year was just doing these like pipelines of like, okay, we'll stitch a TTS model with a voice thing and a lip sync thing and what have you. Nope, just one giant model.
Yeah.
I have like a two-part question. So one is, you know, it's still open. It seems like open source is still very core to what you guys do. And you know, I just have to plug your paper. So, uh- ...
DAN 2024, this is the only Mistral experts like
Very fundamental research on how to do good MoEs. Paper comes out. Very good paper for anyone, but, you know, that's, that's just side tangent with all this- No, this thing, this thing caused- We bring back, we bring back.
I mean, eight by 22 was like the nuclear bomb for open source. Like... I think it takes 7B more. S-seven B more? Yeah. Okay. Yeah. But this is a bigger for 7B. Yeah. Uh, yeah. I, I don't remember this.
I, I rememb- I don't think it was January, right? It was like NeurlPS, uh. It was- It, it dropped during NeurlPS, and then everyone in NeurlPS was- Yeah, it was December of '23. But I think, yeah, the model was- The date is wrong.
I think there was a big build. Yeah. It's just a little update, probably. Yeah, no, but you have a point to make. No, I mean, you gotta check that. But then, I mean, I just wanna hear more broadly on open source for you guys.
And- Yeah ... when you had asked earlier about what's next, what are the other, you know, side tabs working on, you, you put out Leanstral. Um, this- This one is a surprise. Yes. I was like, "I don't... This doesn't fit my mental model of Mistral."
Yeah. I mean, um, I mean, first, for open source in general, I think it's really something which looks to the, the journey of the company. I think we started it around this. We, uh... Sorry. We have been open sourcing with this since the beginning and even before this.
So before this, so me and Timo at Meta, we released Llama. And I think what was really nice to see that before this, for most researchers, like universities, it was impossible to, to work on LLMs. There was no LLM outside.
And if you look at many of the techniques that were developed after, for instance, Llama was open sourced, like, uh, all these post-training approaches, uh, like even DeepOD, like performance optimization. All of these were done by people that had access to this model, and it would have been impossible to do, uh, without this model.
So it's really, uh, making sense move faster. So we really want to contribute to this open source ecosystem. Uh, I think like the DeepSeek and also like very lot of impact. All these papers that are, uh, I think in the open source community are really helping the science community as a whole to move faster.
So we want to contribute to this ecosystem. That's why we are releasing very detailed technical reports. So Mistral and our first reasoning model. And, I mean, abolition of the reasons, things that worked, things that did not work as well, that I think helpful.
Uh, on the... Yeah, so for the audio model, we're also going to share a lot of details, share a lot of them for the real-time model on the... Yeah, so we really want to continue this, basically belong to this, uh, community of people who share science.
Uh, I think we really don't want to be, you know, living in a world where the, the, the smartest model, the best models are only behind, you know, like, uh, closed doors, uh, only accessible to like a few companies that we, you know, have the power to decide who can use them or not.
I think it's kind of a scary future we don't want to live in. We really want this model to be accessible to, to anyone. We want intelligence to be used and accessible by anyone that can use it. So yeah, so that's why we are pushing this mission and source model on the...
Yeah, so not... So yeah, Voxtral TTS, so it's open source, not the first model, so not the best. Uh, on the... Yeah, Leanstral, I think is also one step, uh, into this direction. So it's kind of, yeah, a bit different than what we are usually releasing.
But, um, we have a small team internally working on, on formal proving, formal math. So I think a subject we care about in general, uh, and we are working on reasoning. I think we started too early before LLMs.
Doing reasoning without LLM is very hard, especially when you work with formal systems because, uh, the amount of data you have is negligible. It's a very small, uh, community of people writing like formal proofs. Uh, but the reason why we like it is because I think there is...
Uh, if you look at what people are doing with reasoning, is there... I mean, the problems that you can use are usually going to be problems where you can verify the output. So for instance, all this AIME problem, where the solution is a number between, uh, one and like a thousand.
Training40:25
So you can verify, compare this with a reference, uh, or it's an expression, you can actually compare the output expression, generate your model with a reference. But there are many-- Most of the math problem and most of the reason problems, there is no like way to easily verify the solution.
If the question is, uh, show that F is continuous, you cannot compare any reference, right? If it's like prove that this is true or prove this property, there is no way to... You cannot like simply verify the correctness of your proof.
So it's hard to apply the... There is no verifiable reward here. So what you could provide is, of course, like a judge, an judge that will look at your proof, but it's very hard and it's very... You could also have some reward hacking happening there, so it's difficult.
Uh, you could provide like a reference proof, but then there are also many ways to prove the same thing. So if the model says give negative reward because, uh, it's a different proof, but maybe it was still a legit proof, just different, so, so it's not going to work well.
So, um, what's nice with Lean, uh, and with formal proving is that you don't have to worry about this whatsoever. We just- They all function... As long as it compiles in Lean, it's functionally the same, right? Exactly. It's like a program.
If it compiles- Yeah ... it means it's correct. Yeah. It's very easy. Uh, and you can apply this on any kind of- It's just way too small to like no human will actually go and do it. Yeah. That's kind of...
Exactly. So- ... the only people who can do it, it's like a very small community of people doing a PhD on that. So it's super small. Uh, and it's kind of sad because it's actually very useful on not just math, but also in software verification.
So for instance, software verification today, it's a tiny market. Uh, very, very few industries work on this, and will need that. It's usually going to be like companies, uh, like building airplanes, robotics, like, uh, things where they absolutely want to be sure.
I mean, like life depend on this, but it's very, it's very rare that people formally verify the correctness of their software. Uh, but I think one reason for this is simply that it's just super hard to do. Uh- Are you thinking of TLA+?
It's the language that some people do for software verification. No, I- Hard ...
Yeah,
yeah. One of my theories is that, um, you know, because the proofs take so long, it's actually just a proxy for long horizon reasoning.
And, um, coherence and planning maybe. A lot of people will say, like, "Okay, it's for people who like math. It's, it's for lean. Okay, it's like a niche math language. Who cares?" But actually, and you use this as part of your data mixture for post-training and reasoning, actually it might spike everywhere else.
Yeah.
And, and like I think that's un-underexplored or like no one's like really put out a definitive paper on how this generalizes.
Yeah, absolutely. And I think even, uh, that's kind of what we are seeing already. For instance, if you do some reasoning on math, then the mechanism to do reasoning on code even. Yeah, just code that with the, in the, in the earlier stage.
So it deputies there is some transfer, uh, some sort of an emergence that happens. Uh, and I think some... Well, it's also interesting, it's not just, uh, I think the topic in general, but it's, there is a lot of connection with this on encoding agents because, uh, sometimes the model can see like a theorem that it has to prove it's very complex, but then it can take the initiative to say, "I'm going to prove this three lemma.
I'm going to like suggest three lemmas, and I'm going to, in parallel, prove each lemma." So three of them in parallel with sub-agents, but I'm also going to prove them in theorem as three, the three lemmas are true.
So you can do this sub-agent approach. So pretty interesting. You can even if you fail to prove one of the lemma, you can actually maybe you'd succeed to prove the normal lemma too. So you get some, also reward here.
So it's a bit less sparse than if you just get to zero one for the entire thing. So it's pretty interesting. I think we can actually, um, bust up here.
Yeah. It's also an interesting case just for specialized models in general, right? Like the cost thing you show is pretty interesting. Like, um-
Forward Deployed44:19
Yeah
... similar score-wise, you're, you know, thirty, seventy, a hundred and fifty, three hundred bucks comparison.
Yeah, small model.
I think cost is a bit unfair, right? Because this one is at like inference cost. It's always there on top with their margins on top of it. But, you know, we don't know anything else, so we gotta, we gotta figure it out, right?
Uh-
I, I did wanna actually push on that more, so, um, not on cost, but you, you mentioned about, okay, it's a great way to have verifiable long context reasoning. Uh, what are other frontiers that, you know, I'm sure you guys are working on internally.
There's a lot of push of people pushing back on pre-training, um, scaling RL, pushing compute towards having more than half of your training budget all on RL. Uh, where are you guys seeing the frontier of research in that?
You mean with ZL, uh-
Just in foundation model training in the next... You know. One thing that you guys do actually is you, you do fundamental research from the ground up, right? So you probably have a really good look at where you can forecast this out.
Yeah. I, I think for us, uh, we are still working a lot on the pre-training side. I think we are very, very far from any sort of situation on the pre-training. Uh, I think ML for pre-training will be like big step up compared to everything we have done before, so we are pretty excited about this.
And I think on the RL side, I think now we have, uh, more and more to think about this algorithm that will actually support this very long trajectories. I think, uh, when it was, for instance, GPU, for instance, doesn't really work with this tiny bit of policy.
Which was okay initially, because you are solving math problem that can be solved in like a few thousand tokens. So the model can like regionalize them pretty quickly. So when you do your update, the model is never too far off.
Still, it's still not too far off. But now when you are moving towards these kind of problems where something takes hours, like six hours to get a reward, then your model is completely, uh, policy, so you have to actually build new architecture, uh, new infrastructure that supports this, but also new algorithms.
So now everything we are doing internally, we are trying to, uh, build some infra that we actually anticipate this, uh, what we have in like a six months, one year, which is this extremely long scenarios on the platform they can.
I, I think when we started Mistral, part of me and maybe also Timothy, we wanted to, you know, recreate this very nice environment where people are there, they can do research they like, uh, with like a lot of resources.
So, so it was nice. Yeah. I think things, things changed a lot when, uh, I think when, when ChatGPT came out. Uh, I think after that everything was really different. Uh, this time around it was the same again.
Uh, but, uh, but yeah, but it was nice. And I think we also want to recreate part of this culture we had before.
Coming to the end, uh, we're, we're just-- Obviously, I think you guys are doing incredible work. Uh, you've laid out a very impressive, uh, vision for open source and for voice. Um, what are you hiring for? What's the, uh, next like, you know, what are you looking for that you're trying to join the company?
Outro47:01
Yeah. So we are, um, hiring a lot of people in our science team. We are hiring, uh, in all our offices. So we have like a-- Our HQ is in France, uh, in Paris. Uh, we have like a, a small team in London, like a team in Palo Alto as well.
Recently, we opened some offices, uh, in, uh, in Warsaw, in Poland, uh, so one in Zurich. Um, we also have like some presence in New York as well, and soon one in, in San Francisco. So we are a bit everywhere, also like hiring people remotely.
So we are growing the team, uh, trying to hire like very strong people. Uh, I think we want to stay. So the team is not... Still a fairly small team, and I think we want to keep it that way, uh, because we, uh, we find it quite efficient.
We like a small team, uh, very agile, so, uh, so yeah.
Okay, let's, let's focus on science and the forward deploy. Um, we actually are strong believers in science. We started the, uh, our new science pod that focuses specifically on AI for science. Uh, what areas do you think are the most promising?
What we are pretty excited about right now, and something we have already started doing, and we will probably be able to share more about this in a couple of months, is that we are exploring AI for science. Um, there are a lot of areas where we think that, uh, you could get some extremely promising results, uh, if you are to apply AI in these domains.
Uh, there are a lot of low-hanging fruits. You just have to find these domains where actually AI has not been yet applied. Uh, and it's usually hard to do because the people working in those domains don't necessarily know the capability of these models.
They don't know how AI will-
You just have to pair them with-
Yeah, exactly. You have-
... researchers
... matching, which is actually hard to do. Uh, but this matching, we are kind of doing it to naturally with our customers. So we have, uh, some company we work very closely with. So for instance, ASM Industries are one of our partners, so we are doing some research with them.
And there are, there are like tons of extremely interesting problems. Problems in physics, in science, material science that they are essentially the only ones to, to, to work on because they are doing something no one, no one else is doing.
On the... Yeah, so there are many domains where AI can actually revolutionize things. It's just you have to, to think about it and be familiar with what it can do well now and to apply it. So, so yeah.
It's something we are more modeling with our partners, with our customers. Uh, so AI for science is like, like one big thing.
Yeah. Uh, okay, and then for deployed, um, what, uh, makes a good for deployed engineer? What do they need? Where do people fail?
I think it's, um, um, usually you need people, uh, that are very familiar with the tech, and not necessarily, uh, with a lot of research expertise, but that are actually, uh, pretty good at using this model that can actually like, uh, you know, that know how to do fine-tuning, that know how to like start some RL pipeline.
Uh, and it's, uh, it's not easy. It's something that most cus- majority of companies will not be able to do this, uh, on their own. So here, I think we need people that has-- that are, you know, that like to solve problems, that are excited about like solving some complex, very concrete problem.
It's applied science, basically, on the... Yeah. So I think it's not too different, I think, from the skills you need when you do research, because essentially you are trying to find solutions to problems that customers have not yet solved.
Sometimes it's easy, sometimes you, you have to do the work. You have to like, uh, create synthetic data, like, uh, like find some edge case. So it can be... Yeah, depends on the problem. Uh, but, uh, but yeah, you have to...
I think you need also a bit of patience on the, yeah, be creative. Uh, I think very similar skill as researcher.
The diversity of the work they do, it always surprises me. It's kind of, it's, it goes all the way from the kind of stuff they encounter in industries. It's just very interesting, I think. Uh, they-
Any, any fun, like, success anecdotes?
I mean, yeah, it can be like really training this small model on edge that just do one specific thing. It can be like training some very large model with some specific languages as well, making models very good at, uh, some to-dos, like for instance, computer ID design, these kind of things.
Is that in pairing with Vision as well?
Yeah.
A defect detection, uh, for, for like, uh, chips or like in, in factories identifying things. Like it-- the diversity could be anything where you can deploy these foundation models.
Yeah.
So the, the work, uh, to make it work in that specific setting, basically whatever it takes to make it like add value, uh, in that specific workflow.
Yeah.
And it goes kind of across the stack, right? Like even just pulling up the website, like you have-
How did you... It's so broad
... on compute is so broad. Uh, we didn't even touch on Mistral Vibe. We have a- ... live coding CLI tool. Uh, one thing you guys were actually like, I think the first to was agent-
Mistral Agents
... Mistral Agents. Yeah, the agent builder, you can serve it via API and all that. Like, uh, and you know, I'm, I'm guessing forward deploy people will-
Yeah
... help build that out and stuff.
It's also why we are-- So we are doing many things, but I think that's also part of the value proposition that sometime new customers are always very extremely careful about their data, and they don't want to-- They don't like, you know, trusting so many partners.
Tru-trusting one partner for code, giving your data to another third party for like audios, and another one. So they don't like this. So here, what they really like with our approach is that we can help them on anything.
Mm.
Uh, so they don't have to like send their data out to so many clouds. So yeah.
I think that there can be many orders of magnitude more FDEs than research scientists, and like they don't need your full experience, but they're still super valuable to, to customers.
I mean, in practice, these two teams are still quite, uh, intertwined. I mean, very often-
Yeah
... so first of all, they are using the same tools, the same, uh, data pipeline and everything. Uh, and the-- It's, um, it's very helpful for the science team to get the feedback and the solution team, 'cause they can say, "Look, these customers are trying to do this.
This is not working. Can we be sure that in the next version-"
Yeah.
This is basically a real world eval.
Yeah.
Yeah, exactly.
It's a real world eval, and it's not something, for instance, if you are just working in a lab, it's just ships model, but you don't do this work of deploying the model for customers. You have no idea whether your model is good at this edge case.
Or like, for instance, you know, even in your performance, right?
Yeah.
There is a very gap, big gap between, uh, the public benchmarks that are very like academic on the, the real cases.
The real cases are just very diverse and-
Yeah
... uh, in the specific context of a customer, you can fine-tune and make it like, uh, first like evaluate, but create a solid eval benchmark, and then measure in the context of their... The kind of audios, like for instance, one use case is literally just there's a word for kids, and they have to just say it out.
It's a very specific thing. You're just saying one word, and then you have to like, you'll, you'll, you'll grade the kid, uh, whether they did it right or not. It's like RL for kids. But, uh, so there are very diverse use cases, and the idea is that the, the applied scientist engineers will go and make it better, and then, uh, from the learnings we incorporate it into the, uh, base model itself.
So it's, uh, it, it's just better out of the box.
Yeah. It, it's a good full circle system, you know? Like, the, the foundation model evals are all just proxies of what you really need. You're never gonna have one that's just... It, it doesn't make sense for there to be, you know, a one-word transcription like that.
It's, it's not something you want to bet on. Perfect. Well, everyone should go check out, uh, everything that Mistral has to offer and try the TTS model, which we'll link in the show notes. But thank you so much for coming.
Th-thank you. It's such a pleasure.
Thank you, guys.





