LALatent SpaceOct 11, 2025· 45:05

Building Jamba 3B: the tiny Hybrid Transformer State Space Reasoning Model - Barak Lenz, CTO of AI21

Barak Lenz, CTO of AI21, presents their Jamba 3B model as a tiny hybrid transformer-state space model that brings long context capabilities to edge devices, and argues enterprises need AI systems like Maestro over standalone models. Lenz explains that the 1:8 ratio of attention to Mamba layers emerged from extensive ablations, with attention placed in the middle of the block working best. The Jamba 3B model uses a 1:12 ratio to maximize efficiency, fitting the same context length as larger models with a fraction of memory. He notes that pure Mamba underperformed on some tasks but hybrid models resolved those deficiencies, and that images quickly become long context problems (4 images can be thousands of tokens). For enterprise, Lenz advocates for model-agnostic orchestration layers treating models as "actions" with statistical properties, enabling continuous learning and cost optimization without vendor lock-in. Drawing from his algo trading background, he compares training frontier models to developing trading algorithms, emphasizing the importance of world-class engineering and avoiding brute force reasoning.

  1. 0:00AI21 Origins
  2. 4:23Hybrid Jamba
  3. 12:02Jamba 3B
  4. 15:44Scaling Training
  5. 21:15Hiring Philosophy
  6. 24:25Path to AI
  7. 28:22AI Systems
  8. 41:50Future Direction

Powered by PodHood

Transcript

AI21 Origins0:00

Host0:03

Okay, so we're here in the remote studio with Barak Lenz, CTO of AI21. Welcome.

Barak Lenz0:09

Hi, nice to meet you. Nice to be here.

Host0:12

I actually... I was gonna say CTO and chief scientist, but, uh, you don't have that on your LinkedIn, so I just said CTO.

Barak Lenz0:17

No, no, no, I'm, I'm just a CTO, not the chief scientist. I'm, I'm not an academic, so calling me chief scientist would be stretching it too much, so let's stick with CTO.

Host0:27

Yeah. Awesome. So, uh, a lot of people may not-- may have heard of AI21. Some of you may have confused, uh, AI21 with, like, AI2 out there. Can you, can you sort of introduce, like, what is A to-- AI21 from the, from the top?

Barak Lenz0:42

Yeah, sure. So, so, so we were founded, like, seven plus years ago by Ori Goshen and Yoav Shoham and Amnon Shashua. Uh, I think Amnon does not need a lot of introduction. People know him, CEO of, of Mobileye.

Yoav was a professor at Stanford for a long time, and Ori is very, uh, accomplished in, in Israeli, Israeli startups. And, and the-- I think the general premise was that, you know, deep learning is super cool and super useful, but it's not enough.

We, we wanted to bridge classical AI with, with new AI, and it just happened to be that we started the company, you know, just before BERT came out. So we were there to see everything unfolding and be a part of it.

So we were training models as early as 2018. Um, I'll try to speed the timeline up just to, so that it's not so... And from the top, we always was-- we always wanted to do both something very scientific, very sound, but also have an application side.

And, and we had an application, and we still have a really nice application called Voitune that some people know. It's not our, our main focus, but it's, um... It was very useful for us getting, getting traction. And while working with Voitune, we were again, and also training our models, GPT-3 came out, and then we were very heavy users, and we decided to train our own GPT-3.

Then we came out with Jurassic-1, which is probably the way at least some people knew AI back then. So we had a model that was slightly bigger in our performance metrics and was slightly better, but pretty much on par with GPT-3, and it was called Jurassic-1.

Since then, we've released several models. Recent model lines is called Jamba, which I think the fascinating part about it is it's the first hybrid model. It's not just attention. We've, we've mixed attention, and maybe we can talk about why and how we did it.

And we're actually releasing new versions of our Jamba model soon, including a 3B dense model that's aimed for long context on edge devices, and we can talk about again how and why this fits in. And we were also, uh, pioneers of AI systems.

We went out with JIX, again, a while back, which was the first AI system that, that tried to use tools back when no one was talking about tools. And actually, when LangChain came out, they, they, like, said, you know, Miracle Systems were, was our inspiration.

So we were dealing with this, again, because we wanted to bridge the gap between deep learning and, and classic AI. And for that, you need to go out to code and do tool usage and be more rigorous than just, you know, running a model.

Um-

Host3:23

Is Ma- is Maestro also part of that group of systems?

Barak Lenz3:26

So Maestro is a, a enterprise offering and, and how we see the, the, the future of, of this industry, and maybe we can get to it a little later. Why do we need AI systems? What is Maestro, and, and how does it do the job?

But again, this is a bird's-eye view, so we've been in it for, like, eight years. We've been both training models, doing real first B2C applications, and now we're doing B2B applications with Maestro, which was our intention to begin with.

But in a way, the market wasn't ready for B2B generative AI, and I think just now it's starting to, to warm up to this. It's, again, the, the, the adoption is still not what people would want it to be, but, but it's getting there.

And I think with systems like Maestro, enterprises could really integrate models into their workflows. But, but it's, again, it's still early days, I think.

Host4:23

Yeah. Awesome. So we'll, we'll get to the AI systems part, which I, I know you're very passionate about. I think the, uh, we should definitely start with Jamba. That's, that's where I, uh, started to, to say that, "Oh, okay, AI21, uh, has something new in, like, the sort of post-ChatGPT era that is, like, really frontier."

Hybrid Jamba4:23

Host4:42

Uh, the way-- When I, when I, when you came out with it last year, I said that basically it, it dethroned Mixtral. As I remember, a lot of people at the time were-

Barak Lenz4:51

Yeah

Host4:51

... very interested in, like, MoEs and, like, just, like, the, the sort of efficiency conversation. Like, I think, like, researching hybrid LLMs is something that basically only Tri Dao and, and Albert Gou were kind of, like, looking into.

Barak Lenz5:03

Yeah.

Host5:04

What's the story of you guys exploring hybrid LLMs?

Barak Lenz5:08

That's a great question. I'll tell you, I'll tell you from my point of view. So we were working on a project that I'd call J3, the, the, the third version of our Jurassic models.

Host5:17

Jurassic.

Barak Lenz5:17

And we knew it's gonna be an MoE. And I think MoE, the, the main reason to do an MoE is it is way more cost efficient in terms of training budget. And then for inference, we wanted to make it work.

So we designed J to have a version that, that fits on a single GPU, a single A100 or H1-- 80 gigabytes, and then a bigger version to, to fit in a single pod. So we had inference in mind to begin with.

But then as we were doing the, the ablations, I, I came across the Mamba paper, and it actually was pointed out to me by, by several people, and I started reading it. And, and I already knew, like, the S4 versions, you know, previous version, that they weren't up to par.

And what I really, really liked about the paper is that unlike some other papers that In order to look good, compare themselves to baselines that are not the best baselines to compare to. They straight off the bat compare themselves to, to latest, uh, attention architecture introduced by Llama.

Th-that introduced a lot of corrections that cause things to work, because before Llama, before all the, the layer norms, there were a lot of problems in getting models to train. So, so maybe in, in some other meeting, we can talk about what it took to get J1 to train before anyone was doing models training a GPT-3 scale model.

But the Llama architecture made it more approachable because they made a lot of the adjustment that they made the optimization work and the activations not explode. So the paper is straight off the bat, you know, compared to the best architecture out there and, and looked competitive and released custom kernels and code.

So I was, like, telling my, my engineers, "Let's just try it out and then see how it goes." And we compared it, and we ran our entire, like, evaluation dashboard, which was already big back then. It's huge now, but it had tons of tasks.

And what we saw is it's really competitive. In terms of perplexity, it was pretty much on par across the board. But there were a handful of tasks, few, few short tasks, s- in which it underperformed. And when trying to investigate why, we came to the conclusion that it's, it's missing some attention, or at least potentially missing some attention.

And when we started experimenting with hybrid models, meaning interleaving attention with Mamba, we just saw that the results were just better across the board. You know, perplexity was better, but also all of the deficiencies that we saw in pure Mamba, some tasks degrading, just we did not see that degradation anymore.

And then we kept on advancing head to head. And in advancing, I mean, you know, larger sizes and, and more training data. What we call the vanilla architecture, like a vanilla MLV, and what we call Jamba, or what is now called Jamba, but I actually called it Jamba to, to begin with, and it ended up sticking.

And in all the comparison, the Jamba was just better. And, um, and so we, we, we decided that we're gonna try and train our models, and at that point, no Mamba model was scaled to the extent of the, the mixed model that we released.

Jamba Mini was the first that we released. It has 13 billion active models and around 52 total parameters. Mamba was not scaled above 3B at the time that we were training it. So we had some optimization issues that we had to go through, and that's always fun, fun to do.

Debugging a model is super interesting. How do you dissect it and see the explo-- We were able to, to make it work, the optimization itself, and we ended up with a model that's at least as good as, as attention model, I think.

But having said that, until today, the biggest deficiency is that the industry is very geared and ready for attention models. You know, everyone has custom kernels-

Host9:02

Yes

Barak Lenz9:02

... for attention, no matter how-

Host9:03

The hardware lottery

Barak Lenz9:05

... your, your hardware is. And so, so there are a lot of obstacles that are still existing for people to, to use these hybrid models. But if I had to guess, hybrid models are here to stay, just because the, the efficiency without sacrificing the performance is, is too much to, to give up.

You know, the attention is so expensive, the quadratic cost and the linear memory cost is so much that I think for long context use cases, for sure we're, we're gonna see hybrid models, and I'm happy that, that we've seen hybrid models come out that are not AI21.

So I think NVIDIA released hybrid models, and other companies started following. And since then, also a lot of other companies figured out that they don't need the full attention, even if they're not doing hybrid archit. So they have a cross between sliding window attention or-

Host9:53

Yeah. Which I think, uh, Noam, I was gonna bring up that Noam Shazeer published basically the same ratio of 1:8.

Barak Lenz10:01

Yeah.

Host10:02

Global versus local.

Barak Lenz10:04

I think in our ablations, we, we were trying different things, you know, not just the ratio, but where should we put it? You know, is it the middle of the eighth? It is the start, the end? And actually the middle-

Host10:15

Wow

Barak Lenz10:15

... worked much better, interestingly. It's not like we had time and resources to investigate everything, but putting it the first or the last performed worse than in the middle, and, and 1:8 was good enough. You know, y-you might get very slight improvements with 1:6, but, but it was marginal, maybe within the standard deviation.

And again, you pay for each transformer layer in your KV cache and quadratically whenever you wanna do something in long context. So, so Mamba has a larger fixed cost, like until K, you're slightly worse. But when you do wanna do long context, and I think you wanna do long context, certainly for agentic use cases, but and, and RAG, enterprise RAG use case and, you know, personalization, memory, w- there are a lot of different use cases, so I can definitely see sequence length rising, and I can't see full attention models being as prominent as they are today.

So, so, so at least they'll have less full attention layers, and I hope they'll have more innovations like Mamba. It's not like I'm saying, like, Jamba is the end of all, of all things, but I think it's, it's starting to become clear that you don't need full attention in more than one of eight or, uh...

At, at least this is how it looks empirically. Again, I'm not talking about this from the theoretical standpoint. We, we, we're examining it only from a results basis.

Host11:40

Yeah. So I, I'm gonna-- I like to do a little bit of visual aid, so I like to show people. Uh, if people want more information, they can go to the blog post where they talk, uh, about the rise of hybrid LLMs.

I like the reference to the original Black Mamba, uh, uh, in, in, from basketball. But yeah, I think there's a, there's a really interesting and strong history that's, is coming. So I mean, let's talk about the new model a little bit.

Jamba 3B12:02

Host12:02

When we release this, it will be out. So, you know, what, what are we talking about it? It's an efficiency play. It's talking about, like, a very long context. What other useful features, uh, would you highlight?

Barak Lenz12:12

So I think the most useful features is that it can really fit long context into edge devices or very memory-restricted environments. I think that that's the biggest feature because if you look, let's say, at artificial analysis mini models out there, none of them is, is hybrid.

Um, and, and just to, to, to, to clarify what that means, if you have a 3B model, I think something like 8 or 16K context length has the same amount of KV cache as, as a 3B model. So it gets into the-- it gets the same memory requirement.

So you c-you can't actually, even if you can fit a model in your edge device, you can't really put a long context. And I think people look at long context as something that's very rare, but, but if you think about images as an example, you know, even four images is already long context because usually when you turn images into tokens, it's a few thousand tokens per, per image.

So if you wanted to do something local on your phone to search your images, as an example, you can't do that without a hybrid architecture or, or without doing drastical changes because the model plus KV cache won't fit.

So I think this is the first really interesting thing about it, and we do want to release it with the entire stack that allows people to train it. Because as I see it, one of the reasons that it's hard for people to adopt Jamba is that it starts with-

Host13:38

Fine-tuning

Barak Lenz13:39

... with a large size. You know, Jamba Mini is mini for enterprises, but it's not mini for the, for a developer that, you know, has a T4, has his own GPU and he wants to try stuff. And I think it's also important to have a model that people could really try working with and see where it excels and, and where it doesn't.

Because it's not like we know everything about hybrid architecture, right? We want the, the, the open source community to be also be able to experiment with these kind of things, including long context RL training. And again, I don't know if this will be out in time for the podcast, but it will be out with the...

And, and, and I think if you're able to get a model, even if, and again, for 3B models, it's hard to get them, you know, perfect at all the tasks at the same time. It's hard to balance them.

But once you can get a training recipe out there that, that really works and, and within a few clicks you can get your own, let's say, long context or non-long context use case and run it with Jamba models, there's, there's a much better, you know, chances of you really using it.

So I think again, the, the differentiator fits long context in very small memory constraints, but it also allows you to, to experiment with hybrid architecture. While Jamba Mini, again, people that worked with Mistral could do it, but most people preferred Mistral just because it was very manageable, right?

A 7B dense model is much easier to work with than a 52B total params MoE. So this is a 3B dense with only two attentional layers. It's 1 to 12 and not 1 to 8 because we wanted to maximize the efficiency.

It has very few attention heads, so everything is geared to have, you know, long context with very little memory. And the fact that you could fine-tune it and do reinforcement learning on it means that you could get it to excel on your, on your task, even if the model that we release is not, is not perfect at it.

Because again, 3B models, they need to be tailored for tasks. It's not like you could squeeze whatever you wanted into them.

Host15:44

Yeah. Um, oh man, there's so many questions I want to ask. So, okay, when you experiment with architectures, what scale do you find it meaningful to experiment at, right? Because even 3B is a bit la-a bit large.

Scaling Training15:44

Barak Lenz15:57

No, I think that's, that's a great question, and it doesn't have a super short answer. I think most my answers are usually not super short. So we try to experiment with a minimum size that we can. Obviously, not all tasks respond to these kind of models in the sense that let's, let's just say tasks like MMLU Pro, which is like A, B, C, D, you'll get random.

So, so you won't really know what's working or not in too small sizes. So we did a lot of work in trying to find predictors that would scale to larger models, so, so, so that we'll be able to test things in smaller models and then get to the mo-the bigger model and make it work as well.

But these tests are not perfect. You can't predict everything. So, so we do try to do everything as small as we can for the velocity to test a lot of things, you know, data ablation, architecture ablations for a lot of things.

And sometimes you need to develop evaluations that will respond even in smaller models. So, so, so you can't do the, the, let's say, few-shot or zero-shot evaluations for everything. But if you're creative, you can find ways in which you'll see the response from, from smaller models.

So, so just as-

Host17:08

Yeah

Barak Lenz17:08

... an example, if, if my model didn't train on a certain sequence length and I test on it, I can expect the, the, the log prob to be, like, random. So even if I give the, my model a, a long sequence task, I may be able to tell the difference between a model that sees the entire context and responds in perplexities even before it is good at the task in, in zero-shot or few-shot.

Does it make sense to you? So, so, and sometimes you need to be creative with evaluations and find the right evaluations that will scale with size, but it's never gonna be bulletproof. So it's not that I think you can craft the perfect, you know, one trillion model architecture just by testing it on 100 million parameters, but you can gradually increase in size and then change your test, and this is usually how we, how we do the architecture, so.

Host18:04

But it, it, it's basically like a exponential step up, right? Like 100, 200, 400-

Barak Lenz18:08

Yeah, you can look at it something like this. I would, I would look at it-- I would call it the other way. It's like a log space thing, right?

Host18:14

Log, log space.

Barak Lenz18:15

So you reach sizes in the log space, which is exactly like what you said with an exponent. But, so, so you'd go from 100 million to a billion to, to let's say seven billion 60-

Host18:27

I see

Barak Lenz18:28

... and then, you know, the 1 trillion

Host18:29

Your scale, your scale is much bigger than I'm used to.

Barak Lenz18:31

Well, you know, frontier models, scale still matters. So I think, you know-

Host18:35

Exactly, right. Emergence is a thing.

Barak Lenz18:38

Yeah. And, and size is still a thing. It's not as big a thing when you have open source models that have this size, so you can distill from them, and you can do a lot of things because they already exist.

But Jamba Large is a big model, and it, it takes a lot of effort to train. And again, this is still very far from commodity training large, because they, they have different problems than smaller models. So in all the large models that we trained, we also had to deal with optimization problems.

Host19:05

Yeah. Can you share, uh, I, I just realized I don't really know. What is less publicly known, and we can scrap this if, if it's not, if it's a sensitive question. What is publicly known about your team size and, y- you know, amount of GPUs that, that you, that you guys have?

What does it take to, to run a frontier lab these days?

Barak Lenz19:21

I can tell you that the models themselves, you need thousands of GPUs to train. And in terms of team size, it varies. It depends on your team. You know, there are individuals that can do a lot of things, and I don't think team size is, is what we should look at.

I think it's more the velocity of ablations. So if I had to, in a way, dodge your question, but answer it-

Host19:46

Sure

Barak Lenz19:46

... as well. So, so you need-

Host19:48

That's fine

Barak Lenz19:48

... thousands of GPUs, and you need to do a lot of different ablations. How you do them and how you get to a point when you can try a lot of different things is, is how you manage your, your teams, right?

So it's, again, it doesn't have to be a huge team, but you need to test a lot of, of different things. And, and so you need an infrastructure that's, that's built for that, and, and we're using our own infrastructure to train our models.

There still isn't an open source infrastructure that I could say, "Use this to train your very, very large model." I'm guessing it, it will happen as well. I'm not, I'm not putting this as a... But, but there are a lot of different parallelisms that you need to implement, a lot of different things that you wanna try out and-

Host20:31

Yeah, you know, like, uh, some people, like the DeepSpeed team from Microsoft, like, uh, is, like, doing good stuff. I think, uh, even the Imbue team has, has open sourced quite a lot of their stuff for, uh, for-

Barak Lenz20:42

Yeah

Host20:42

... training as well.

Barak Lenz20:42

Yeah. Yeah. But, but there's still a difference between taking a recipe that works, let's say in a specific version of Qwen in, I don't know, Verve something, and then getting it to really work in a robust way with any model, with your framework of how you do evaluations, how you do rewards, how you...

So you need to have very, very good engineering. I think one of the things that people don't get is how important is good engineering. So algo is important obviously, but you need world-class engineers to...

Host21:15

What, what's an... Uh, okay, uh, since, since you're, you know, CTO and all, I'm sure you're always hiring. What is the kind of a problem that, you know, if you're, if one of our listeners is an engineer, uh, and they, they are, they're, they're know the s- they know the solution, you would hire them on the spot?

Hiring Philosophy21:15

Barak Lenz21:29

Wow. Again, that's a good question, but it's, it's very far from me because the way that I interview is that I, I don't have a fixed set of questions. It's just an open conversation. And I try to-

Host21:40

Oh, man

Barak Lenz21:40

... ask about, you know, because, uh, I think the, the, the mystery for me is that I know people prepare for interviews for a very, very long time these days.

Host21:49

Yeah.

Barak Lenz21:49

And I wanna be able to ask them things that they couldn't prepare for. I wanna know really who they are, and, and I'm looking for people that try to solve problems on their own. Because if you think you're gonna find the answers online and think you're, you're not, uh, in the frontier.

So I think this is the main characteristic. I want people that know how to work on their own, that are not afraid of very, very tough problems and have a can-do mentality that's like, "I don't care if no one did Jamba before.

Why couldn't we be the first ones to do it?" Because, because if, if, if in everything you say, "Let's look at the references, let's look what other people did," and, and that, that usually works less well for me.

But, but again, that's a personal preference. It's not like in company. Opinions are my own. Let's put it like this. I'm not, I'm not, I don't claim to be an expert.

Host22:42

That, that is a valid opinion. Uh, you know, the, the pushback I would say as a podcaster who interviews a lot of founders is you, you, you help yourself by mentioning, like the first question, uh, which gives you the sort of embedding neighborhood, right?

Or the first principle component, uh, but you don't review all the rest, right? That, that actually comes in interview. But you attract the right talent by, by having the strongest message about, like, what is, what is the rough neighborhood of people that you're looking for, and here's everyone else that works with us.

So, like, you, you, like, that, that is the recruiting message that I'd like to help people with.

Barak Lenz23:15

So I don't have it as crisp as, as, as you would like it. Everything is a little longer for me, but I can say that one of the first thing that I really wanna know is how would you like to spend your day?

Let, let's say I gave you an ideal day, you know, day of your choice. How would it go? You know, how, how much reading would you do? How much coding? How much would you work with other people? How, how much do you prefer advancing on your own?

And I think this tells me a lot about because you do want team players, team players are super important, but it's also super important to have people that are not afraid to, to, you know, if they're stuck at something, push through until they're successful on their own.

And again, the way you want to distribute your time in, in your ideal setting tells me a lot about who you are and, and what you really wanna do, right?

Host24:08

Yeah. Yeah. Totally.

Barak Lenz24:11

Let me-

Host24:11

I mean, I, I think one, one last... Sorry. One way of... Yeah, everyone has the same 24 hours a day, and, uh, the time is the one resource you will never get more of, right?

Barak Lenz24:20

Right. Right.

Host24:20

So it's the most, it's the most restricted.

Barak Lenz24:21

So how you ideally spend it tells me a lot about what you wanna do.

Path to AI24:25

Host24:25

Yeah. Uh, I think one way, one way I also tie it back to, like, the, the sort of technical discussion is, like, kind of your personal journey into AI. Like, uh, you don't have, you know, a PhD in AI.

Like, you, you, you come from more of like a sort of CTO, a startup CTO background. And, like, I guess how... what was, like, the critical points for you in ramping up on, uh, just AI in general and, and s- sort of LLMs more specifically?

Barak Lenz24:49

I don't, I don't think I have a great answer for you, but I'll try. So be- before joining AI21, I was actually doing algo trading for, for something like eight years. And I was doing algo trading because I wanted to do something that the bottom line is me, my algo and engineering.

It doesn't have sales and marketing and, you know. It's just you and the numbers, and I've always been a numbers person. I, I like numbers and, and we can talk to each other. We, we understand each other. So, so I've been doing algo trading for a long time, and then I decided that that's not the lifestyle that I wanted, mainly the fact that the trading live in US hours while living from Tel Aviv was not the right lifestyle, et cetera.

And then, you know, I, I, I met Ori, like, 20 years back and, and he talked to me about th- this opportunity. He, he knew me, you know, from, from before and, and he knew-- he thought I had the right skill set to do it.

And in a way, AI is very close to algo trading if, at least from my perspective-

Host25:47

Yeah. There's a lot of, uh... I'm also ex- ex hedge fund. There's a lot of finance people.

Barak Lenz25:52

So, so it's very, it's very similar in terms of how you interpret results. You wanna treat everything as a black box. You wanna establish your bounds, you know, what are you... And, and I've, I've had tons of experience in algo trading, both from making money and not making money and, and figuring out what my mistakes were, right?

And, and how does something that really has a predictive power looks, and I think AI looks like it. So you asked me about ablations for, for a new foundation model. In a way it's, it's very similar, at least for me, to, to how you develop a new, a new trading algorithm.

So to some extent, the transition was smooth in that sense, but in, in terms of knowing the material, there were tons that I had to know, but that doesn't bother me. You know? Learning new things is always fun, and I was very lucky to start early on.

So, so it took me more than a year to really feel like I'm starting to grasp, to get a grasp of it, but it was still 2019. So people were not, you know, on this ride yet, and then they were, you know.

So, so there are a lot of people interested in it now, so I, I'm lucky to have a few years, you know, starting a few years early, making my transition from algo trading to this. But, but again-

Host27:05

Yeah. Amazing

Barak Lenz27:05

... I see it as very closely related algorithmically and very different technique.

Host27:10

Yeah. The, the comment I'll make there is that it's, it's a lot of, uh... in, in both ways you're kind of looking at correlations, and you're trying to think about causation.

Barak Lenz27:18

It's correlation, cau- causation, but then you go through the deeper derivatives, the covariance matrices of how you balance-

Host27:24

Right

Barak Lenz27:24

... the entire portfolio and... Right? So, so you need-- But yeah, I, I think if you had to put it in a line, you're, you're looking at correlation and trying to say, "What is the causing factor, and can I find it?"

Host27:37

I would say it's much easier, uh, as a, as a former trader because there's no regime change, right? The, the data is the data. It's, it's static.

Barak Lenz27:45

You're not looking for breaking points in the data and, and thinking about algorithms that could say, "How can I predict, you know, the Swiss government moving the, the Swiss currency by 20% overnight?" Which is obviously not predictable unless you had-

Host27:59

Most annoying thing is you have a strategy that works, but then, like, it no longer works going forward. It's not because you made any mistake at all. It's just-

Barak Lenz28:06

Right

Host28:06

... the, the underlying rules change, and you had some secret data in there that was, like, exposed to something else that w- that you didn't know about.

Barak Lenz28:11

But that's the nature of every algo trading strategy, to lose alpha. Have a theta on the alpha, let's put it like this. You know, the alpha decays over time, like in options.

AI Systems28:22

Host28:22

Yeah. Oh, totally. Okay. So, so not to get too finance-y, uh, I also just wanted to bring it back to close on the idea of the AI systems. I think the way that you use it, I think we would say in the US, we, we'll say harness for agents, uh, but maybe there's a difference.

So, so what is your definition of AI systems? Why is it so important? Why are you so passionate about it?

Barak Lenz28:43

So I think just, just calling a model, in a way, I, I, I've still yet to have seen a model that gets a perfect score at anything, literally anything. So you can train a model for a long time, but it will still not be perfect.

It's... We sample it with a temperature, and we do a lot of different things, but a model is not a piece of code. You know, a piece of code, you do A plus B, you always get the right answer.

But whatever model you take, give it the right A and B, and you're not gonna get the right answers. So I think that's one reason to, to move to systems, because you want reliability, and models are not very reliable, I think by design.

They're, they're statistical beasts, and we sample them statistically, so they're not meant to be reliable. Um, and you need AI systems, first of all, to get reliability. And I think when you look at enterprise use cases, it gets pretty clear that it's very hard to do it with a model because you weren't trained on the data that you need.

And even if you were trained on tool usage, it's not like showing the model, you know, a description of where my, my enterprise documents are stored is gonna get it to get a very good score in trying to get data from them and, and, and, and answering correctly, et cetera.

So, so in order to get the reliability and robustness that you really want, I, I think you don't have a choice but moving to AI systems. And in a way, this move is already happening. It's just under the hood a lot of times.

So GPT-4o is already an AI system. It, it's not calling a model directly. It can do certain, you know, it can do tool calls. It can orchestrate this entire thing, like a, a simple API that I give it max tokens and temperature, right?

And And I think the move there is necessary because again, you need tools to compensate for the model even without the enterprise. You definitely need Python and then web search to, to compensate. But when you look at enterprise use cases, it gets way more complicated.

You want to connect to a lot more systems, and you want to do things that weren't necessarily there in your pre-training or in your fine-tuning even, right? So, so models that you fine-tune are trained on the task that, that you train them.

It's not like they're gonna ace tasks that they've never seen. Maybe that's an impression that some companies would, would wanna make, but at least that's not my experience. So it's not the models don't generalize, but it's not that, you know, I, I taught it base 64 and then it can do any other encoding, right?

Maybe that's not the best example, but-

Host31:16

It can do, it can do a bit of in-context learning, but yeah, you s-- it still has to be somewhat within the trajectory.

Barak Lenz31:20

Yeah, but the in-context learning is just part of it. There's a reason we train our, our models. So, so i-it's good in the trajectories that we trained on, but then at inference time, it's just trying to mimic trajectories that it has seen.

If you talk about, you know, an LLM or even a, a reasoning model, right? It's just trying to... And, and AI systems, at least in my perspective, they don't, they don't just try to mimic what they've seen. They try to rigorously answer your question.

Okay. So, so when, so when I say rigorously, I think AI systems need to be model agnostic. They shouldn't care about which model that they, they use, and they should look at what I call actions, which is a combination of a model with a prompt and maybe a set of tools that it can use, and say, "What can an action do for me?"

And then look at it as a black box statistic. Say, "What is the average value that I'm gonna get from this action? What is the average cost that I'm gonna get from these actions?" Right? And then say, "How can I do this in the most cost-efficient way?"

And I think cost is a major factor in AI systems. You see it a little bit now with reasoning models with token budget. But I think in Maestro, we show that, that it's more important. And again, there's a lot of papers coming out recently that are trying to balance with this budget, with performance as well.

So, so for an AI system to really work, you can't just say, "I'm gonna use the best, most powerful model," first of all, because this model doesn't exist. The best model only has certain languages that it supports and, you know, it's not necessarily the best in vision as well as text and other modalities.

So, so I think my assumption is y-you're gonna use a lot of different models, and you need to use a lot of different models. Some of them may be proprietary for a specific use case, other can be general, and you wanna be very cost-effective at doing this.

Otherwise, the latencies, which is most enterprises don't really wanna use reasoning models. The latencies is too high. But maybe for some use cases you wanna use them, but figuring out exactly why and how to use them is something that the system needs to train.

The system needs to be able to learn when should it use different actions and what's the expected value that I can get. And then we can dive deeper into the things we said about algotrading, you know, the covariance as well.

But, but first of all, the first tier, costs and values and, and different correlations between them. And, and the correlations are also important because if you wanna do a majority vote that's really significant, you wanna make sure that the different, let's say, models that you use for majority really give you different answers.

So I think it, it connects to a lot of very fundamental statistical things that in other fields people are alread-already doing, bounds and bandits and... But I think these things, this is where, where it's going. This is where we as a company are also trying to take it, is, and I guess other companies would try as well.

So AI systems that are model agnostic, they don't suffer from the illness that we have today. I think mainly two illnesses, I would say. One of them, there's no continuous training. And in an AI system, you can pretty easily see how you can continually learn because you're not coupled to models, right?

You're just estimating their values and costs. So you can scale with the number of, of, of models in action. And then the life cycle, which I think is again, a big pain point, you know. A big company comes with a new version.

Should I switch to it or not? The answer right now is, "I don't know. Let's do all of our tests again and see if it's really an improvement." But in an AI system, this looks very different.

Host35:09

But there's a, there's a blog post on planning actions, not predicting tokens, which I think is-

Barak Lenz35:13

Right. Right

Host35:14

... a be-a better elaboration of what you just said.

Barak Lenz35:16

Right. And we also have blog posts about Maestro, and, and Maestro is... I think it's super cool, and I think people can, maybe not today, but by, again, by the time the post does comes out, they should probably be able to play with Maestro on their own.

And Maestro is already an AI system, still early version, but it's already an AI system in the sense that you can give it requirements, and it will try to optimize them while being model agnostic, doing internal planning and internal search space, and you can give it budget.

But again, the budget won't be in tokens. The budget is actually in the total latency and cost that, that you will have. So we haven't had tons of times to talk about this, but it's not just thinking in actions, not tokens.

There's also Maestro, uh, material out there, and at least enterprises should be able to, to try it out. I don't know how open it will be for anyone to try out, but, you know, model agnostic system that does things more explicitly, tries to enforce what you want it to enforce.

Host36:18

Yeah. I think the main challenge I have with these, with sort of, um, the RL search, uh, space is that there's a, there's a lot of ... issues with calibration of models on how much they know that they don't know, therefore, they need more information, right?

Uh, I think that's like the core problem to solve because I, I always describe this as basically like there's a fog of war and sometimes you have to take an action to discover more information in order to take action.

Barak Lenz36:47

Right.

Host36:47

So then there's the implied value of information that you have to go and get before you actually can take any action at all. And I think, like, there's definitely a lot of, like, experimentation around this. Right now people are biasing towards just greedily grab as much information as possible, but actually that's very expensive.

Barak Lenz37:02

Right. But I think the right approach is a bandit approach and I'm, I'm not gonna dive into exactly how you do this bandit approach, but I think bandits literally tackle this exact issue. You know, what's your exploration? How do you get to estimate your values?

Costs are pretty straightforward to, to estimate. And I think the right approach is a variant of a bandit approach. And again, brute forcing is not the way. And I think you can already see this in how companies are doing RL today.

You know, the RL is very brute forcing. The number of rollouts per example is fixed and- ... the data is not really changing while you train although some of the data is already saturated and you just waste all of these rollouts.

And, and once you've trained a few hundred steps of let's say GPO, most of your training is just wasted on example that are either too hard for you and you didn't get any success on them or too easy and everything was a success.

You just use a group filter to get them out. So I see a, a, a direct connection between how you wanna do these rollouts and what Maestro needs to do for you at inference time. Because the problem already exists when you wanna do these rollouts.

And right now if, if your pockets are big enough, you can just brute force these rollouts and say like, "Oh one, we're doing 64 rollouts for each input with like 20K max tokens on a huge model." But if you look at the price tag for that, this is not something that anyone can use to adapt to their enterprise use case.

But if you are much smarter in those rollouts and much smarter and looked at cost as a principal component, not just an afterthought, things will look very, very differently. And I think people have different requirements. So cost is not just to save you money, but it's also to allow you to say, "I wanna spend more."

Or "I wanna check every little detail in this and, and the fact that some webpage cited this and say this is true is not enough for me. I want you to keep on checking things because I want a, a higher certainty."

So, so the controllability that you want a user to have is something that I, I don't see enough these days, right? Because it's like a one size fits all thing. While I may only ask a casual question when I want you to quote some answer from the internet, but sometimes we were talking about the Black Mamba.

Sometimes I'm going to ask you a question and for you to calculate the real data. You know, how many, you know, baskets did, did Kobe Bryant- ... score in the last three minutes of away games? I don't think I'm gonna find this, but this is calculatable in today's technology.

It's just no one's doing this because you'll need to go over a lot of different games and do aggregation and do things that are not very natural for models, but are certainly within reach. And I don't wanna do this for you whenever you ask a question, but I do think if you want to dig deeper, you should be able to and, and this is part of what AI systems could, could, could bring you.

And just one thing, I know I, I might sound all over the place, but just, just as an example, just if you look at the difference in behavior between Groq and, and let's say Claude and OpenAI's models, you could see why in a way an enterprise would want a system and give them their own policy.

Right now if you're using the models, you're taking in their own policy. Even if I want to use GPT OSS, I've, I've taken in a lot of different policies about what to abstain from, what's considered dangerous and not dangerous, how I should behave and et cetera.

And, and I don't think model makers are the right people to, to make these calls. I think we need to provide a technology that allows you to define your policies and enforce them as well as Claude and OpenAI do it for their policies.

But this is not something that the, at least from the best of my knowledge, is something enterprises can do today with-- how-- without very heavy machinery. Because you don't just need GPO, you also need judgments on the GPO and...

Never mind. I won't go into all the technicalities, but it takes a lot to really get a model to follow a policy as well as being accurate on code and math and, and all the other stuff.

Host41:28

Yeah. I totally buy that. Uh, okay. So y-you've been really generous with your time. We're, we're a little bit over time, but, uh, I, I think it's just very encouraging, um, how you're, you know, viewing progress. I, I guess maybe the, the final question I'll leave you with is like, uh, you know, if you were to have your personal hat on, um, where, where does, where does AI21 go in, in like a year or two years?

Like, like, what's the overall direction that people should be listening out? Is it more enterprise focused, more on device? Like where, where is, where is it going? More systems?

Future Direction41:50

Barak Lenz41:58

There's one thing I learned in algo trading is that given predictions that are more than two days in the future has more standard deviation than-

Host42:05

My God.

Barak Lenz42:07

But, but having said that, I think AI21 is, is heading for the enterprise with AI systems with, with Maestro and I think, I think this is where the industry needs to go. I think enterprises need-

Host42:19

So Maestro model agnostic including other, other people's models, right? Like you have-

Barak Lenz42:23

Right. But also, but also including your own models and the ability To get them to be the best you can in your specific use cases. So it's not that I'm saying we don't need to do RL in training.

I'm saying we do, and we need this to happen in one system. This is why I was referring to continuous training. As I see this, Maestro can be model agnostic, but also gather all the information and data from successes and failures and be able to do better the next time by also fine-tuning and RL'ing models.

Whether or not, you know, AI21 models or not, I, I do think that you'll need models with different attributes, and Jamba is one of them. If you wanna do long context, you better do it on a hybrid. If you wanna do it in the edge, you need a model that, that fits on the edge.

But I think enterprises would want continuous improvement. They, they would wanna see how by trying out your use cases and understanding when I did well and not well, I could in some way incorporate the signal to get better, either by selecting better actions like, like in the bo-- but also training my actions when I'm missing actions, action that's either as good as I want or as cheap as I want.

But I, I wanna have this option, and I think, again, I think it's important because most enterprises will have use cases that would-- they will also want to, to use this. So, so I'm-- I don't think model training and, and, you know, fine-tuning, RL'ing, whatever, is something that you shouldn't do.

I think that should be a part of one integrated system that also have long-term benefits, not just you use it, and then whenever there's a new system, you're gonna throw it out. You're gonna learn from your experience and get better at specific actions and that ex-- and in exploring your action space.

Host44:16

Okay. I love it. Well, I think we'll just leave it there. Th-th-thank you so much for joining us, and, uh, all the best on your launch and, uh, I'll be keeping watch as well as, as you, as you roll out Maestro.

Actually, I, you know, I wanna try it, so.

Barak Lenz44:27

Yeah, yeah. I think, I think it's, it's, it's super cool and, and I'm excited. Now, hopefully, this, this podcast went, went well. I, I don't know how you, you see this. I know I've ranted on so many topics.

Host44:40

Rants are good. My job is to get your best rants, you know?

Barak Lenz44:43

It's too, it's too easy to get rants out of me. That's, that's my problem. But- ... but anyway, it was a lot of fun, and thank you. Thank you for having me, and then maybe we should do it again sometime.