AlphaFold2 Revolution0:00
Actually, we only trained the big model once. Uh, that's how much compute we had. We could only train it once. And so, like, while the model was training, we were, like, finding bugs left and right.
Yeah.
Uh, a lot of them that I wrote. And, like, I would-- I remember, like, us, like, sort of like, you know, doing, like, surgery in the middle, like stopping the run, making the fix, like relaunching, and, um, yeah, we never actually went back to the start.
We just, like, kept training it with, like, the bug fixes along the way- ... uh, which was, uh-
So it's impossible-
Yeah
... to reproduce now.
Yeah. Yeah, yeah, no, that model is, like, has gone through such a curriculum that, you know, it's learned some weird stuff. Uh, but, uh, yeah, somehow by miracle it worked out.
It's a pleasure to have with us today Gabriele Corso and Jeremy Wolwen. They are-- They recently founded Boltz, a company trying to democratize and bring our structure prediction in biology to, you know, the masses. Uh, they are both, uh, recent PhD grads from MIT and have been working on all sorts of foundational papers in, like, generative biology.
Um, anyway, uh, pleasure to have you here. Thanks for coming.
Thank you.
Thank you.
Uh, I guess we're maybe, what, six years post AlphaFold2 right now, which was, like, kind of a big moment. Um, is that right? Yeah.
I think, what was it? Twenty twenty-one.
Yeah.
So yeah, on-- going on five years.
Five years.
Yeah.
Five years.
Yeah, yeah.
Um, yeah, so maybe for the audience, like, can, can-- Let's go back to that moment in time and explain, like, what was this big moment, and why was it interesting? Why was everyone so excited? And I think you two were probably quite excited, so why were you personally excited?
I would start on kind of why that was interesting, kind of, you know, from a scientific standpoint. So, well, AlphaFold-- So maybe first as a kind of introduction for, uh, the ones in the audience who are not structural biologists.
So the idea of structural biology is that, you know, we want to try to understand how, you know, proteins and other molecules take shape inside our cells and, you know, how they interact. And structural biology is sort of this beautiful discipline, uh, where we are somehow able to understand this minuscule structure at, you know, kind of atomic details using, uh, these incredibly, um, complex methods like, you know, X-ray crystallography.
And, you know, the, the dream has always been of computational biology, can we understand kind of the structures without having to, you know, resolve this crystal, you know, shoot X-rays, and so on. And so AlphaFold, um, was a real breakthrough in this problem of protein folding, which is trying to understand the structure of a single, uh, protein.
And to me, it was exciting across kind of many dimensions. One, I was a computer scientist, I was working a lot on machine learning, and I saw kind of the impact that kind of the work similar, somewhat similar to what I was doing could have on, like, a long-standing scientific problem.
And on the second perspective, from a more, you know, personal side, the seeing kind of the structures coming out of these models where, you know, you see kind of this beautiful, you know, um, creation of life is something that was, was very inspiring to me.
And so that was kind of one of the things that led me to start, uh, working on, uh, structural biology and in particular with machine learning.
Were you a structural biologist before AlphaFold came out? I mean, did you-- You did machine learning, but it was not in structural biology, so that actually shifted your career quite dramatically.
Yeah, very dramatically. I was, I was working on some pretty kind of theoretical methodological things, and I was starting to see kind of, you know, some of the challenges in, you know, kind of doing somewhat theoretical or methodological work and, you know, seeing kind of the potential impact of, you know, um, doing excellent, you know, AlphaFold was really a machine learning breakthrough, but, you know, an applied machine learning, and so that led me to, uh, want to start working in applied ML.
Our, our group at the time was, um, working a lot on, like, small molecules already and, um, I think AlphaFold is kind of what triggered, I think, this shift to, like, working on, on biologics. Um, and at the time, I think it, like, opened as many questions, you know, as it answered in a sense.
Like, we, um... the immediate follow-ups were, "Okay, like, can we do this on other things than proteins? Can we do, um, you know, interactions of small molecules with proteins, nucleic acid with proteins? Can we model more complex protein systems?"
And I think, yeah, very rapidly, I think after AlphaFold, uh, people realized, I think, that there was, you know, machine learning could have a-- could really, yeah, sort of target this problem very differently than, you know, than previous methodologies.
Clarification. So what, what does small molecule mean? What does protein mean? What is-
Yeah
... you know, the terms that you just mentioned.
Yeah. Maybe we can start with protein. Um, so, you know, protein is, it's maybe the most fundamental one. It's what gets decoded out of our DNA.
Mm-hmm.
Um, it's, uh, essentially a sequence of, uh, amino acids. Each amino acid you can kind of consider as a, what we call a small molecule. Um, and there's twenty of them in the-- at least in the human body.
Um, and, you know, any compositions of these twenty amino acids in a sequence, um, you know, creates a, a different form of a protein. Um, and so, you know, obviously there are a very large number of those sequences that you can create.
Um, small molecules are sort of- You know, following the name, um, typically considered to be, um, you know, a much smaller number of atoms. Um, and the atoms that compose them, I think, are also generally a bit more diverse, right?
I mean, amino acids have, um, you know, this composition, and it's always the same. Uh, with small molecule, you know, there's a, a larger set of possible atoms that also we have to consider that also make the problem, uh, pretty challenging.
And then we have nucleic acids, so DNA and RNA, uh, which are also very interesting to model the structure for. And those a little bit more similar to proteins. You know, they're sort of composed of four, um, uh, nucleic acids, and you form sequences from them.
And, um, any, uh, codon, which is, like, three, uh, nucleic acid, um, translates into a specific amino acid. Um, so yeah, different forms of molecules. At the end of the day, just a bunch of atoms, uh, you know, that are bonded together, uh, that we try to understand the interaction of.
Going back to the AlphaFold2 moment, like, um, I remember this very well. I was at, uh, NeurIPS when I guess the results of this famous competition came out. So, um, can one of you talk about CASP and, like, what it is and why it is-- it was so interesting and exciting?
Yeah, I think every... So, um, um, every couple of years, um, I think the goal has always been to, you know, find, uh, protein structures that are a little bit different from what's known. So CASP over the years has, like, you know, put in a lot of effort to, like, gather structures from, you know, academic groups and, uh, even industry groups, uh, to try to create sort of a test set that would be, um, difficult, um, for, uh, different methods.
And CASP, uh, 14 was when, uh, AlphaFold2 really, you know, blew everything out of the water. Um, and the, the improvement was so large over, you know, the previous, previous method and also over the previous competitions. Um, and now CASP continues.
You know, we've had CASP 15, we have CASP 16, and, you know, sort of what's happened now is that it's really expanding to also all these other modalities, like I was mentioning, like protein, small molecule, nucleic acid, and, uh...
But the goal remains to, like, you know, really challenge the models. Like, how well do these models generalize? And, you know, we've seen in some of the latest CASP competitions, like while we've become really, really good at proteins, especially monomeric proteins, um, you know, other modalities still remain pretty difficult.
And it's really essential, you know, in the field that there are, like, these efforts to, um, you know, to, to, to gather, um, you know, benchmarks that, that are challenging. So keeps us in line, you know, about what the models can do or not.
Yeah.
Yeah.
It's interesting you say that. Like, in some sense, CAS-- you know, at CASP 14, a problem was solved and, like, pretty comprehensively, right? But at the same time, it was really only the beginning. So can you explain, like, what was the specific problem which argue is solved and then, like, you know, what is remaining, which is probably quite open?
I think, I think we'll, we'll steer away from the term solved-
Solved.
... because we have many friends in community who get pretty upset at that word, and I think, you know, fairly so. Um, uh, but the, the problem that was, um, you know, that a lot of progress was made on- ...
um, was the ability to predict the structure of single chain proteins. So proteins can, like, be composed of many chains, and single chain proteins are, you know, just a single sequence of amino acids. And, uh, one of the reason that we've been able to make such progress is also because, um, we take a lot of, uh, hints from evolution.
So the way the models work is that, you know, they sort of decode a lot of hints, um, that, that comes from evolutionary landscapes. So if you have like, you know, some protein in an animal and you go find the, uh, similar protein across like, you know, different organisms, uh, you might find different mutations, um, in them.
And as it turns out, if you, uh, take a lot of these sequences together and you analyze them, you see that, uh, some positions in the sequence tend to evolve, um, at the same time as other positions in the sequence, sort of this like, uh, correlation between different positions.
And, um, and it turns out that that is, uh, typically a hint that these two positions are close in three-dimension. So part of the, you know, part of the breakthrough has been, like, our ability to also decode that very, very effectively.
Uh, but what it implies also is that, you know, in absence of that co-evolutionary landscape, the models don't quite perform as well.
Mm-hmm.
And so, you know, I think when that information is available, maybe one could say, you know, the, the problem is, like, somewhat solved from the perspective of structure prediction. When it isn't, it's, it's much more challenging. And I think it's also worth also differentiating the, um, sometime we confound a little bit structure prediction and folding.
Folding is the more complex process of actually understanding, like, how it goes from, like, this disordered state into, like, a structured, like, state. And that I don't think we've made that much progress on. But the idea of like, yeah, going straight to the answer, uh, we've become, uh, pretty good at.
So there's this protein that is, like, just a long chain, and it folds up.
Yeah.
And, and so we're good at getting from that long chain in whatever form it was originally to the thing, but we don't know how it necessarily gets to that state.
Yeah.
And there might be intermediate states that it's in sometimes that we're not aware of.
That's right. And, and that relates also to, like, you know, our general ability to, um, model, like, the different... You know, proteins are not static. They move. They take different, uh, shapes based on their energy states. And I think we are also not that good at understanding the different states that the protein can be in and at what frequency, what probability.
Yeah.
Um, so- Uh, I think the two problems are quite related in some ways. Um, so yeah, still, still a lot to solve. Um, but I think the-- it was-- Yeah, I, I think, I think it was very surprising at the time, you know, that, uh, even with these evolutionary hints that we were able to, you know, to make such, such dramatic progress.
So I wanna ask, why does the, you know, sort of like intermediate states matter? But first, I kinda wanna understand, why do we care what proteins are shaped like?
Yeah. I mean, the proteins are kind of the machines of, uh, our body. You know, the way that all the processes that we have in our cells, you know, work is typically through proteins, sometimes other molecules sort of intermediate, you know, interactions.
And through that interactions, we have all sorts of cell functions. And so when we try to understand, you know, a lot of biology, how our body works, how disease work, we often try to boil it down to, okay, what is going right in case of, you know, our func-- um, normal biological function, and what is going wrong, uh, in case of the disease state.
And we boil it down to kind of, you know, proteins and kind of other molecules and their interaction. And so when we, uh, we try predicting the structure of proteins, it's critical to, you know, have an understanding of kind of those, those interaction.
It's a bit like, um, seeing the difference between having kind of a list of parts that you would put in, uh, in a car and seeing kind of the car, uh, in its final form.
Right.
You know, seeing the car really helps you, uh, kind of understand what it does.
Yeah.
Uh, on the other hand, kind of going to your question of, you know, why do we care about, you know, um, how the protein folds or, you know, how the car is made-
Mm-hmm
...uh, to some extent is that, you know, sometimes when it-- something goes wrong, you know, there are, you know, cases of, you know, proteins misfolding-
Mm-hmm
...in some diseases and so on. Um, if we don't understand, uh, this folding process, we don't really know how to, uh, intervene.
Okay. And so do proteins when they're m-- in the body, do they-- are they typically in that folded state, or are they kinda just like, you know, doing whatever until they're in a location where they need to interact with something?
That's a great question. Uh, and it really depends on the protein. Uh, it depends on basically the stability of the protein. There are some proteins that are very stable, and so once they are produced, you know, from the ribosome, they sort of fold into shape.
Then more or less they keep that shape with, um, minor variations.
The ribo- ribosomes is part of the cell that actually translates and, and turns DNA to RNA to proteins.
RNA to proteins.
RNA, RNA-
That final part-
Yeah
...of RNA to proteins.
Sorry. Yeah.
And so once they come out, they're pretty stable. Uh, but then on the other hand, there are some that, you know, for example, have multiple states that they switch to depending on their environment, you know. Uh, the bi-biologists really figure out some incredible machines.
Uh, there are machines where, you know, proteins where, you know, depending on whether, for example, another molecule is present or not, they will take different shapes, and that different shape will give it a different function.
Mm-hmm.
And so we have this, you know, so-called fold switching, uh, proteins that take multiple... And we have some proteins that are completely disordered, and these disorder proteins are actually pretty important in kind of many diseases, and those are kind of ones of the ones that we have, you know, the least understanding of.
There's this nice line in the, um, I think it's in the AlphaFold2 manuscript, where they sort of discuss also like why we even hopeful that we can target the problem in the first place, and then this, this notion that like, um, well, for proteins that fold, um, the folding process is almost instantaneous-
Mm-hmm
...which is a strong like, you know, signal that like, yeah, like we sh- we might be able to, um, you know, predict that this very like constrained, uh, thing that, that the protein does and so quickly. Um, and of course, that's not the case for, you know, for, for all proteins, and there's a lot of like really interesting mechanisms in the cells.
But, um, yeah, I remember reading that and thought, "Yeah, that's somewhat of an insightful, insightful point." Um, yeah.
I think one of the interesting things about the protein folding problem is that it used to be actually studied, and part of the reason why people thought it was impossible, it used to be studied as kind of like a classical example of like an NP problem.
Yeah.
Uh, like there are so many different, you know, type of, you know, shapes that, you know, this amino acid could take. And so, uh, this grows combinatorially with the-
Yeah
...size of the sequence. And so there used to be kind of a lot of actually kind of more theoretical computer science thinking about and studying problem, problem, protein folding as an NP problem. And so it was very surprising also from that perspective, kind of seeing machine learning so clear.
There is some, you know, signal in those sequences, uh, through evolution, but also through kind of other things that, you know, us as humans were probably not really able to, uh, to understand, but that these, uh, models have, have learned.
And so An-Andrew White, we were talking to him a, a f- a few weeks ago, and he said that he was following the development of this and that there were actually, uh, ASICs that were developed just to solve this problem.
So, um, yeah, that like-- and that there were many, many, many millions of computational hours spent trying to solve this problem before AlphaFold. And just to be clear, um, one thing that you mentioned was that- Uh, there's this kind of co-evolution of, um, mutations and that you see this again and again in different species.
So explain why does that give us a good hint that they're close by to each other?
Yeah. Um, like, think of it this way that, you know, if I have, you know, some amino acid that mutates, um, it's gonna impact everything around it, right, in three dimensions.
Yeah.
And so it's almost like the protein, you know, through several probably, you know, random mutations in evolution, like, um, you know, e-ends up sort of figuring out that this other amino acid needs to change as well for the structure to be conserved.
Right.
Uh, so this whole principle is that the structure is probably largely conserved, you know, because there's this function associated with it.
Mm-hmm.
Um, and so it's really sort of like different, yeah, different, um, different positions compensating for, for each other.
I see. So the, the-- those hints in aggregate kinda give us a, a lot of information about what is close to each other, and then you can start to look at what kinds of folds are possible given the structure, and then what, wh-where, where, what is the end state, and therefore you can make a lot of inferences about-
Yeah
... what the actual f-total shape is.
Yeah, that's right. It's almost like, you know, you have this big, like, three-dimensional valley, you know, where you're sort of trying to find, like, these, like, low energy states, and-
Yeah
... um, there's so much to search through that's almost overwhelming. Um, but these hints, they sort of maybe put you in an area of the space that's already, like, kind of close to the solution. Maybe not-
Yeah
... quite there yet. And, and there's always this question of, like, how much physics are these models learning, you know, versus, like, just pure, like, statistics. And, like, I think one of the thing, at least I believe, is that, um, once you're in that sort of approximate, you know, area of the solution space, then the models have, like, some understanding, you know, of how to get you to, like, you know, to lower energy, uh, uh, lower energy state.
And so maybe you have some, some light understanding of, of, of physics, but maybe not quite enough, you know, to, to know how to, like, navigate the whole space well. So we need to give it these hints to, like, kind of get there.
So you get it into the right valley, and then it finds the, the minimum-
Yeah
... or something.
Yeah.
Yeah.
One interesting, uh, explanation about how AlphaFold3 works that I think it's quite insightful, of course, doesn't cover kind of the entirety of, of what AlphaFold does, that is, um, that I'm gonna borrow from, uh, Sergey Chychikov at MIT.
Uh, so he sees kind of AlphaFold, and the interesting thing about AlphaFold is it's got this very peculiar architecture that we have since, you know, um, used, and this architecture operates on this, you know, pairwise contacts between amino acids.
And so the idea is that probably the MSA gives you this first hint about what potential, uh, amino acids are close to each other.
MSA is mult-
Multiple sequence alignment. Exactly.
What you were talking about.
Exactly.
Yeah. Yeah.
This evolutionary information.
Yeah.
And, you know, from this evolutionary information about potential contacts, then it's almost as if the model is sort of running some kind of, you know, Dice roll algorithm, where it's sort of decoding, okay, these have to be close.
Okay, then if these are close, and this is connected to this, then this has to be somewhat close. And so you decode, uh, this, that becomes basically a pairwise kind of distance matrix, and then from this rough pairwise distance matrix, you decode kind of the actual potential structure.
Interesting. So there's kinda two different things going on in the, the, the kinda coarse grain and then the fine grain optimizations.
Exactly.
Interesting. Yeah. Very cool.
AlphaFold3 Shift21:28
Yeah. You mentioned AlphaFold3, so maybe a good time to move on to that. So yeah, AlphaFold2 came out, and it was, like, I think fairly groundbreaking for this field. Everyone got very excited. A few years later, AlphaFold3 came out, and, um, maybe for some more history, like, what was the difference between AlphaFold-- what were the advancements in AlphaFold3?
And then I think maybe we'll, after that, we'll talk a bit about the, uh, sort of how it connects to Boltz. But anyway.
Yeah. So after AlphaFold2 came out, I mean, um, you know, Jeremy and I got into the field, and with many others, you know, the clear problem that, you know, uh, was, you know, obvious after that was, okay, now we can do individual chains.
Mm-hmm.
Can we do interactions, interaction different proteins, proteins with small molecules-
Mm-hmm
... proteins with other, other molecules? And so-
So, so quick. Why, why are interactions important?
Interactions are important because to some extent, that's kind of the way that, you know, these machines that, you know, these proteins have a function. You know, the function comes by the way that, uh, they interact with other pro-- uh, with other proteins and other, uh, molecules.
Actually, in the first place, you know, uh, the machines, the individual machines are often, as Jeremy was mentioning, not made of a single chain, but they're made of the multiple chains, and then these multiple chains interact, uh, with other molecules to give, uh, the function to, uh, those.
And on the other hand, you know, when we try to intervene of these interactions-
Mm-hmm
... think about like a disease, think about like a biosensor or many other ways we are trying to design molecules or proteins that interact in a particular way with what we would call a target protein or target, um.
And so, you know, this problem after AlphaFold2, you know, became clear kind of the, the big, uh, one of the biggest problems in the field to, to solve. Uh, many groups, including kind of ours and others, you know, started making some kind of contributions, uh, to this problem of trying to model these interactions.
And AlphaFold3 was, um, you know, put a, was a significant advancement on the problem of modeling interactions. And one of the interesting thing that, uh, they were able to do while, you know, some of the rest of the field that really tried to, tried to model different interactions separately, you know, how protein interacts with small molecules, how protein interacts with other proteins, how RNA or DNA, um, have their structure They put everything together and, you know, train, uh, very large models with a lot of advances, including kind of changing kind of some of their key, uh, architectural choices, and managed to get a single model that was able to set this new state-of-the-art performance across, um, all of these different kind of modalities, whether that was protein-small molecules that is critical to developing kind of
new drugs, ah, protein-protein, understanding, you know, interactions of, you know, proteins with RNA and DNAs and so on.
Just, uh, to satisfy the AI engineers in, in the audience , what were some of the key architectural and data changes that made that possible?
Yeah. So one, uh, critical one that was not necessarily just unique to, uh, AlphaFold3, but there were actually, um, a few other teams, including ours in the field that proposed this, was moving from, you know, modeling structure prediction as a regression problem, so where there is a single answer and you're trying to shoot for that answer, to a generative modeling problem where you have a posterior distribution of possible structures and you're trying to sample, uh, this distribution.
And this achieves two things. One is al- starts to allow us to try to model, um, more dynamic systems. As we said, you know, some of these structures can actually take multiple, um, multiple structures. Uh, and so, you know, you can now, you know, model that, you know, through kind of modeling the entire distribution.
But on the second hand, from more kind of core modeling questions, when you move from a regression problem to a generative modeling, uh, problem, you are really tackling the way that you think about uncertainty in the model in a different way.
So if you think about, you know, I'm undecided between different answers, what's gonna happen in a regression model is that, you know, I'm gonna try to make an average of those different kind of answers that I had in mind.
When, uh, when you have a generative model, what you're gonna do is, you know, sample all these different answers and then maybe use separate models to analyze those different answers and pick out, um, the best. So that was kind of one of the, uh, critical improvement.
The other improvement is that they significantly simplified, to some extent, the architecture, especially of the, um, final model that takes kind of those pairwise representations and turns them, uh, into an actual structure and as-- and now looks a lot more like a more traditional transformer than, you know, like a very, um, specialized equivariant architecture that it was, uh, in AlphaFold3.
So this is a bitter lesson a little bit?
There is some aspect of a bitter lesson, but the interesting thing is that it's very far from, you know, being like a simple transformer. I think one of, um, this field is one of the, uh, I would argue, very few fields in, uh, applied machine learning where we still have kind of architecture that are very specialized and, you know, there are many people that have tried to replace these architectures with, you know, simple transformers.
Yeah.
And, you know, there is a lot of debate in the field, but I think kind of the, uh, most of the consensus is that, you know, the performance that we get from the specialized architecture is vastly superior than what we get through a single transformer.
Interesting. Yeah.
Okay. Can you talk a bit about that, like, specialized architecture? Um, I assume you're referring to triangle layers probably as the core idea or-
Yeah. There's something, uh, maybe s- probably quite fundamental about the fact that we sort of modeled this in like a, you know, sort of second order. So, like, instead of just the sequence, we model every single pair, and then to update every pair, then we need to have these like, well, sort of triangular type operations.
And, um, I think what's interesting about it is, is, is, is couple of things, like, one, I think it relates a little bit to what the input is. You know, we talked about these multiple sequence alignments before and kind of this notion that like, um, you know, we need to look at pairs of residues to try to understand, you know, maybe this initial, like, distance matrix like Gabri was talking about.
Um, and that's something that is very natural, right, to, to model, um, in/to the, um... And, and I think also there's something about the output as well, I think, where I think supervising, you know, over these pairs, I think is also quite powerful.
You know, it's this idea of telling the model, "Hey, like these two things are close to one another. These two things are not." Um, and doing that, I think in, in 3D is, is maybe a bit more challenging.
Um, you know, when I say 3D, so I mean like, so it's like 1D where we model the coordinates in three dimensions, but doing that in like one dimension I think is probably more challenging for the model. And, and yeah, I think to-- it's, it's really survived the test of time.
I mean, you know, this thing came out in 2021 and it's largely the same. I mean, there's been this change to the, the structure module that's been like largely simplified, but where a lot of the magic happens, you know, I think is still, is still in the same place, um, with these like large like pairwise interaction modeling.
Um, that's maybe like the most differentiated portion. The other part I think that's in AlphaFold3, uh, is sort of this moving away from modeling just at the amino acid level to actually sort of having, um, the model sort of alternate between, um, you know, sort of atomic resolution modeling and then more like token level, which is like at the amino acid level.
That's also something that was introduced, um, that I think, you know, was particularly helpful in like, you know, modeling these other modalities like small molecule, et cetera. And like this idea of like coarse grain, like, um, finer grain is I think that's actually quite popular I think in other areas as well.
So that's maybe like not too surprising. But yeah, I think this, the, the fact that you, for some reason, you know, the models that- So I have so much more inductive bias when you, when you go, you know, into this 2D representation, I think is, is, is very interesting.
So you, you mentioned coarse and fine grain, and that brings to mind the sort of ribbony diagrams of proteins-
Yeah
... that I've-- that everyone has probably seen. Can you actually pull up a-
Yeah
... like, a molecule and kinda talk about what, you know, what the different components of that protein are-- we're looking at, like, with the spiral and the arrows and all those little components, and those, like, what, what level of, uh, granularity are we looking at?
H-like, h-how do we think about that? How does the model think about that?
Yeah. So, um, there's actually a little image of our, of our own BoltzLab platform. Um, I have a protein here. Um, and you actually, like, sort of see both, uh, the coarse grain and the finer grain here. So we have the, the sort of ribbon-like structure here that is, um, you know, representing these, these different amino acids in the protein.
But then, like, when we zoom in over this, like, interaction with this small molecule, um, then you see, like, sort of this at the atomic level and how these things, like, you know, interact with one another. There's even, like, the actual, like, bond interactions, uh, here that are, like, uh, shown.
Um, and yeah, like, you know, we, we go from, like, this very abstract representation of these things, you know, like the, the sequence, the graph of the molecule and, and the goal is, like, every single atom should have a coordinate and, um, and, you know, ends up looking like something like this.
It's actually pretty, pretty elegant. I think this is, like, something that's nice, really nice that this field has done, is like it's made really beautiful visualizations of stuff, which is, like, really nice to look at. And yeah, I mean, this is, this is, this is one example.
And so the, th-there's like... Okay, um, the, there's, like, ribbons, there, there's, like, the coily ribbons, there's arr-arrows, there's, like, some sort of, like, not coily ribbons. Like, what do those mean? How does someone think about those?
Yeah. So, um, maybe we can zoom into a few different areas of the protein. This one's actually a good example because there's a few different, uh, secondary structures here. So, um, here you have, you know, what we call an alpha helix.
Um, essentially, like, sort of three categories. There's the alpha helices. Um, this is where it takes a little bit of, like, this, like, ribbon, uh, shape. Um, there's here what we call a, a beta sheet, um, which actually, you know, as the name says, uh, sort of like ribbon going like this, forming, forming a bit of a, of a sheet.
And then you have, um, these more, like, loopy regions, uh, which look, like, more unstructured. And those are, you know, the parts of the protein that are most flexible. They are super important. Uh, you know, maybe, like, one of the most, like, canonical, you know, drug modalities are antibodies.
And antibodies have, like, you know, six of these loops that are, like, largely flexible, but when they interact, you know, kind of come into this, like, fixed structure when interacting with the, um, you know, with, with its target.
Um, so harder to model and really critical to interactions. Um, and yeah, those are largely the three sort of big families.
Okay. And when-- And as a, you know, as a structural biologist or just a biologist, when you look at that, so you, you're basically looking, okay, here's the, the sheet part, here's the... And then you're, you're kinda saying, "Okay, that-- So that'll be bendy, and then I have, like, these coils."
Those, like, what do, what do those mean to you when you look at them?
Yeah, I mean, you know, I, I should say I am not a structural biologist in any way, shape, or form. Um, but you know, there's certain types of interactions that are more canonically associated with these different types-
I see
... of structures.
Yeah.
Yeah.
Okay.
Um, I think, uh, a more well-versed structural biologist-
Okay.
... you know, could give you a more thorough answer than that.
Yeah, yeah.
I don't know if you know anything more than I do, but yeah. Um-
Okay
... yeah. And, and, and you know, like, we've seen, for example, this, this may be related to that point. Like, we've seen, um, you know, some of the early successes of protein design, um, being able to design a binder, you know, to, to any, any target.
Um, a lot of the early success was, like, these, like, very, you know, alpha helix-centric type peptides, which I think are, um, almost like bricks, you know? Um, and the models had, like, a pretty good understanding of those, like, kind of interactions.
I see.
And so, like, there was, like, good success with that, and then took a little bit of time to, like, go from that to, like, you know, more exotic, uh, binders and, and things like that. And so, um, yeah, there's certainly a lot of, um, yeah, a lot of important, um, interaction behaviors associated with, with these structures.
Yeah.
Another interesting thing that I think on the staying on the modeling machine learning side, which I think is somewhat counterintuitive seeing some of the other kind of, uh, fields and applications, is that scaling hasn't really worked kind of the same, uh, in this field.
Um, now, you know, models like AlphaFold2 and AlphaFold3, uh, are, you know, still very large models, but at the same time, they, in terms of parameters, they're actually not very big.
Yeah.
They are definitely below a billion parameters. You know, if you hear these days in, uh, LLM space, you know, a model with less than a billion parameters, you'd think can't do anything.
Yeah.
But on the other hand, when you look at the computational cost of running these models, they're actually a lot more expensive than, uh, it is to run, uh, language models because, as Jeremy was saying, we go from instead of having sort of like quadratic operations, you know, a cubic operation.
And, and so it's interesting how right now in the field, and, and this may be related to, you know, having kind of less data or, you know, needing more inductive biases, but we have, um, this ratio of, you know, amount of computation to parameters that is much, much higher than in other, in other place.
Yeah. If I recall, AlphaFold2 was, like, what, seventy million parameters? Something like that?
Um, yeah, it's, it's something like that. It's quite... Yeah, it's quite small. It's less-- Around a hundred or so.
Mm-hmm.
Yeah.
It, it-- So, like, these, these decisions of triangle layers and, like, these, uh, for AlphaFold2, this, like, interesting equivariant architecture, like, really were priors that- It baked in a lot of the physics of the system, and also co-evolution data is, I think people have argued that is kind of like a, almost like a database lookup of some sorts.
Yeah, yeah.
It also sort of-- So that provides in some sense more parameters as well.
Yeah, I mean, it's, uh, it's more definitely the amount of like, you know, pure like compute FLOPS-
Yeah
... right? Is, is very high and it's almost like more, yeah, more, almost more like reasoning based maybe than like more just like information extraction. You know, I think one of the things that the, uh, part, part of the reason the LLMs are so large isn't just because of their reasoning capability, but it's also because of like, like the sheer quantity of information that they store.
And I think here there's a little bit less of that, you know, and I think it's more about like, you know, decoding this input rather than maybe like memorizing as much of it.
So is there like a loop in the architecture that allows it to compute more for per parameter? Like how does that work?
Part of it is just, you know, exclusively this fact instead of, you know, having operations that operate on the pa- uh, on the single chain, they operate on the pairwise, and so you instead of having like quadratic number of, of up, uh, kind of interactions, you have a cubic number of interactions.
And so that on its own, you know, leads you to have, you know, smaller kind of representation sizes, but more representation that leads to more FLOPS but fewer parameters.
Yeah.
On the other hand, you know, there is actually also this idea of, you know, they somewhat similar to, to reasoning where you recycle kind of this operation. So from AlphaFold2 but also kind of AlphaFold3, they have this interesting framework where, you know, you start with, as we were discussing kind of the input to the model is sort of like this initial understanding of the interactions either from the evolution of the multiple sequence, but also potentially from what we call templates that are basically database lookup of similar structures.
And so how the model works is that, you know, it decodes these and tries to understand a good, you know, potential raw structure of the pairwise interaction. And then what you can do is basically do this recycling where you feed this, uh, kind of understanding back to the input of the model and then try to decode it again.
And people do this, you know, three or four times and, you know, in some cases, you know, I've even tried to do it, uh, tens of times. And so you can see it as a very, very early version of kind of, uh, reasoning, uh, or, you know, trying to, uh, to get.
Yeah. So you, you know, uh, AlphaFold2, really cool. AlphaFold3, really cool. Um, but AlphaFold3 came with a catch, and I think this catch was important for the development of-
Yeah
... you know, Boltz and so on, so.
Yeah. The catch was that it was an amazing paper, Nature, uh, paper, but unfortunately they, uh, decided not to release the model. Uh, you know, AlphaFold2, uh, was open source and since then was, was used, I think the reported numbers is, you know, more than a million scientists.
AlphaFold3 for, you know, commercial reasons that, you know, um, DeepMind has since spin off Isomorphic Lab that is now trying to become sort of like a new pharmaceutical company, uh, had decided to keep this model internal and, and only use it internally.
And now, uh, both, you know, we were in the field and, you know, building on top of models like AlphaFold and so now we no longer had, you know, kind of the base starting point, uh, to build on top.
But even more importantly, everyone in, uh, both kind of academic research and in industry no longer had access to these incredible models that, you know, was, you know, really useful to try to understand, um, kind of biologists but also try to develop new therapeutics.
And so we, um, decided that, you know, to, to take the, the matter in our own hands and decided to kind of, um, try to obtain a model that was of similar accuracy and so largely also, you know, using a lot of, you know, the, uh, information that was in the AlphaFold3 manuscript.
Open-Source Boltz40:01
We went ahead and built Boltz1, which was, um, the first fully open source kind of model to approach the, the level of accuracy of, of AlphaFold3. And, you know, along the way and, and, you know, uh, we can talk about it more, but, you know, we realized that, you know, it was probably too ambitious to have, you know, to see this as, uh, an academic project.
And, you know, there were a lot of things that were kind of missing. And so, um, we decided to also start a, a public benefit company to push kind of this, this mission of, you know, democratizing access to these models that we started with, uh, Boltz1.
Quick interjection. I mean, I remember this. It was actually shocking how fast you got Boltz1 out. Like it was just like two or three months, right? It was-
I think we started in late May and it came in November-
Mm-hmm. Yeah
... if I remember correctly. So slightly longer, but yeah. Yeah, it was relatively quick. I mean, for what it's worth, like, you know, we were working on some of the, some similar ideas at the time. I think like we, you know, for example, this idea of like having a diffusion model on top of, um, this like more, this pairwise strong was something that we were, we were exploring independently.
Um, now when the paper came out it was like really clear that especially for example on the data pipelines there was like so much that we were like not really doing and so um, there was a lot to like catch up on.
Um-
But we were already in a place, I think, where we had, you know, so-some experience working in, you know, with, with the data and working with these type of models. And I think that put us already in like a, a good place to, you know, to, to produce it quickly.
And, you know, and I would I would even say like, I think we could have done it quicker. The problem was like for a while, we didn't really have the compute. And so we couldn't really train the model.
And actually, we only trained the big model once. Uh, that's how much compute we had. We could only train it once. And so like while the model was training, we were like finding bugs left and right.
Yeah.
Uh, a lot of them that I wrote. And like I would-- I remember like I was like sort of like, you know, doing like surgery in the middle, like stopping the run, making the fix, like relaunching and, um-
Wow.
Um, yeah, we never actually went back to the start. We just like kept training it with like the bug fixes along the way-
Right.
Uh, which was-
It was impossible to reproduce now.
Yeah. Yeah, yeah. No, that model is like has gone through such a curriculum that, you know, it's learned some weird stuff. Uh, but, uh, yeah, somehow by miracle, it worked out.
The other, uh, funny thing is that the way that we were training, uh, most of that model was through, uh, a cluster from the Department of Energy.
Yeah.
Uh, that's sort of like a shared cluster that many groups use. And so we were basically training the model for two days, and then it would go back to the queue and stay a week in the queue.
Oh, man.
And so it was, it was, it was pretty painful. And so we actually kind of towards the end, um, I, uh, caught up with, with Devon, the CEO of, of Genesis. And, and basically, I was telling him a bit, a bit about the project and, you know, kind of telling him about this frustration with the compute.
And so luckily, you know, uh, he offered to kind of help. And so we, uh, we got the, the help from Genesis to, you know, finish up the, um, the model. Otherwise, it probably would have taken a couple of extra weeks-
Wow. Nice
... of waiting.
Yeah.
Yeah.
Boltz1, how did that compare to AlphaFold3? And then, and then there's some progression from there.
Yeah. So I would say kind of the Boltz1, but also kind of these other kind of set of models that came, um, around the same time were kind of approaching were a big leap from, you know, kind of the previous kind of open source models, uh, and, you know, kind of, uh, really kind of approaching the level of AlphaFold3.
I would say, still say that, you know, even to this day, there are, you know, some specific instances where, uh, AlphaFold3, uh, works better. I think one, one common example is antibody-antigen, uh, prediction, where, you know, AlphaFold3 still seems to have an edge, uh, in, in many situations.
Obviously, these are somewhat different models. They are-- You know, you run them, you obtain different results. So it's, it's not always the case that one model is better than the other, but kind of in aggregate, we still, uh, especially at the time, saw AlphaFold3 as, you know, still having a bit of an edge.
We should talk about this more when we talk about BoltzGen, but like how do you know one is-- one model is better than the other? Like you-- So you-- I make a prediction, you make a prediction. Like how do you know?
Yeah. So the-- easily, you know, the, the great thing about kind of structure prediction, and, you know, once we're gonna go into the design, uh, space of designing new small molecule, new proteins, this becomes a lot more complex.
But a great thing about structure prediction is that a bit, uh, like, you know, CASP was doing, basically the way that you can evaluate them is that, you know, you train, uh, the model on a structure that was, you know, released across the field up until a certain time.
And, you know, one of the things that we didn't talk about that was really critical in all this development is the, uh, PDB, which is the Protein Data Bank. It's this, um, common resources, basic common database where every, uh, biologist can, uh, publishes their structures.
Mm-hmm.
And so we can, you know, train on, you know, all the structures that were put in the PDB until a certain date.
Mm-hmm.
And then we basically look for recent structures. Okay, which structures look pretty different from anything that was published before?
Yeah.
Because we really want to try to understand generalization. And on this new structure, we evaluate all these different models. So this is-
So you just know when AlphaFold was three-- three was trained, you know when you're-- you intentionally train to the same date or something like that.
Exactly.
Right. Yeah.
And so this is kind of the way that you can somewhat easily kind of compare these models. Obviously, that assumes that you know the, the training set.
You've always been very passionate about validation. I remember like DiffDoc, and then there was like DiffDocL and DocGen. Like you, you really thought-- You've thought very carefully about this in the past. Like, um, yeah, I mean, uh, actually, I think DocGen is like a really funny story that I think, um, I don't know.
I don't know if you wanna talk about that. It's an interesting like, uh-
Yeah, I think one of the amazing things about putting things open source is that, you know, it, um, we get a ton of feedback from, from the field and, you know, sometimes we get kind of great feedback of people really liking the model.
But honestly, most of the times, uh, you know, to be honest, that's also maybe the most useful feedback-
Yeah
... is, you know, people sharing about where it doesn't work. And so, you know, at the end of the day, it's critical, and this is, you know, also something, you know, across other fields of, of machine learning. It's always critical to set, uh, to do progress in machine learning, set clear, uh, benchmarks.
And, you know, as, you know, you start, you know, doing progress of certain benchmarks, then, you know, you need to improve the benchmarks and make them harder and harder. And this is kind of the progression of, you know, how the field operates.
And so, you know, the example of, of, uh, DocGen was, you know, we, um, published this initial, uh, model called DiffDoc, um, in my first year of PhD, which was sort of like, you know, one of the early, um, models to try to predict, uh, bio-- kind of interactions between proteins, small molecules, um, that we-- about a year after, uh, AlphaFold2 was published.
And Now, on the one end, you know, on these benchmarks that we were using at the time, uh, DiffDock was, was doing, uh, really well, kind of, you know, uh, outperforming kind of some of the traditional physics-based methods.
But on the other hand, you know, when we started, you know, kind of giving these, uh, tools to kind of many, uh, biologists and, uh, one example was, uh, that we collaborate with was the group of Nick Palitzia at Harvard.
Uh, we noti- started noticing that there was this clear pattern where for proteins that were very different from the ones that we're trained on, uh, the models was, was struggling. And so, you know, that seemed clear that, you know, this is probably kind of where we should, you know, put our focus on.
And so we first developed, you know, with, uh, Nick and his group a new benchmark and then, you know, went after and said, "Okay, what can we change and kind of about the current architecture to improve this, uh, pattern generalization?"
And this is the same that, you know, we, uh, we're still doing today, you know, uh, kind of where does the model not work, you know, and then, you know, once we have that benchmark, you know, let's try to, uh, throw everything we, uh, any ideas that we have at the problem.
And there's a lot of like healthy skepticism in the field, which I think, you know, is, is, is great. And I think, you know, it's very clear that there's a ton of things the models don't really work well on.
But I think one thing that's probably, you know, undeniable is just like the pace of, pace of progress-
Yeah
... you know, and how, how much better we're getting, you know, every year. And so I think if you, you know, if you assume, you know, any constant, you know, rate of progress moving forward, I think, you know, um, things are gonna look pretty cool at some point in the future.
ChatGPT was only three years ago.
Yeah. I mean, it's wild, right? Like-
What?
Yeah. Yeah. Yeah, it's one of those things like even being in the field, you don't see it coming, you know? And like, I think, yeah, um, hopefully we'll, you know, we'll, we'll continue to have as much progress we've had the past few years.
So this is maybe an, an aside, but I, I'm really curious. You get this great feedback from the, from the community, right, by being open source. Um, my question is partly like, okay, yeah, if you open source then everyone can copy what you did, but it's also maybe balancing priorities, right?
Where you-- like all my customers are saying I want this. Like there's all these problems with the model, yeah, yeah, but that like my customers don't care, right? So like how do you, how do you think about that?
Yeah. So I would say a couple of things. One is, you know, part of, uh, our goal with Boltz and, you know, this is also kind of established as kind of the mission of the public benefit company that we started, is to democratize the access to these tools.
But one of the reason why we realized that Boltz needed to be a company, it couldn't just be an academic project, is that putting a model on GitHub is definitely not enough to get, you know, chemists and biologists, you know, across, you know, uh, both academia, biotech and, and pharma to use your model to, uh, in their therapeutic programs.
And so a lot of what we think about, you know, at Boltz beyond kind of the just the models is thinking about all the layers that come on top of the models to get, you know, from, you know, those models to something that can really, uh, enable scientists, uh, in the industry.
And so that goes, you know, into building kind of the right kind of, uh, workflows that take in kind of, for example, the data and try to answer kind of directly that-- those problems that, you know, the chemists and the biologists are asking, and then also kind of building the infrastructure.
And so this to say that, you know, even with kind of, you know, models fully open, you know, we see kind of a ton of, um, potential for, you know, um, you know, products in the space. And, um, the critical part about a product is that even, you know, for example, with an open source model, you know, running the model is not free.
You know, as we were saying, these are pretty expensive model and especially, and maybe we'll get, uh, into this, you know.
Yeah.
These days we're seeing kind of pretty dramatic inference time scaling, uh, of, of these models-
Yes
... where, you know, the more you run them, the better the results are. Uh, but there, you know, you start getting into a point that compute and compute cost becomes a critical factor. And so putting a lot of work into building the right kind of infrastructure, building the optimizations and so on, really allows us to provide, you know, a much better service potentially to the open source models.
But that to say, you know, even though, you know, with the product, we can provide a much better service, I do still think and we will continue to put a lot of our models open source because the, um, the critical kind of role I think of open source models is, you know, helping kind of the community progress on the research and, you know, from which we, we all benefit.
And so, you know, we'll continue to, on the one hand, you know, put some of our kind of base models open source so that the field can, can build on top of it. And, you know, as we discussed earlier, we learn a ton from, you know, the way that the field, uh, uses and builds on top of our models.
But then, you know, try to build a product that gives the best experience possible to scientists, um, so that, you know, like a chemist or a biologist doesn't need to, you know, spin off a GPU and, you know, set up, you know, our open source model in a particular way, but can just, you know, uh, a bit like, you know, I even though I am a computer scientist, machine learning scientist, I don't necessarily, you know, take a open source LLM and try to kind of spin it off.
But, you know, I just maybe open, um, the ChatGPT app or a Cloud Code and just use it as an amazing product. Uh, we kind of want to give the same experience to scientists around the world.
I heard a good analogy yesterday that a surgeon doesn't want the hospital to design a scalpel, right?
Surgeon just buy the scalpel.
You, you wouldn't believe like the number of people even like in my, uh, short time, um, you know, between A-AlphaFold3 coming out and, and the end of the PhD, like the, um, number of people that would like reach out just for like us to like run AlphaFold3 for them.
Oh, really?
You know, or things like that, just because like, um, or Boltz in our case, you know, just because it's like not that easy, you know, to do that com- you know, if you're not a computational person. And I think like part of the goal here is also that, you know, we continue to obviously build an interface with computational folks, but that the, you know, the models are also accessible to like a larger, broader audience and, and then that comes from like, you know, good interfaces and stuff like that.
I think one, like, really interesting thing about Boltz is that with the release of it, you didn't just release a model, but you created a community.
Yeah.
Did that community-- It grew very quickly. Did that surprise you? And, like, what is the, the evolution of that community and how has that fed into Boltz-
Uh, if you, if you look-
... as Boltz company?
If you look at its growth, it's-
Yeah
... it's like very much like when we release a new model, it's like there's a big, uh, big jump. Um, but yeah, it's, I, I mean, it's been great. You know, we have a Slack community, uh, that has like thousands of people on it.
Um, and it's actually like self-sustaining now, which is like the really nice part because, you know, it's, it's almost overwhelming, I think, you know, to be able to like, answer everyone's questions and help. It's really difficult, you know, with the, the few people that we were.
But it ended up that like, you know, people would answer each other's questions and like sort of like, you know, help one another and so the, the, the Slack, you know, has been like kind of, yeah, self-self-sustaining and that's been, it's been really cool to see.
Um, and, um, you know, that's, that's for like the Slack part, but then also obviously on GitHub as well, we've had like a nice, nice community. Um, you know, I think we also aspire to be even more active on it, you know, than we've been in the past six months, which has been like a bit challenging, you know, for us.
But, um, yeah, the, the community has been, has been really great and, you know, there's a lot of papers also that have come out with like new evolutions on top of Boltz and, um, it surprised us to some degree because like there's a lot of models out there and I think like, you know, sort of people converging on that was, was really cool.
And, you know, I think it speaks also I think to the importance of like, you know, when, when you put code out like to try to put a lot of emphasis on like making it like as easy to use as possible and something we thought a lot about when we released the, the code base.
Um, you know, it's far from perfect, but, you know.
Do you think that that was one of the factors that caused your community to grow is just the focus on easy to use, make it accessible?
I think so, yeah. And we, we've, we've heard it from a few people-
Okay. Yeah
... over the, over the, over the years now and, um, you know, and some people still think it should be a lot nicer-
Oops
... and they're, and they're right. Uh, and they're right. But, um, yeah, I think it was, you know, at the time maybe a little bit easier than, than other things.
The other I think part, uh, I think led to, to the community and to some extent I think, you know, like the, uh, somewhat the trust in the community and kind of what we, what we put out is the fact that, you know, it's not really been kind of, you know, one model by and maybe we'll talk about it, you know.
After Boltz1, you know, there were maybe another couple of models kind of released, you know, uh, or open source kind of soon after. We kind of continued kind of that open source journey and released Boltz2, where we are not only improving kind of the structure prediction but also starting to do affinity predictions-
Mm-hmm
... understanding kind of the strength of the interactions-
Mm
... between these different models, which is this critical component, critical property that you often want to optimize, uh, in discovery programs. And then, you know, more recently also kind of protein design, uh, model. And so we've sort of been building this suite of, of models that come together, interact with one another, where, you know, kind of there is almost an expectation that, you know, we, we take very at heart of, you know, always having kind of, you know, across kind of the entire suite of different tasks the best or across the best model, uh, out there so that it's sort of like, uh, our open source tool can be kind of the go-to, uh, model for everybody in the, in the industry.
BoltzGen Design58:04
I really wanna talk about BoltzGen, but before that, one last question in this direction. Was there anything about the community which surprised you? Were there any like someone was doing something and you're like, "Why would you do that?
That's crazy," or, "That's actually genius and I never would've thought about that"?
I mean, we've had, you know, many contributions. I think like some of the interesting ones like, I mean, we had, you know, this one individual who like wrote like a complex GPU kernel, you know, for part of the architecture, um, on, on a piece of-- The funny thing is like that piece of the architecture had been there since AlphaFold2.
Mm.
And I don't know why it took Boltz for this, you know, for this person to, you know, to decide to do it, but that was like a, a really great contribution. We've had a bunch of others like, you know, people figuring out like ways to, you know, hack the model to do cyclic peptides.
Like, you know, there's... I don't know if there's any other interesting ones come to mind.
One, one cool one and, and this was, you know, something that initially was, uh, proposed as, you know, as a message in the Slack channel by, by Tim, Tim O'Donnell was basically he was, you know... There are some cases especially for example we discussed, you know, antibody-antigen interactions where the models don't necessarily kind of, uh, get the right answer.
What he noticed is that, you know, the models were somewhat stuck into predicting kind of the, the antibody to interact with the part of the antigen that was incorrect. And so he basically, um, ran the experiments. In this model you can condition, uh, basically you can give hints.
And so he basically gave, uh, you know, random hints, uh, to the model basically, "Okay, you should bind to this residue, uh, you should bind to the first residue or you should bind to the 11th residue or you should bind to the 21st residue," or, you know, basically every 10 residues scanning the entire antigen and-
Residues are the, the-
The amino acids
... the amino acid, yeah.
So the first amino acid, the 11th amino acid-
Mm
... and so on. So it's sort of like doing a scan and then, you know, conditioning the model to predict all of them and then looking at the confidence of the model in each of those cases and taking the top.
And so it's sort of like a very somewhat crude way of doing kind of inference time search, but surprisingly, you know, for, for antibody-antigen prediction it actually kind of helped quite a bit. And so there's some, you know, interesting ideas that, you know, as a, um, obviously as kind of developing the model you, you say kind of, you know, "Wow, this is...
Why would the model, you know, be so dumb?" But, you know, it's, it's very interesting and, and that, you know, leads you to also kind of, you know, start thinking about, okay, how, why can I do this, you know-
Yeah
... not, you know- With this brute force-
Yeah
... but, you know, in a, in a smarter way. And so we've also done a lot of work on that direction.
And that speaks to like the, you know, the power of scoring. Uh, we're seeing that a lot. I'm sure we'll talk about it more when we talk about BoltzGen, but, um, you know, our ability to like take a structure and determine that that structure is like good-
Yeah
... you know, like somewhat accurate, uh, whether that's a single chain or like an interaction, is a really powerful, you know, way of improving, you know, the, the models. Like sort of like, you know, if you can sample a ton and you assume that like, you know, if you sample enough, you're likely to have like, you know, the good structure-
Mm-hmm
... then it really just becomes a ranking problem.
Yeah.
Um, and you know, now we're-- you know, part of the in-inference time scaling that Gabriel was talking about is, is very much that. It's like, you know, the more we sample, the more we like, you know, the ranking model ends up finding something it really likes.
Um, and so I think our ability to get better at ranking, I think is also what's gonna enable sort of the next, you know, next big, big breakthroughs.
Interesting.
I guess there's a-- my understanding, there's a diffusion model, and you generate some stuff, and then you... I guess it's just what you said, right? D- then you rank it using a score, and then-
Yeah
... you finally, um... And it was-- So like, can you talk about those different parts?
Yeah. So first of all, like the-- one of the critical kind of, you know, beliefs that we had, you know, also when we started working on Boltz1 was sort of like the structure prediction models are somewhat, you know, our field version of some foundation models.
You know, learning about kind of how proteins and other molecules interact and, and then we can leverage that learning to do all sorts of other things. And so with Boltz2, we leverage that learning to do thing-- uh, affinity predictions.
So understanding kind of, you know, if I give you this protein, these small molecules, how tightly is their interaction? Uh, for BoltzGen, what we did was taking kind of the kind of foundation models and then fine-tune it to predict kind of entire new proteins.
And so the way basically that that works is sort of like instead of, uh, for the protein that you're designing, instead of filling in an actual sequence, you fill in a set of blank tokens, and you train the models to, you know, predict both the structure of kind of that protein and with the structure also what the different amino acids, uh, of, um, that proteins are.
And so basically the way that, uh, BoltzGen operates is that you feed, uh, the st- a target protein that you may want to kind of bind to or, you know, another DNA, RNA, and then you feed, um, the high level kind of design specification of, you know, what you want your new protein, uh, to be.
For example, it could be like an antibody with a particular framework, could be a peptide, could be many other things.
And that's with natural language or with this-
And that's, you know, basically, you know, prompting and we have kind of this sort of like spec w-
Okay
... that you, you specify.
Uh-huh.
And, you know, you feed kind of this, this spec to the model, and then the model translate this into, you know, a set of, you know, uh, tokens, a set of conditioning to the model, a set of, you know, blank tokens.
And, and then, you know, basically decodes as part of the, um, diffusion models decodes a new structure and a new sequence, uh, for your, your protein. And, you know, basic- and then we take that, and as Jeremy was saying, you know, trying to score it and you know how good of, you know, a binder it is to that original target that you-
You're using Bol- basically Boltz2 to, well, predict the folding and the affinity to that molecule. So and then that is your-- that kinda gives you a score. Is that-
Exactly.
Yeah.
So you, you use this model to predict the structure, and then you do two things. One is that you predict the structure-
Mm-hmm
... and with something like Boltz2, and then you basically compare that structure with what the model, uh, predicted, what BoltzGen predicted. Um, and this is sort of like in the field called consistency. It's basically you want to make sure that, you know, the structure that you're predicting is actually what you're trying to design, and that gives you a much better confidence that, you know, that's a good design.
And so that's the, the first filtering. And the second filtering that we did, uh, as part of kind of the, the BoltzGen pipeline, um, uh, that was released is, uh, that we look at the confidence that the model has in the structure.
Now, unfortunately, kind of going to your, your question of, you know, predicting affinity, unfortunately, confidence is not a very good predictor of affinity. Um, and so one of the things that we've actually done a ton of progress, you know, since, uh, we released BoltzGen and, uh, kind of we have some, uh, new results that we are gonna kind of announce soon is kind of, you know, the ability to get much better, uh, hit rates when instead of, you know, trying to rely on confidence of the model, we are actually directly trying to predict the affinity, uh, of that interaction.
Okay. Just backing up a minute. So your diffusion model actually predicts not only the protein sequence but also the folding of it.
Exactly.
Oh.
And actually kind of the way, uh, one of the big, um, kind of different things that we did compared to, uh, other models in the space, and, you know, there were some, uh, papers that had already kind of done this before, but we, uh, really scaled it up, was, you know, basically somewhat merging kind of the structure prediction and the sequence prediction into almost the same task.
Mm-hmm.
And so the way that BoltzGen works, uh, is that you are basically the only thing that you're doing is predicting the structure. So the only sort of supervision is we give you a supervision on the structure. But because the structure is atomic- And, you know, the different amino acids have a different atomic composition basically from the way that you place the atoms.
We also understand not only kind of the structure that you wanted, but also the identity of the amino acid that, you know, the models believed was there. And so we've basically instead of, you know, having these two supervision signals, you know, one discrete, one continuous that somewhat, you know, don't interact well together, we sort of like build kind of like an encoding of, you know, sequences and structures that allows us to basically use exactly the same supervision signal that we were using to Boltz2 that, you know, uh, you know, largely similar to what AlphaFold3 pr- uh, proposed, which is very scalable, and we, we can use that to design new proteins.
Oh, interesting.
Maybe a quick shout-out to Hannes, uh, Stark on our team who, like, did all this work. Um, yeah.
Yeah, it w- that was a really cool idea. I mean, like, looking at the paper and there's this just like encoding where you just a-add a bunch of, I guess, kind of f- y- atoms which can be anything, and then they get sort of rearranged and then basically plopped on top of each other so that-- and then that encodes what the amino acid is, and there's sort of like a unique way of doing this.
That was, that was like such a really, such a cool, fun idea. Yeah.
I think that idea was-- had existed before. I think it wasn't-
Yeah, there were a couple of papers-
Yeah, yeah
...that had, had proposed this and-
Yeah
...and Hannes really took it to, um, to the large scale.
Yeah.
In the paper, a lot of the paper for BoltzGen is dedicated to actually the validation of the model. In my opinion, we talk a-- all the people we basically talk about feel that this sort of like in the wet lab or whatever the appropriate, you know, sort of val-- like in real world validation is the whole problem almo-- or not the whole problem, but a big, giant part of the problem.
So can you talk a little bit about the highlights from there that really-- 'cause, uh, to, to me, uh, the results are, uh, impressive both from a, the perspective of the, you know, the model and also just the, the effort that went into the validation by a large team.
First of all, I, I think I should tr- start saying is that both when we were at MIT and Thomas Jiakolis and working at Barzilay's lab as well as at Boltz, uh, you know, we are not a, we're not a bio lab and, you know, we are not a therapeutic company.
And so to some extent, you know, we were first forced to, you know, look outside of, you know, our group, our, um, team to do the experimental validation. And so one of the things that, um, really Hannes, uh, and the team pioneered was the idea, okay, can we go not only to, you know, maybe a specific group and, you know, trying to find a specific system and, you know, maybe overfit a bit to that system and, and trying to validate, but how can we test these models across a very wide variety of different settings so that, you know, anyone in, uh, in the field and, you know, protein design is, you know, such a kind of wide, um, task with all sorts of different applications from therapeutic to, you know, biosensors and, uh, many others
that, you know-- So can we get a validation that is kind of goes across, uh, many different tasks? And so he basically put together, you know, I think it was something like, you know, twenty-five different, you know, academic and industry labs that, you know, sort of like committed to, you know, testing, uh, some of the designs from the model and some of this testing is, is still ongoing.
Yeah.
Uh, and, you know, giving, uh, results kind of, uh, back to us in exchange for, you know, hopefully getting some, you know, new seq- new great sequences for, uh, their task. And, and he was able to, you know, coordinate this, you know, very wide, uh, set of, you know, uh, scientists and, uh, already in the paper, I think we, uh, shared, uh, results from I think, uh, eight to ten different, uh, labs, uh, kind of showing kind of results from, you know, designing peptides, uh, designing, uh, to target, you know, ordered proteins, peptides, uh, targeting disordered proteins.
We showed results, you know, of, uh, designing proteins that bind to small molecules. Uh, we showed results of, you know, designing nanobodies and across a wide variety of different targets. And so that sort of like gave to the, to the paper a lot of, you know, validation and to the model a lot of validation that was kind of, uh, wide, uh-
Just, uh, again, um, for our n-non-biologist, uh, audience, peptides, nanoparticles, what are these things that are being designed? Why-- Like, what is interesting about these particular things? Why are the-- is there focus on them?
Yeah. So largely, you know, they're all proteins. It's just different shape the proteins take. Peptides, uh, is, is a small protein, uh, and is, you know, a relatively common type of, of therapeutic. The very common examples these days-
Yeah.
...are the, you know, GLP-1, uh-
That's Ozempic, right?
Ozempic and so on. They're all peptides formed by both canonical and non-canonical amino acids. Um, the-- When we think about kind of larger proteins, they also can take different shapes. There is, you know, maybe one of-- We have this term called mini-proteins, which is like a very vague term to say kind of any sort of shape.
Uh, but then there are some specific shapes that, you know, proteins can take and, um, so one very common one is antibodies. And antibodies are a particular type of protein in our body that is involved, uh, in our immune system, and it's formed by, um, a Basically a set of, you know, four different, um, protein chains, you know, two, uh, two heavy, we call heavy, which are longer, and two light that come together and form kind of this interesting structure.
And those are, you know, very common type also of therapeutic because of, you know, their function that they have in our immune system. And finally, there are kind of what are called nanobodies that I mentioned that are sort of like the equivalent of antibodies, but on specific, in specific animals.
So there are some animals, and I think some examples are llamas-
Llamas, sharks
... camels, and sharks that instead of having kind of this more complex set of, you know, four proteins coming together, it's a single protein. And, and so recent, in recent years, it's been also a relatively common type of therapeutic, uh, that people are trying to design.
And so these are sort of like have a similar function to antibodies, but, uh, are simpler in terms of, uh, structure.
And so those would be therapeutics for those animals, or are they relevant to humans as well?
They're relevant to humans as well.
Okay.
Uh, obviously, you need to do some work into, uh, quote, unquote, "humanizing them," making sure that, you know, they have the right characteristics to, so they're not, uh, toxic to humans and so on. Uh, but there are, um, some approved medicine in the, uh, in the market that are nanobodies.
Yeah.
There's a general pattern, I think, in, like, in trying to design things that are smaller. You know, like, it's easier to manufacture. Um, at the same time, like, that comes with, like, potentially other challenges, like maybe a little bit less selectivity than, like, if you have something that has, like, more hands, you know?
Yeah.
Um, but the... Yeah, there's this big desire to, you know, try to, yeah, design mini proteins, nanobodies, small peptides, you know, that just are just great drug modalities.
That-that's because they're more selective?
No.
Uh-
Generally, I think it's largely a manufacturing thing.
Oh, okay. I see.
Yeah.
So it's-
You know, the bigger, the bigger the protein, the more complex-
Got it
... potentially.
Yeah.
Yeah.
I put a pin in. I wanna understand, like, how do you actually build a protein? Like, I know you guys are not wet-
Sure
... wet lab technicians, but I wanna, like, hear more about the validations that you've done.
Um, I can try.
Yeah.
Essentially, like, so we work with, uh, us organizations to do the, the lab evaluation. Um, we, we don't do any of it ourselves, and typically what we send those people is we send them, you know, the sequence of the target and the sequence of the binder.
In this case, for example, a nanobody, it would be like a single, single chain. Uh, we have to order the DNA. That's yet another company that produces that DNA, sends it over, and then you have to, like, express the proteins.
So you have to, like, express the target. You have to express the binders. And by expressing, what I mean is, like, you put, you know, the, the DNA, uh, typically in, like, you know, either, uh, some capsid, or you put it in, like, a, and then you express it in like a, a yeast, for example, or you can, like, do it, uh, in like cell-free systems now.
But, um, essentially, you know, you, you use, like, typical biological mechanism, um, to-
So you're kinda like hijacking the yeast to create-
Yeah. You give it-
Create-
... the extra DNA, and you're like-
Yeah
... "Okay," like, "create this thing."
So, so you wanna... This is like, um, you're, you're amplifying the DNA in some way. Is that basically it?
Yeah. Yeah. So there, yeah. So they're jumping some steps. There's like, uh, yeah, amplifying the DNA. There's all this stuff, and then, um, but at the end of the day, the, you starts to produce a lot of this protein.
Then you need to purify it because there's a lot of other stuff in there. So typically you have, like, a tag on the protein, and then, like, based on that tag, you can sort of purify. Um, once you have, like, your pure target, your pure binder, you can then, like, run your binding assay, and then that's where my knowledge stops.
But there's various methods to do that.
Yeah. Yeah.
Uh, you know, you put things in a well, and then you can measure, um, the, you know, the, the, the binding strength. Um, and then we get the results back. Um, and we also get to know, like, you know, whether it was a binder, but also, like, how strong of a binder it was.
Um, and that's generally the, the process.
Right. So the, so that you kinda, you specify the m-molecule. They create some DNA that can create the RNA.
Yeah.
They create some RNA from that. That creates the, uh, the protein. You take the protein, and you use lab voodoo to-
Yeah.
... measure the, the binding-
Yeah
... strength of that, or the two proteins or two molecules.
Yeah.
Yeah.
Yeah, that's right.
Okay. So we were, I think we were left off, we were talking about validation in the lab, and I was very excited about seeing, like, all the diverse validations that you've done. Can you-
Yeah
... go into some more detail about them?
Yeah.
The specific ones.
Yeah. The nanobody one, I think we did, what was it? 15 targets? Is that correct?
14.
14 targets, um, testing. Um, so we-- typically the way this works is, like, we, uh, make a lot of designs-
Yeah
... right? On the order of, like, tens of thousands, and then we, like, rank them, and we pick, like, the top N. Uh, in this case, N was 15, right, um, for each target, and then we, like, measure sort of like the success rates-
Okay
... both on, like, how many targets we were able to, um, to get a binder for, and then also in, like, more generally, like, out of all of the, you know, binders that we design, how many actually prove to be, uh, good binders.
Some of the other ones I think involved, like, yeah, like a, we had a, a, a cool one where there was a small molecule. We designed a protein that, uh, you know, binds to it.
Mm-hmm.
Um, that has a lot of, like, interesting applications. You know, for example, like Gabby mentioned, like biosensing and things like that-
Okay
... uh, which is pretty cool. Um, we had, we had a disordered protein I think you mentioned also. Um, and yeah, I think some of, maybe some, those were some of the, the highlights.
Yeah.
So I would say that the way that we structure kind of some of those validations, uh, was on the one end, we have validations across a whole set of different problems that, you know, the biologists that we're working, uh, with came to us with.
So we were trying to, uh, for example, in some of the experiments design peptides that would, uh, target the RexC, which is a target that is involved in metabolism. Um, we had, you know, a number of other, ah, applications where we were trying to design, you know, peptides or other modalities against some other therapeutic relevant targets.
Um, we designed some, um, proteins to bind small molecules. And then some of the other, ah, testing that we did was really trying to get like a more broader sense of how does the model work, especially when tested, you know, on somewhat generalization.
So one of the things that, you know, we, we found, ah, with the field was that a lot of the validation, especially outside of the validation that was on specific problems, was done on targets that have a lot of, you know, known interactions in, in the training data.
And so it's very-- always very hard to understand, you know, how much are these models really just regurgitating kind of what they've seen or trying to imitate what they've seen in the training data versus, you know, really be able to design, uh, new proteins.
And so one of the experiments that we did was to, ah, take nine targets from, um, the PDB, filtering to things where there is no known interaction, ah, in, in the PDB. So basically, the model has never seen kind of this particular protein bound or a similar protein bound to another protein.
So there is no way that the model, ah, you know, from its training set can sort of like say, "Okay, I'm just going to, ah, kind of-"
Tweak something.
"... tweak something and, and just imitate this particular kind of interaction." And, and so we, we took these nine proteins. We worked with, ah, Adaptive, a CRO, and basically tested, you know, fifteen mini-proteins and fifteen nanobodies against each one of them.
And the very cool thing that we saw was that on two-thirds of those targets, we were able to, from these fifteen designs, ah, get nanomolar, ah, binders. Nanomolar, roughly speaking, just a measure of, you know, how strongly kind of the interaction is.
Roughly speaking, kind of like a nanomolar binder is approximately the kind of binding strength of binding that you need for a therapeutic.
Okay.
BoltzLab Platform1:21:37
Yeah. So maybe switching, ah, directions a bit. Um, so I, I-- Boltz Lab was just announced, um, this week. Or was it last week? Yeah.
Yeah.
Um, ah, this is like your first, I guess, product, if, if that's the-
Yeah
... if you want to call it that. Um, can you talk about what Boltz Lab is and, um, yeah, you know, what you hope that people take away from this?
Yeah. You know, as we mentioned, like I think at the very beginning, is the goal with the product has been to, you know, address what the models don't on their own. Um, and there's largely sort of two categories there.
Um, you know, or let's say-- let's say I'll, I'll split it in three. Um, the first one is that, you know, it's one thing to predict, you know, a single interaction, for example, like a single structure. Um, it's another to like, you know, very effectively, ah, search a space, a design space, you know, to, to produce something of value.
And, you know, what we've, what we've found, like sort of building of this product is that there's a lot of steps involved, you know, in that, that we sort of need to like, you know, accompany the user through.
Um, you know, one of those steps, for example, is like, you know, the, the creation of the target itself. You know, how do we make sure the model has like a good enough understanding of the target so we can like design something?
And there's all sorts of tricks, you know, that you can do, ah, to improve like a particular, you know, structure prediction. And so that's sort of like, you know, the first stage. And then there's like this stage of like, you know, designing and searching the space efficiently.
You know, for something like BoltzGen, for example, like you, you know, you, you design many things and then you rank them. For example, for small molecule, the process is a little bit more complicated. We actually need to also make sure that, ah, the molecules are synthesizable.
And so the way we do that is that, you know, we have a generative model that, um, learns to use like appropriate building blocks such that, you know, it can design within a space that we know is like synthesizable.
And so there's like, you know, this whole pipeline really of different models involved, you know, in being able to, um, to design a molecule. And so that's been sort of like the first thing. We call them agents. We have a protein agent and we have a small molecule design agents.
And that's really like at the core of like what powers, you know, the Boltz Lab platform.
So these agents are, are they like a language model wrapper or they're just like your models and you're just calling them agents-
No. Yeah
... because they, they, they sort of perform a function-
Yes
... on behalf of you. Okay.
They're more of like a, you know, a recipe if you wish. And I think, ah, we use that term sort of because of, you know, sort of the complex pipelining and automation, you know, that goes into like all this plumbing.
Um, so, so that's the first part of the product. Um, the second part is the infrastructure. You know, we need to be able to do this at very large scale for any one, you know, group that's doing a design campaign.
Um, you know, let's say you're designing, you know, I'd say a hundred thousand possible candidates, right? To find the good one. Um, that is, you know, a very large amount of compute. Ah, you know, for small molecule that's on the order of like a few seconds per, ah, per design.
For proteins it can be a bit longer. And so, you know, ideally you want to do that in parallel otherwise it's going to take you weeks. Um, and so, you know, we've put a lot of effort into like, you know, our ability to
Have a GPU fleet that allows any one user, you know, to be able to do this kind of like large parallel search.
So you're amortizing the cost over your, your users basically.
Exactly. Exactly. And, you know, to some degree, like it's-- whether you do, uh, you use ten thousand GPUs for like, you know, a, a minute, is the same cost as using, you know, uh, one GPUs for God knows how long, right?
So you might as well try to parallelize if you can. So, you know, a lot of work has gone, has gone into that, making it very robust, you know, so that we can have like a lot of people on the platform doing that at the same time.
Um, and, and the third one is, is the interface. And the interface comes in, in two shapes. One is, um, in form of an API, and that's, you know, um, really suited for companies that want to integrate, you know, these pipelines, these agents directly in existing, you know, workflows that they have or like existing user interfaces that they have.
And we're already like partnering with, you know, a few distributors, you know, that are gonna integrate our API. And then the second part is the, the act- the user interface. And, you know, we, we've put a lot of thoughts also into that, and this is when I, I mentioned earlier, you know, this idea of like broadening the audience.
That's kind of what the, the user interface is about. And we've built a lot of interesting features in it, you know, for example, for collaboration. Um, you know, when you have like potentially multiple medicinal chemists are going through the results and trying to pick out, okay, like what are the molecules that we're gonna go and test in the lab, it's powerful for them to be able to, you know, for example, each provide their own ranking and then do consensus building.
And, um, so there's a lot of features around, you know, launching these large job, but also around like collaborating on analyzing the results, um, that we try to solve, you know, with, with that part of the platform. So Boltz Lab is sort of a combination of these three objectives into like one, you know, sort of cohesive platform.
Who is this accessible to?
Everyone. Uh, you do need to request access today. We're still like, you know, sort of ramping up the usage. Um, but anyone can request access. Um, if you are an academic in particular, uh, we, uh, you know, we provide, um, a fair amount of free credit, so you can play with the platform.
If you are a startup, uh, or a biotech, you may also, you know, reach out, and we'll typically like actually hop on a call just to like understand what you're trying to do and, um, also provide a lot of free credit to, to get started.
Um, and, uh, of course, also with larger companies, uh, you know, we, uh, we can deploy this, this platform in a more like secure environment. And so that's like more like custom, you know, uh, deals that we make, you know, with, with the partners.
Um, so, you know, and that's sort of at the ethos of, of Boltz. I think this idea of like, uh, servicing everyone and not necessarily like going after just, you know, the, um, the really large enterprises. Um, and that starts from the open source, but it's also, you know, main-ma-- a key, key design principle of, of the product itself.
Yeah. One thing I was thinking about with regards to infrastructure, like in the LLM space, you know, the cost of a token has gone down by, I think, a factor of a thousand or so-
Yeah
... over the last three years, right?
Yeah, yeah.
And is it possible that like, essentially, you can exploit economies of scale and infrastructure, that you can make it cheaper to run these things yourself than for any person to roll their own system?
Hundred percent. Yeah. I mean, we're already there. You know, like running Boltz on our platform, especially on, on a large screen, is like considerably cheaper, um, than it would probably take anyone to put the open source model out there and, and run it.
And, and, you know, on top of the infrastructure, like one of the thing that we've been working on is, is accelerating the models. So, you know, our, our small molecule screening pipeline is ten x faster on Boltz Lab than it is in the open source.
Um, you know, and that's, that's also part of like, you know, building, building a, a, a product, you know, of something that, that scales really well. Um, and yeah, uh, we really wanted to get to a point where like, you know, we could keep prices, uh, very low, um, you know, in, in a way that it would be a no- no-brainer, you know, to, to use Boltz through, through our platform.
How do you think about validation of your like agentic systems?
Yeah.
Because, you know, as you were saying earlier, like we're-- AlphaFold style models are really good at, let's say, monomeric, you know, proteins where you have, you know, co-evolution data. But now suddenly the whole point of this is to design something which doesn't have-
Right
... you know, co-evolution data, something which is really novel. So now you're basically leaving the, the domain that you thought was-
Yeah
... you know, that you, you know you are good at. So like how do you validate that?
Yeah. Um, I mean, I, I like every complete, but there's, um, there's obviously, you know, a ton of computational metrics that we rely on, but those are only take you so far. Um, you really gotta go to the lab, you know, and, and test, you know, okay, with this method A and this method B, uh, how much better are we?
You know, how much better is my, uh, my hit rate? How stronger are my binders also? It's not just about hit rate, it's also about how, how good the binders are. And there's really like no way, no way around that.
And I think we're, you know, we've really ramped up the amount of experimental validation, um, that we do so that we like really track progress, you know, as scientifically s-sound, you know, as, as possible. Um, I don't know if there's anything.
Yeah, no, I think, you know, one thing that is unique about us and maybe companies like us is that because we're not working on like maybe a couple of therapeutic, uh, pipelines where, you know, our validation would be focused on those, we-- when we do an experimental validation, we try to test it across tens of targets.
And so that on the one end, we can get a much more statistically significant, um, result and, and really allows us to make progress from the methodological side without being, you know, steered by, you know, overfitting on any one particular system.
And of course, we choose, you know, we always try to choose, uh, targets and problems are sort of like at the frontier of what's possible today. So, you know, you don't want something too easy, you don't want something too hard, otherwise you're not gonna see progress.
And so, you know, this is a somewhat evolving set of targets. We talked earlier about the targets that we looked at with, with BoltzGen, now we are even trying kind of, you know, even harder targets, both for small molecule and proteins.
And so we try to keep ourself on the, on the boundary, uh, of what's possible.
So do you have like infrastructure, or this is like you just have a lot of different partnerships with academic labs, and you're just gonna keep pushing on these and driving these?
We do partially this through academic labs. Uh, more and more we do this through, uh, CROs just because of, you know, to some extent is also we need kind of replicability often kind of, you know, going after the same targets, you know, multiple times and, you know, to see the, the progress from, you know, one month to the next.
Um-
And speed.
And speed.
You know, speed of execution. Yeah.
And-
So what happens if you start getting a bunch of like really strong biters against therapeutic targets? What do you do?
Um, I mean-
Release them.
Yeah.
Put that, put that in the line.
Do you release them in open source? Like you-
Yeah, I mean, you know, I mean, we're, we're-- when we say we have no interest in making drugs, we're serious. Like, you know, uh-
Yeah.
I mean, when it, when it was with the academic labs, basically the, you know, it was they keep it, they do whatever they want with it. And with the, with the CROs so far, yeah, we've been, we've been very, yeah, releasing, releasing them.
I, you know, I, I will also say, and, and I think this is a bit been a bit of the issue that I, I have with, with some of kind of the things that have been said in the field is like when we say that we design new proteins or we say that we design new molecules, you know, uh, it, that, you know, go and bind these particular targets, we should be very clear, you know, these are not drugs.
You know-
Yeah, yeah.
These are not things that are ready to be put-
Yes, yeah
... into, into a human, and there is still a lot of development that, that goes with it. And so this is, this is kind of to, to us, you know, we see ourself as, you know, building tools for scientists.
You know, at the end of the day, you know, it really relies on the scientist having a great therapeutic hypothesis and then pushing through kind of all the stages of development. And, you know, we try to build tools that can accompany them, uh, in that journey.
We-- it's not like a, a magic box where, you know, uh, you j- can just turn it and get-
Get FDA-approved drug.
... dr- FDA-approved drugs, uh, FDA-approved drugs. Um, yeah.
But, but actually that brings up an interesting question that I have, I've been wondering about is do you guys see yourself like staying in this, like this s- for lack of a better way of saying it, layer? Or do you think that you'll start to like either on a physical sense, looking at different layers of the virtual cell, so to speak, or, um, also, you know, so there's the, like the development process that goes, you know, sort of like design, preclinical, clinical approval, and thinking about im- like improving the performance throughout that process based on the designs.
Is, is that a direction that you guys are pushing?
Yeah. So one of the things, as Jeremy said, you know, we are not a therapeutic company, and we want to kind of stay not to be a therapeutic company, always be at the service of, you know, all the different, you know, companies, including therapeutic companies that we serve.
And, you know, that to some extent does mean, you know, that we need to try to, you know, go deeper and deeper in getting these models better and better. One of the things that we are doing across, you know, uh, many other, uh, in the field is, you know, now that we are really-- they're starting to be good both for small molecule and for proteins to design kind of binders, design relatively tight binders.
It's starting to look at all these other properties, you know, they call developabilities or ADME that, you know, we care about when developing a drug and try can we, uh, design them from, uh, from the get-go. And the thing about those properties in some of them, you know, um, you need to, you know, s-start having an understanding of, of the cell.
And, and so that's on the one end kind of why we need that understanding. But also, you know, the way that we also think about kind of, you know, um, all different and complex diseases is that these models then these tools that we're building have a good understanding of kind of, you know, biomolecular interactions and kind of their interactions.
Now, at the same time, every disease is often kind of unique, and every therapeutic hypothesis is unique, and so you maybe want to have something that, um, needs to, uh, hit the particular, you know, um, let's say target, uh, in, in a virus in a particular way, but you don't maybe know exactly what, uh, way you want to do.
And so maybe in the first set of designs, you're gonna try to target different epitopes in different ways, and then you're gonna test them in the lab, maybe directly in vivo, and you're gonna see which ones work and which one don't.
And so then you need to bring those results back into the models, and then the models can start to have a more, uh, wider understanding, you know, not just of the biophysical of the, uh, antibodies interacting with that target, but also how that is shaped i- within the entire, uh, the entire cell.
And so first of all, you know, that means on the one end that we need, you know, kind of these loops, and this is also partially how we, we design the platform to be. But that also means that we also need to start understanding more and more kind of higher level things.
And, you know, I wouldn't say that we're working in any way on like a virtual cell like, uh, others are, but we're definitely thinking kind of very deeply about kind of, you know, how does, you know, kind of the way that we target, um, certain proteins interfere, interact with, you know, maybe pathways that are existing in the cell.
One question that has come up is you talk a lot about user interface and so on, and I think this is really important. But like my experience with dealing with medicinal chemists, when you give them machine learning models, is they are the most superstitious, skeptical, like pseudo-religious people I've ever talked to when it comes to doing science.
So-
Sorry for the medicinal chemists listening.
How do you solve the-- Yeah. They're, they're amazing. Like, they're absolutely... So I've worked with some spectacular medicinal chemists who just pull magic out of their hat again and again, and I have no idea how they do it.
But when you bring them a machine learning model, it is sometimes quite tricky to get them to deal with it. How, how has your interaction been with this, and how have you thought about, like, building Boltz Lab to Work with the skeptics.
One of the great value unlocks for us and for our product has been when we brought to the team, um, a medicinal chemist.
Mm-hmm.
His name is, uh, Jeffrey. So I think kind of like on the one end, you know, day one, you know, he obviously had a lot of, uh, opinions on kind of a lot of the ways that we should, uh, uh, change, you know, both kind of-
Mm-hmm
... the way that the agents worked, the way that the platform worked. Uh, but it's been really amazing kind of, you know, once also we started kind of shaping kind of the platform, uh, in a better way with, with his feedback, how we went from, you know, some extent, you know, a fair skepticism to him, you know, actually using a lot more compute than any of our- ...
computational, uh, folks in the team. You know, um, at times that, you know, he's, you know, running, you know... He has all these sort of ter- uh, hypotheses. "Okay, maybe I can hit this protein this particular way. I can hit it in that way.
Actually, let me look at for this particular molecular space, uh, let me try to optimize for these particular interactions." So he ends up, you know, running several screens in parallel, you know, using hundreds of GPUs, you know, on his own.
And, you know, so this has been, you know, pretty incredible to see kind of how, you know, maybe the way that I was more thinking about a problem, which is okay, you're just trying to design a binder, a small molecule to a particular protein.
The way that he thinks about it is, you know, much more deeply and, you know, trying all these different things, these different hypotheses. And then, you know, once he gets the results from the, um, from the model, he doesn't just, you know, take the top 15.
Mm-hmm.
Uh, but he, he really kind of looks over and, you know, kind of tries to understand, you know, the different things. And then when we, uh, select, you know, maybe some designs to, uh, to bring forth, you know, he has, you know, something where, you know, both the models understand that something's good, but all himself as well.
And that's why we also built kind of the platform to be, uh, an interface for, you know, this kind of-
Mm-hmm
... uh, this kind of chemist and, you know, also like a collaborative experience.
I, I think at the end of the day, like, you know, for people to be convinced, you have to show them something that they didn't think was possible.
Mm-hmm.
And until you have that aha moment, you know, I think the skepticism will remain. But then when, you know, every once in a while I think there's like a, a result that like really surprises people and then it's like, "Oh wow, okay, this actually- I can do something with this."
Yeah. So you just get it in their hands, have them try it out, and they'll be convinced.
Yeah. Or like at maybe once-
Some of them
... the lab results come back.
Or their, their friend, yeah, or maybe one of their colleagues is convinced and-
Yeah
... yeah.
I think it, it, it takes going to the lab-
Yeah
... I think at some point. There's no avoiding that. You know, as beautiful as the platform can be, as nice as the molecules might look, you know, that the model predicted, I think what really convinces people is like, you know, hits.
Yeah.
Yeah.
You see the results and yeah.
Exactly.
Yeah. Cool. Thank you for, you know, taking the time to chat with us.
Yeah.
It's been really interesting.
Um, you know, is there anything that you would like your audience to know?
I mean, first of all, you know, uh, we're just getting started, you know, uh, continuing to, to build a team and so, um, definitely always looking for, uh, great folks both on the kind of, you know, software side-
Mm-hmm
... you know, machine learning side, but also scientists-
Mm-hmm
... uh, to join the team and help us, you know, uh, shape.
On the infrastructure side too, like-
Indeed. Uh-
If you, if you think that if you want a new challenge, because this is not just next token prediction, this is really a new engineering challenge-
Exactly
... that hasn't been done before.
If you-- if no matter, you know, how much experience you have-
Mm-hmm
... with, you know, biologists and chemistry, if you want to come, you know, help us, you know, shape what, you know, biology and chemistry hopefully will look like-
Mm-hmm
... in five, 10 years, um, we'd love to hear from you. And so, um, go to Boltz.bio and, you know, come join the team.
Cool. Thank you.
Awesome.
Yeah. Thank you so much.
Thank you so much.
Thanks so much.
Thank you.





