Opening0:00
MD was supposed to be the protein folding solution. There is a great counterexample. The counterfactual is basically a group called Desres, DeeSha Research. They had, you know, similar funding to DeepMind, um, probably more actually. They tested the hypothesis to death that MD could fold proteins.
Yeah.
They built their own silicon, they built their own clusters, they had them taped out all themselves. They burned into the silicon the algorithms to run MD. They ran MD at huge speeds, huge scales. I remember David Shaw came to a conference once on MD, and he flew in by helicopter and just like- ...
to this, this pretty famous guy.
Wow.
Kinda rich.
Yeah.
And, um, he, he gave, uh, an amazing presentation about the special computers and special room and out- out, sort of Times Square, and, like, what they can do with it. It was-
Yeah
... beautiful, amazing, and I always thought that protein folding would be solved by them, but it would require a special machine. Maybe the government would buy, like, five of these things, and we could fold, you know, maybe one protein a day or two proteins a day.
And when AlphaFold came out and it's like you can do it in Google Colab, you know, or on a GPU or desktop- ... it was so mind-blowing. I forget, like, that protein folding was solved. I always thought that was inevitable, but the fact that it was solved and on, like, your desktop, you can do it, was just completely floored, changed everything.
This is the first episode of the new AI for Science podcast on the Latent.Space network. I'm Brandon. I work on RNA therapeutics using machine learning at Atomic AI.
My name is RJ Honicky. I'm the co-founder of MiraOmics, where we build spatial transcriptomics AI models.
The point of this podcast is to bring together AI engineers and scientists, or bring together the two communities. These are two communities which have been developed independently for quite some time, but there's been some attempt to combine them, and only now, after, you know, many years, are we starting to see some of the big developments start to play out in the real world and start to solve, you know, key scientific problems.
There's no, like, one size fits all solution. You need domain expertise. You need k- people on both sides of the aisle who can really talk to each other and really work together and understand both the modeling and all of the real subtleties of the system you're actually trying to work on.
We hope that we can connect these communities and that we can provide a starting point for this new era of AI and science to move forward.
So without further ado, let's get started on the first podcast. We're really happy to have in the studio today Andrew White, co-founder of Future House and newly formed startup Edison Scientific. Um, rather than introduce him, I'll let him introduce himself.
Uh, hey, I'm Andrew from San Francisco, former professor, now running two startups, uh, one that's a nonprofit research lab and one that's a for-profit venture backed company, and we're trying to automate science. We're gonna get into all those points.
Academic Roots2:50
Yeah, really happy to be here. Thanks for having me on. I wanna know personally about jump from academia to industry and- Yeah ... or quasi industry. Yeah. So, like, I would love to hear that story. Yes, um, I guess, like, that's the whole story, right?
Uh, so I did my PhD at University of Washington, and, uh, I worked in a group with, um, I think 19 people doing experiments and, like, two people doing simulations, and I was working on a topic, uh, called molecular dynamics, um, which I think is actually suddenly becoming interesting again as everyone's looking for ways to generate data from first principle simulation, and molecular dynamics, you know, covers, uh, basically everything that's molecules moving around in dynamic systems, so, like, biology, things like that.
Of course, the complement in material science is things like density functional theory, where you can model chemical reactions in these, like, solid systems. So I was working on that, and we were working on biomaterials. And so the goal of my, my PhD was trying to find what are called non-fouling materials.
So in biological systems, whenever you put, like, a, um, foreign object into the body, it will trigger a response, and that response is called the foreign body response. Basically, it encapsulates it in, like, this, um, layer of collagen.
And this actually is exploited for some implants. Like, if you get a heart, uh, sorry, pacemaker installed, like, it coats it with this collagen so that if you go to change the battery, you can almost change the battery out, like, without even bleeding because, like, the body has, like, completely encased it, and this is great for pacemakers.
But for, like, a glucose sensor or, like, a, you know, brain- Yeah ... cognitive interface, BCI is what they call it now. Yeah. Yeah. There, it's not so great, and so that's why some of those things have, like, a limited lifetime because eventually your body treats it as, like, a wound and, and heals.
Rejects it. Yeah. It's, it's kind of like-- So some rejection is, like, immune-based. Okay. And so that's where, like, if the body can see anything on it, like- Yeah ... if it can see, like, some, some, um, uh, ligand that it can bind to- Yeah ...
with, with antibodies- Right ... then you get this, like, inflammation, which is like a rejection response that you see- Okay ... in organ transplants. Yeah. But with materials, the body's just like, "Oh, there's just, like, a wound," or, "There's just some-" Oh, it's-- "...
something here," and it just covers it up. It's like a cut. Yeah. I think, you know, the research in that field has gone on a long time since I left my PhD, and there was a lot of theories about it's related to the mechanical properties in the material, like if it's spongy.
Mm-hmm. There's things like if it's trabecular, like, does it have a bunch of little pores in it? We worked on the theory that it had to do with how hydrophilic the material was. Okay. But anyway, so I was only one working on computers in this group.
I couldn't figure out, like, how to connect what's on the computer with what's done in the lab because you can make, like, a simulation of whatever, 10,000 part, 10,000 particles, 10,000 atoms. It's like, well, this is not gonna model a human body- Yeah.
... in an implant. You know, there's a lot more- It's a little too small. ... a lot more atoms involved. So I had a good time. We did some cool stuff, some bioinformatic stuff. I learned a lot, but then when I did my postdoc, I was like, "Okay, we're gonna try to merge experiments and, and simulations."
So I worked on this, this theory called, um, maximum entropy, and it's about, like, how do you take complex simulations and match them to limited observations? Mm-hmm. And it's like the inverse of machine learning. Yeah. Machine learning is, like, you have simple models, you're m- going to a lot of data- Right ...
where I had, like, complicated models I'm trying to fit to very little data. Yeah. It was fine. It was great. We wrote some papers. It was useful, and then I wrote-- Like, I started my research group at University of Rochester on applying these methods to model peptides.
Yeah. I'm always, like, too early for things. We studied peptides for, I don't know, four or five years, and it was a cool niche field, not that popular. Yeah, yeah. Now peptides are, like, the hottest thing ever. Right.
I think there's even, like, a peptide rave I heard about a couple of- Oh, yeah ... weeks ago. I heard about that, too. But, but- I didn't go though ... when I was an assistant professor, nobody cared about peptides.
Um, so we worked a lot on, on different ways to combine them. We looked at, like, uh, different experimental methods that we could do these molecular dynamic simulations of peptides. And then in, in 2019, I was, like, out on a sabbatical at, at UCLA.
Um, they have a place called the Institute of Pure and Applied Mathematics there, which is, like, this institute where people can go and do a sabbatical and learn new methods, and they were happen- to be doing, they happened to be doing, um, machine learning for physics.
Mm. I think the name of it was like some, like, symmetric thing. It's like machine learning for physics and physics of machine learning. Okay. It's kind of a cool- Yeah, yeah ... cool concept. Right. But, like, Jan Lekan was there and, like, um- Nice.
Yeah ... uh, Frank Noe was there, who's a big guy in Europe in, in this field. Uh, I don't know, it's just... Terrence Tao came by. Just a great group. Yeah. And everyone's kind of jamming. It was, like, 2019, so, like, there'd not really been the big hit- Yeah ...
um, especially in, in, in, uh, in non-computer science fields. Right. And then I came back from that and was like, "Well, I gotta teach a class on this." So writing the book about, like, how you can apply these methods in chemistry.
It was very kind of niche field because, uh, every machine learning class that my PhD students could take at the time, this is when I was professor at University of Rochester, it was always would end in like, okay, this is an RNN, and this is, like, what you need to know.
Or, like, this is how you do image classification. Yeah. But in chemistry, it's all about graphs, right? It's all about how do you represent these graph structures. It's all about symmetry and geometry. Yep. And that was, like, not a thing.
It, it was very popular, but you had Max Welling on before- Yes. ... and so I, like, you know, the, the godfather of, of geometric deep learning. Yeah. So I wrote this textbook about, like, this, these methods, and there was a bunch of interesting mathematics to it.
I had a good time and stuff. And then, um, uh, I think I was following, you know, the news in the space and Codex, the original Codex came out. And, um, I had been looking at transformers for a while.
I just tinkered with them. And we started trying them on doing some chemistry tasks, and we were really impressed actually. And we wrote a benchmark, and this is like 2019 or something. We wrote a benchmark of verifiable rewards in 2019.
Wow. Maybe it's 2020 by then, but, like- Yeah, yeah ... here's a- Little ahead of the curve. Little ahead of the curve, yeah. It's like, here's a, here's a function and, um, uh, sorry, there's, like, a, a task which is like I have a body of a function for, like, a Markov chain Monte Carlo simulation.
It's missing some pieces, complete it. And then we had, like, a verifier that would see did it... is it a valid MCMC simulation? Yeah. Yeah. So it's Markov chain Monte Carlo. Right. We wrote this paper. It ended up coming out, I think, in 2021, 2022, um, because it took a long time to bank enough questions.
ChemCrow8:49
But I wrote an opinion piece about how transformers could change, um, how we think about chemistry and things like this and how we teach it. And then OpenAI f- some people there, um, Lama, uh, w- was there. She saw this paper, and they reached out like, "Hey, we're building this new model, and we think it'd be great to red team it to see, like, what could happen with these models if they're applied to chemistry or biology."
And so I was a red teamer for GPT-4, and I was using it, like, nine months or something before release. It was, like, August. Yeah. So GPT-4 came out in March, and I was using it in August. Yeah.
And then, like, the ReAct and Miracle paper came out, um, I think Shen Yu, he wrote that paper, and- Yeah ... I plugged it in with GPT-4, like, in the fall, you know. Right. Right. And I was like, "Wow, there's so much stuff coming out."
With ReAct? Yeah. Yeah. And it was really exciting. And then so when GPT-4 came out, I released this paper called ChemCrow. I worked with Philippe Schwaller in, in, um, Switzerland on this and IBM. So that was, like, ReAct applied to chemistry?
Yeah. And what we had- Yeah ... is we had, like, there was a cloud lab that IBM built in Switzerland. Yeah. So we had, like, GPT-4 operating the cloud lab, and then it was, like, I had written a literature research agent that did, like, agentic RAG.
Again- Yeah ... nobody knew what agentic RAG was at the time. Yeah. Um, uh, a- and I think actually, uh, uh, Harrison Chase had, like, written a blog post about some ideas there. Yeah. And so I, I, I stole some of the ideas.
A really smart guy. Um, and basically, we applied that, and we saw some really cool stuff. It was really exciting. And then I... And we wrote the paper. It set off this crazy storm of, like, everyone was... had a lot of anxiety about AI progress.
Yeah. And, um, I ended up visiting the White House. Um, I guess my, my paper was, like, the only time a preprint or peer-reviewed paper was presented to the president- Whoa ... on, like, their schedule- ... for, like, a 30-minute block.
Wow. And the national security advisor at the time, um, uh, Jake, uh, God, I always confuse them with the- Epper, Epper. Yeah. No, sorry. One of them is a talk show host, and one of them is, like, the national security advisor.
Oh, yeah. I forget which is, which is which. Um- That guy. Yeah, that guy. Yeah. He, like, had a presentation about our s- our, our paper, and they- Yeah ... presented to all the... 'Cause there was a big tech CEO summit- Yeah ...
at this time where they sent out Sam Altman and some other CEOs out there, and they- This is the Future of Chemistry is Language or a different one? This is, uh, the ChemCrow paper. It's named- Oh, ChemCrow ...
after the paper that we just discussed. Right. Okay. Yeah, yeah. Yeah, yeah. Sorry. I probably should- Yeah ... name these things this. Yeah, yeah. And, and so it was crazy, and they had me go out there, and then I, I met a lot of three-letter agencies I didn't really wanna meet and, um- Yeah.
And then, like, you know, there's, like, some- somebody from, uh, one three-letter agency is like, "How does this change explosives?" You know- Yeah ... the three letter agencies, like, "How does it change breakout time for nuclear- Right ...
weapons research of- Yeah, yeah. I was like, "Guys, like, I don't- Yeah. "I'm not really sure." And but it turns out that, like, you know, there was not that many people world experts on, on- Yeah ... on AI and science.
Right. So what's the answer? Uh, yeah. Uh, great. Um, good question. Uh, we'll come back to that. Uh- Okay, yeah. Let's come back to that. In the end, um, I, I, you know, had a lot of energy, a lot of, a lot of excitement about this area, so took a sabbatical from University of Rochester, and it was Sam Rodriguez.
And, um, Sam had been talking to Eric Schmidt and Tom Kalil, who was, um, uh, also a, a National Security Council at, under Obama administration, about how to, uh, uh, like, scale up these ideas. And so Sam had this concept of, like, focused research organizations- Yeah ...
Future House11:44
which is how do you do science, like, not in academia, not in, in a, um, in, like, one of these kinda near monopoly tech companies- Yeah ... these, these big labs, and, uh, wanna try this idea out. Yeah.
And I was like, "Hey, we should do this around agents for science or AI for science." I love Sam. He pushes me to come up, you know, with really Lofty ambitions, so we decided to automate science as the goal instead of, like, see what fun stuff we could do with agents and science.
Yeah.
But I think that was maybe the real mission, but of course, automating science is the, the long-term mission.
Yes.
And so that was what led to Future House.
Okay. Yeah.
And that was a very long-winded-
Yeah. No, no, that's great.
Yeah.
So but, so you chose to leave a tenure track position-
So I-
... to do this
... I was on sabbatical, which is a, a beautiful concept.
Yeah.
Um, but then I did resign my tenure position when we co-founded Edison.
Yeah. Okay.
And I think-
I see
... it was, you know, I, I had been on sabbatical for a very long period of time, and so at a certain point, I just had to resign my tenure. So I resigned tenure in June.
Okay.
Um-
Oh, so that's only recently.
Yeah, only recently.
Yeah, yeah. And you just felt like this is, this is the direction of my career. I, I need to do this.
Yeah, yeah. I mean, I got tenure, and I had these, these early career awards, like the NSF Career Award. It was great, and I think academia is really exciting, but, um, I just thought that right now, these, this kind of area, like AR science is, um, just I think, A, difficult to do in academia, and B, so exciting that I think you can take bigger bets.
Yeah.
And, um, I think having a tenured position and writing research grants is maybe not the biggest bet you can take on, on a field.
Yeah.
So now we have a venture-backed startup called Edison-
Right
... which we spun out of Future House.
Right.
And, you know, we took a lot of the ideas, and we're trying to do this at an even bigger scale right now.
Yeah.
And, and so Edison was always kind of the plan. Like, going back to Sam's idea of a FRO or a, like, what's it, fundamental research-
Yeah
... organization. Like, he, he always had this goal of, like, let's do fundamental research in this, like, tightly scoped nonprofit, which can kind of explore, and then you have that as a natural arm for spinning off-
Yeah
... um, you know, venture-backed, uh-
Yeah, I think that's right. Um, I think some things that, like, make that not as clean these days is how expensive AI research is and how expensive GPUs are.
Mm-hmm.
So I don't think we can repeat it many times from Future House. It might be, like, an N of one thing right now. It just, may- maybe not. I don't know, if, if-
Yeah
... venture capital keeps growing, then maybe, I mean-
Yeah
... maybe we can. But yeah, I think we took a lot of the ideas of Future House. Another thing, like, I, I think we expected it to be harder to automate science, and actually, it's really hard. Like, I'm, I feel like I'm always miscalibrated in this domain, but it's always hard to predict progress.
Yeah.
Um, and I think that I overestimate the speed of things on month scale, and I underestimate the things on year scale. So the two years from, like, 2023 to 2025 was an enormous amount of progress.
Yeah.
Um, it always felt like things were not going as fast as I thought, but when you look back on it, wow, like, there's a lot of progress. And so I think the idea that, I think at Future House and our, Sam actually regrets us writing this, but in the, the original marketing or, like, the announcement was like, "It's our 10-year mission to automate science."
And then, like, now it's like, okay, yeah.
It's like two year, yeah.
Yeah. So two years later, we, we had Kosmos, and, um, things are going so much faster, and it, also this kind of, this, this thing you, you notice it in, um, in, in San Francisco is, like, it's actually kind of hard to find problems which are, like, so hard that, like, they are a challenge for language models, but not so hard that they're impossible.
And we're in this, like, gray zone, and actually, I feel like that's where we are right now, is that we can actually automate so much of the scientific method because it turns out, especially in a field like biology, which is very empirical limited, you know, the top 1% guesser of, you know, what, what they think will happen in an experiment, the, you know, the top, you know, quintile or quartile, they're about equal.
And so, you know, even if you had an even, we wait 10 years and get even smarter models, I don't think it's gonna really change the fact that we're ready to automate a lot of science with, with existing LLMs.
Automating Science15:54
I mean, what do you mean by automate science? That's, like, a pretty loaded statement. There's lots of things.
Yeah.
Like, there's-
Absolutely
... many ways of thinking about that, so-
So we try to draw a line between, um, what I would call, uh, uh, groups that are trying to model something like the cell or how proteins fold or, like, how antibodies can be designed or, like, maybe the virtual cells, like an example.
If they're trying to use machine learning or AI to model some very specific system-
Mm-hmm
... we're trying to automate the, like, cognitive process of scientific discovery, making hypotheses, choosing experiments to do, analyzing the results from the experiments, and using it to update your hypotheses or your confidence in those hypotheses, and then leading to, like, a world model which of, like, okay, this is how I understand this process to be, and then that, you know, begets new hypotheses or new experiments.
We want to automate that sort of loop. We thought that we would have to build up, like, a whole new organization from the ground up for agents. So it means, like, automated labs. It means, like, putting all the papers in one spot, like getting APIs wrapped around everything.
But, um, over time, like, the model's gotten better and better that we had to, like, you know, stop and rethink. Okay, we don't actually have to hold their hands so much anymore.
Yeah.
Or, like, they don't actually need to necessarily have an automated lab. They can, like, write an email to a CRO or something, or they can, like, tell you what experiment to do, and you can take a video of you doing it and then show it to the model-
Yeah
... and it can be like, "Okay, well, this is, you know, what happened."
Mm-hmm.
So it's been a really interesting experience of, like, sometimes, you know, we over-engineer things, and sometimes-
Yeah
... uh, uh, well, actually, basically, just mostly over-engineer.
So, so I always think about systems, and scientific is a system, like scientific process is a system. I always think about in terms of constraints, right?
Mm-hmm.
And, like, what is a bottleneck in the system? So that, so what is your hypothesis about this, right? Like, in my mind, not knowing a ton, but in my mind, the constraint of the scientific process is the work you do in the lab, and that's sort of notably missing from...
Well, not entirely. You said auto- you mentioned automating lab and whatever. So, like, how are you thinking about this?
Yeah, I, I think you're right, is that, uh, basically the, the best model, um, whatever, Opus 7 or GPT-10, like, um-
Yeah
... it really can only propose the first experiment, maybe slightly more clever, but at a certain point, you just need information, right? Like, some little calculations you can do that, like, there's more atoms in the brain than you could ever simulate, even if you had all the energy from the sun, right?
Like, I think you could simulate maybe 1,000 brains-
Yeah
... in real time with all the energy in the sun 'cause there's too much information.
Yeah.
So science really hits these, like, bottlenecks where you just actually have to go measure things.
Yeah.
We definitely think about maybe like lab in the loop sort of situations. Like, one of our papers, which was called Robin, is that we, like, had an, uh, one of our agents propose an experiment. We did the experiment, and then we had our agent analyze the experiment, then propose the next experiment.
Yeah.
And that kind of loop, I think, is where you wanna get to.
Yeah.
So what is the bottleneck in that? I don't think it's, like, the intelligence of the first experiment. I think the bottleneck might be something... Like right now, I think the bottleneck is something silly like knowing what's the lead time on all the reagents that you need and what is available in the lab, right?
Like, you know, I-
Yeah, yeah, yeah.
I think whether GPT-5.2 Codex Max or Opus 4.5 is gonna do better is probably doesn't matter.
Yeah.
It's just a matter of, like, which one's gonna have all the information of what's in the lab and how much will it cost, how long will it take.
Right.
Um, and also, I guess the, the kind of frontier that I think about for these models is taste, which is, like, a lot of science, and of course we want to, you know, accelerate technology, we wanna improve the economy, we wanna improve people's life expectancies, we want everyone to be happier.
But a lot of what is done in science is, is based around, like, human preferences. Like, why do people study, I don't know, a particular worm? Well, like, there is a theory that by studying the worm it has led to good medicines or it's led to discovering new genes.
But also people studied it in the past, people's careers depend on that worm- ... and people wanna write papers about that worm.
Yeah.
And, and so there's a human element to some of this.
Yeah.
And I think that, that, uh, models don't capture that so well about knowing what is an exciting result and what is a boring result.
I see.
So I think that's like a, a scientific taste. It's like a broad category of, of all these things.
How do you def- Like, do you try to quantify taste in any way? I mean, I know that I, I have some, like, fun anecdotes about this, but maybe, yeah, I'd just like to hear what you think.
Um, yeah, actually, we sat on this idea, we sat on it, but we, like, argued about it for a long time. Sam and I usually every Monday morning at 8:00 in the morning, Sam and I meet, and we're both, you know, caffeinated and ready, and we argue about stuff like this.
Yeah.
And we had a lot of Mondays where we talked about scientific taste. And in the end we're like, "Okay, let's just, let's just do the dumbest thing," which is to, like, have our agents make hypotheses and put them in front of humans and have them be like, "I like this one," or, "I like that one," right?
So we just did, like, whatever, RLHF on, on, on hypotheses. And, um, we learned a lot about how bad RLHF is with people. Just, like, people pay really attention to, to the tone, to the details, to, like, how many specific facts or figures are on the hypothesis, right?
Like actionability about, like, if the experiment is feasible. But what people didn't really pay attention to is, like, I don't know how to describe this, like, if this hypothesis is true, how does it change the world? If the hypothesis is false, how does it change the world?
This like, um, how much information you gain. It's not really information, but, like-
Yeah
... impact or something.
Impact, yeah.
And that really didn't come through from those things. So then we're like, "Okay, well, this is maybe one strategy," and so we had to go back and think about it more. And, um, we then took a pause from that research, and then we made Kosmos.
And then Kosmos, like, has baked into it taste, right? Like, at the end of the day, there will be some report, and we're working on generalizing this. Basically, at the end of the day, we'll be like, "Okay, I made these discoveries."
Then a person will be like, "Great, I'm gonna download that one," or, "I like that one," right? "I don't like this one." And that rolls up to some hypothesis that came earlier in the process.
Mm-hmm.
And so w- we think we can get to end-to-end on this-
Yeah
... as opposed to human preferences.
So you, you mean the feedback loop is the click?
It could be the click. It could be also, like, the, you know, it... We do an experiment. Sometimes in Kosmos you could ask it to end an experiment, and you can see was the experiment a success or failure-
Yeah
... or something like. But I, I guess, like, we, we brought it out of this kind of like hard to quantify, is this a good hypothesis or a bad hypothesis, and into this, like, you can see some downstream consequences of the hypothesis.
So yeah, humans have, I think, a very, like, strongly well-calibrated nose for science. Like, I mean, maybe you could argue there are sociological effects-
Yeah
... like, um, across the community. But ultimately, oftentimes people really, like, good scientists know right off the bat, like, is this going to be likely to be, um, useful or not? How long, how many attempts did it take before you started to, like, see results that to yourself seemed useful?
Like, you've been working on this for, I guess, two years now. Um-
You know, I think when the Co-Scientist paper came out from Google, I think it was a really interesting idea to do this, like, tournament style or just pairwise ranking of hypotheses, right? So I think Co-Scientist is a very interesting counterexample to what we built.
What we built is something with either lab in the loop or data analysis in the loop or literature research in the loop, where you're, like, iterating on an idea. And then Co-Scientist took this, like, very different approach of, like, let's list all the ideas and then try to come up with a filtration process to come up with the best hypotheses.
Mm.
So Co-Scientist will produce these very long reports of, like, "Oh, we, like, really tested this idea with lots of dialogue," and, and, and it's very interesting stuff. And I was really impressed with the paper that came out. And then we had this Robin paper, and one of the things that came out of the Robin paper is that the hypothesis that people thought was best was not the one that led to success in that paper.
Interesting.
It was in, um, age-related macular degeneration or ocular, uh, age-related macular degeneration. So basically it's like part of the eyes. You're going blind because you have this, like, accumulation of debris in the eye and can't clear it out.
That's one of the major cause of blindness in people over-
Yeah
... uh, 60. Uh, Oli, who works on the hill-
Yeah, yeah.
He'll, he'll cringe when he hears me say that, but-
Something, something like that.
Something like that.
Yeah.
Sorry, Oli. Um, in that one, like we, we went to optometrists or ophthalmologists. Actually get confused on that as well. Sorry, Oli. Um- But essentially, you know, asked them, "What hypotheses do you think are, are, are good hypotheses?
What do you think would lead to, like, um, a good mechanism for treating dry AMD?"
Yeah.
And, um, yeah, there was, you know, they agreed maybe on the top 10.
Yeah.
But, but beyond that, it was, it was kind of noise.
Yeah.
Um, and then, you know, the, the... what we found was ribose dutil was a very good, um, medicine and, and had a, a mechanism that I think is novel, although there was lots of debate on X.
Yeah.
Because in, I think in 2012, there was a master's thesis which proposed this mechanism on like page 38. I actually think it was a typo. I think they meant wet, wet AMD. But anyway, I won't, I won't belabor the point.
I will concede that, that maybe was... there was one reported, uh, example of it in the past. That was a, a really eye-opening experience for me because that was a, the first really serious test where we really went to the lab, and we, you know, spent, like, four weeks on a battery of experiments to see-
Yeah
... what is, uh, which hypothesis led to a good mechanism and a good repurposed drug.
Right.
And it was not as correlated with human opinions as I expected.
Opinions. Yeah, yeah.
And so since then, I think that, uh, I have a lot more faith in these, like, um, verifier-in-the-loop kind of scenarios where you have either data analysis, literature search, or you're running a unit test or whatever. You're, you're, you're going and running the experiment.
Anything like that, I think, is, is gonna give you a higher signal than the sort of vagaries of, like, "Oh, this is a higher opinion," or "That we like this one better."
Yeah, Max, Maxwell called it nature's computer.
Yeah.
It's like-
There you go. Yeah
... it's like you have this computer computational cycle you're running, and the nature is part of that computation cycle.
Yeah, yeah.
I'm curious, so you said that there is a paper which maybe, like, could propose, maybe propose where this, um, molecule came from, but, like, do you have some way of interpreting or, like, understanding where that hypothesis originated in the absence of that?
Like, is there-
Um, oh, from the model?
... a traceable thought train?
Yeah, yeah, yeah. Actually, this is something we pay really close attention to at, at Future House and at Edison, was, um, provenance of, like, information. So our first sort of, like, agent was, was PaperQA. Sorry about the name.
PaperQA sounds like an EL7 that was an agent.
It, it really does.
Yeah, yeah. Um, PaperQA was, uh, has, like, every sentence that it, it outputs has a citation to a page, right? So it's like a lot of provenance. And then we basically built along a philosophy for everything. So Robin would, which is the name of this, I don't know, workflow or something you can call it, that led to this result on repes Gl being a good, um, therapeutic for dry AMD.
Mm-hmm.
It has, like, data analysis that goes, shows you, like, which line Python code-
Mm-hmm
... led to the result here, and then that is like, okay, then it goes to this other model, which says, "Well, based on this literature finding and this result from the data analysis, I believe this is the right thing."
But, you know, where does the original idea come from? Like, going after these ROCK inhibitors, which is the, the mechanism for the, the target, was basically enumeration. And so this is like, if you can't be smarter, you can be, I don't know, you can try more times.
Brute force.
Brute force.
Yeah, yeah.
And I think that was, like, the theory of, of the Robin paper was that we can put out a whole bunch of hypotheses-
Yeah
... and then we can filter them, just like I think is about how CosScientist did, is you go for a filtration process. But the difference is that in CosScientist, their filtration process was other LLMs sort of ranking it with rubrics or, like, personas, and our filtration process was, like, literature search and data analysis.
Um, like, here's some data. Is it consistent with the data? Go see if anyone's discovered it in the paper, in the literature, or if they've disproven it.
Yeah.
And I think that's the easy way to succeed in, in AI over humans is, is you can try more ideas faster.
Something I've, I've heard people say, and maybe I've experienced this in my own life, um, like, sometimes hypotheses are kind of cheap, especially in, you know, biology.
Yeah.
It's many ways actually easy to come up with what you think could be happening.
Yeah.
And, um, it seems like to me verifying is oftentimes a big bottleneck and maybe the biggest bottleneck. Like, if you have lots of hypotheses, you know, and it costs, you know, one-one hundredth of your runway to test each one of them or something, you don't have many shots on goal.
Yeah.
Yeah. So how do you make sure that, like, you, you are actually enriching for good hypotheses?
Literature and data analysis, right? You know, like-
You just, you just do.
Yeah.
Yeah.
There was a time when we used, um, something called tiling trees, and tiling trees is like a literal brute force method invented by Ed Boyden, Sam's PhD advisor, and basically the idea is, okay, I wanna, like, accomplish X.
Okay, I could try these methods. And then, like, once you pick, "I'm gonna try this method," then you split into, like, two different paths. I'm gonna use this method or not use this method.
Yeah.
If I'm using this method-
Mm-hmm
... I need to have, like, I don't know, some kind of substrate. I'm gonna try this substrate or this substrate or this substrate, right? And you can basically try to really, like, tile space of all the ideas. We tried some early experiments there, and you're right.
You run into this thing where some of the hypotheses that come out just don't make any sense, and, like, you are going to waste a ton of effort if you actually test them all. Nowadays, I actually would argue that if you go to an LLM, and you ask it to evaluate, you know, hypotheses, including some garbage ones, it will probably do as good of a job as an expert in the field in, in filtering them out.
That's not always the case.
Yeah, I've actually seen that myself.
Yeah, but there's a lot of gotchas-
Yeah
... and I think people can miss those.
Yeah.
But I think they're actually pretty good. And so I'm not as worried about hypotheses that can fail fast by an expert looking at them.
Mm-hmm.
I think now the filtration process really happens in, in, in literature, and I think the filtration process happens in looking at, like, um, you know, like, uh, a biobank data or, like, you know, what do we know from GWAS or something.
Yeah.
You know, you know, other sources of, of existing data as much as you can draw upon.
Yeah. So yeah, and with regards to existing data, um, another, like, maybe contrarian, uh, take is that, like, oftentimes the, the hardest part is just understanding the, like, the context of data and, like, where it comes from and how do you interpret it.
Like, I can also think from my own life multiple cases where, you know, the, the data in some sense was, like, there, and you had two people who were both experts and very smart people who looked at it and drew very different interpretations.
And in fact, like, when we were interviewing Heather Culic, she, she had some fun stories about using LLMs and, um, she would find that there would be raw data in a paper which wouldn't agree with the conclusions of the actual paper, and it's straight from the paper.
It's not even, like, cross-paper talk or something.
Man, I'm gonna be a really boring interviewer and be like, "Yes, you're right." You know, like, this is, this is a hard question.
These things are hard. Yeah.
Yeah.
Um, I think, you know, to put, give you something concrete, um, we have a, a, a bioinformatics benchmark we call BigBench.
Mm-hmm.
BigBench is like we put it out. We've updated a few times. It's in, it's in some frontier, um, uh, LLMs when they release their system card, they'll mention BigBench as, like, one of the things they test on.
Yeah.
And, um, you know, we're getting to 60%, 70% correctness on BigBench.
Mm-hmm.
And, um, we found that actually we're at the point where humans disagree at this level. Like, humans only agree 70% of the analysis.
Yeah.
And so it's true that, like, when it comes to analyzing data, like humans do not agree 100% of the time. There is a certain amount of, like, choice that goes into it. And, you know, we, we try to, um...
So Edison is a, is a for-profit company. We like trying to sell, uh, some of this stuff to, to companies, and we'll go to some companies that like, "Oh, well, we never impute data. Imputing data is bad." Like, or, you know, whatever.
And then like, okay, well, we'll have to change our agents so we don't impute data for them. And then some of the companies are like, "Oh, yeah, we impute data. It makes everything easier," right? Or, uh... And, and, and, you know, you wonder what the real modern dark arts are, the, like, AI-resistant area of the world is, like, uh, medicinal chemistry.
That is, like, the spot where, like, you know- ... there's so much superstition.
Oh, yeah, yeah.
Yeah.
Everyone, yeah.
It is, it's-
Everyone is, like, pseudo-religious.
Yeah, exactly. You have to be to survive.
Yeah.
And I feel otherwise you get burnt out.
But the religions never agree, too.
They... Yeah, exactly, yeah.
Two medicinal chemists will have completely different viewpoints about, like, a functional group.
Yes, exactly.
Yeah.
And I remember this is, uh... Talking to somebody who works with a CRO, and they're like, "Oh, whenever, like, company X orders anything, we never put boron on any of the compounds- ... because they hate boron," because there was one program-
Yeah
... that was killed because there was a boron, you know, somewhere in the core, and it led to some s- toxic side effects. So no boron for this company. This company, they, like, love things to be fluorinated or something-
Yeah
... because they love... think it's great for, um, the admin properties, right? And so there's, like, all this, this, this stuff where you reach the point where you're at, I don't know, human bias level or human disagreement level, and I think we're getting to that point in data analysis.
And so, of course, you will see then that if I take the raw data from a paper and I analyze it myself, I will get a different conclusion. One of the cool tricks you can do is... This is back to this brute force thing, is that I can go to our agent and I can run it 100 times, and I can take the consensus, like, analysis.
Or I can say, even if you make these three different choices in your data analysis, you get the same conclusion, right? Or this conclusion is somehow s- sensitive to those choices. And then you can... There's even, like, words.
It's like epistemic versus aleatoric-
Oh, yeah, yeah
... uncertainty, right? It's like this is aleatoric, which means, like, I think it's noise from the data, or this is epistemic uncertainty, which means, like, I think there's some choices that are being made. There's some model differences that lead to the, the disagreement.
Anyway, there's, like, a, there's, like, a Donald Rumsfeld formulation of this as well. Like- ... the known unknowns-
Yes, yeah
... and yeah, it is, it's aleatoric, epistemic, uh, uh, uh, debate there.
Interesting. This kind of digging into your Kosmos model-
Yeah, yeah
Inside Kosmos33:05
... a little bit. So I, I glanced at the paper, and one of the things that jumps out is that there were a certain class of problems for which, uh, it was only 50-some percent accurate. And-
Oh, yeah, yeah.
And can you talk a, a little bit about that and how that... Like, okay, so if I'm just raw getting 50% accurate answers, then, and then I'm going to the wet lab and being like, "Okay, try this," and then it's like, "Ah."
Yeah.
Like, "The, the stupid thing did- told me to do a dumb thing."
Well-
How do you-
I would say, first of all-
Yeah, yeah
... that 50%, it's actually pretty good because it's rare that experiments in the lab are actually coin tosses, right? There are usually a lot more outcomes than, than, you know, than, than binary. Um-
Yeah, yeah. Sure. Okay, yeah.
But, but that particular number was, uh, human agreement in the interpretation of the results.
Okay.
And so, uh, we asked people to evaluate different aspects of Kosmos. We had them evaluate, like, the data analysis decisions. We had people ask to evaluate the literature. Like, is this, are these, is... Do you agree with its finding in the literature?
That number that was 50%, that came from Kosmos' interpretation of, uh, some of the analysis.
Yeah.
So, like, it might go in literature and find this result, and it would say, "Wow, this is super exciting. This is amazing."
Yeah.
Or it might do data analysis, be like, "This is a novel discovery. We're really excited about it."
Yeah.
And then people would disagree. "That's actually not interesting," or like-
Yeah
... "I don't agree with the interpretation of it."
So it's like picking bad problems maybe.
Yeah.
In the, the, the negative-
Yeah
... class. Yeah.
And so I think it's like that, that 52 or 55, whatever it is, that's, um, interpretation. And so I agree. I think that's where, like I was saying, I think the frontier right now-
Yeah, this kind of taste
... is scientific taste.
Yeah, yeah. Taste.
Yeah. And so that's what we're working on right now, is-
I see. Yeah, yeah
... how do you get the interpretation to, to match? Um.
Interesting.
Could you just step back and just introduce Kosmos-
Yeah, yeah
... from a high level? Yeah, yeah. Um, it's, it's-
It-- I would actually be in-
Yeah
... even curious to hear starting from, like, ChemCrow and, uh, you know, you have, uh, PaperQA, Avery, Ether 零 .
Zero. Yeah, yeah.
I'd like to hear a little bit of the, the lineage and how those different decisions were made. What were the key learnings, and how did you get to where-
Yeah
... you are now?
Yeah.
Yeah.
So I could retcon and tell a really great story about how we arrived at Kosmos. But I will say that, like, to a large extent, we just try a lot of stuff, and sometimes-
Yeah
... it works, and sometimes it doesn't.
Yeah. Okay.
You know, I'll, I'll, I'll say that we're very... I'm, I'm a, I'm a builder. Like, I like to, like, build things piece by piece. I'm, uh, probably some fancy word for it, but I'm like a Lego guy or something.
Yeah.
My vision was that we would make an agent that does this part of the scientific process, an agent that does this part of the scientific process, whatever. And so we had, like, um, ChemCrow, which is gonna help us with setting up our medicinal chemistry work.
We had ProteinCrow, which we haven't released. I don't know if we will ever release, but ProteinCrow was, like, bu- designing proteins we might need for, for some part of our workflows. Um, or we had a data analysis agent.
Is that LLM, or that's a-
It's an agent, so an LLM plus tools.
Okay.
Or we had Ether 零 . It was like, okay, we noticed that the frontier models can't work with molecules very well, so let's make a, a model with intuition for medicinal chemistry, and that was what led to Ether 零 .
But then, uh, Sam actually really pushed on us to like, "Let's just even do the whole thing. You know, let's, let's just try to build an AI scientist. Let's just try the whole thing." And, and that was what led to Robin.
And, um, Robin was like, "Let's just take these agents we already have, and we'll just put them in, like, a, a, a workf- basically it's like you could express it in a concise Python file of like, you know, try a whole bunch of ideas, then go see if they all filter through literature or if they've been disproven, and then go, like, uh, come up with experiments that you could do in a wet lab."
Yeah.
"And this is our inventory list. And then go analyze all the data, then go back and repeat the process." Right? So that's, like, what Robin was. And, um, and we, like, we came acro- a- across Kosmos. We were trying to, like, understand what is the process that Robin is, is automating?
And it came from this idea of, like, a world model, which is that when we first started Edison, we were thinking, like, what do we, what do we want to change about this? Like, what is new here?
Yeah.
And so we spent some time thinking about, well, the scientific process, like, what is actually going on in, like, my brain, which is that I have some understanding of-
Yeah
... of the world or the-
Yeah
... phenomena I'm studying, and that's my world model. And then a lot of the actions I take are about trying to update that world model. And it's something that changes over time, and so this is, like, this ability to change over time, but it's also something that is practical.
Like, I can use it to make predictions about, I know from this experiment, this will happen. That's why it's like a model and not just, like, you know, a, a memory or like-
Yeah
... a bunch of, like, papers or something like that. It's, like, uh, it's supposed to operate. In Kosmos, we tried this idea out, and actually, um, uh, uh, Ludo, uh, who's the, the, uh, first author on the paper, we tried a whole bunch of ideas around world models.
And, uh, we kind of thought they weren't really appropriate. Like while we tried a lot of different ways to do this, we tried, you know, method A, method B, method C, and, and they were okay. And so we all just had to take a break.
Ludo, like, his project didn't work on trying to do this world model stuff, but he's like, "I'm gonna keep trying it." Ludo's a very stubborn person. So he tried it-
Yeah
... for, like, I don't know, a week or two weeks, then he-
Yeah
... just kind of, like, quietly was like, "Hey, can you guys t- come take a look at this?"
Yeah.
And, and we're like, "Wow, this is actually really cool," and then we, like, started building on it a- and jamming, really. And I think what Ludo figured out is that you have to get this, like, experiment loop thing.
You have to be a- let it... Then the data analysis agent is what got us in the loop.
Yeah.
So if you put that in the loop of, like, it can really update this world model because the... We were trying to build it around literature before. And when you build it around literature, there's just, like, not really experiments you can do and then see the results for.
That was, like, our surrogate was, was literature.
Yeah.
It just wasn't working.
Right.
But data analysis actually really lets you explore ideas. And so that was what led to Kosmos, and so in Kosmos, we basically, we had all the pieces sitting around. We've been working on world models, we've been working on a data analysis agent, working on a literature agent, um, and then we're working on, you know, we built a platform for scientific agents.
So we had things that can write a LaTeX report. We had things that can make nice plots.
Yeah.
Then we put that all together, and, like, a world model was, like, sort of the, the glue that allowed it to, to fit together.
Yeah.
An analogy is, like, um, in coding agents, like, GitHub is sort of the glue. Like, there's some shared repo, and everyone works on the repo, and, like, software engineers have spent whatever, lots of brain cycles thinking about what's the way to coordinate, you know, and organize working on code together-
Yeah
... for a long time.
So the world model is actually like a memory system kind of sort of?
Yeah, you can think of it as a memory system.
Oh.
Um, we, we think about it as, as a model. So, like, it actually-
Yeah
... you can put in-
Yeah, sure
... input, and it will output predictions.
I see. Okay.
And we, we think about calibration.
Yeah, yeah.
But, like, really, it is a set of, like, a big bundle of information-
Right
... that we accumulate over time-
Yeah, yeah
... that's distilled in some way, and, and that is, like, uh, what allows us to do this. And I think you can think about, like, um, a GitHub repo as, like, it's a distillation, right? Like, really, there's a long graph of commits that lead up to it-
Yeah
... and, like, the current file system in that GitHub repo-
Yeah
... or, I keep saying GitHub. I'm-
Yeah
... such a corporate- ... uh, shill here. Git-
Yeah
... your Git repo is like a, a distillation of all of the work that people put in into the PRs, into the, into the, uh, the commits. And so I think there's a nice analogy, uh, between a Git repo and, and what a world model is.
I see. Yeah.
And I think that's just sort of what allows us to automate scientific discovery so well.
Can you talk about, like, kind of how you implement a, a world model, or is that sort of like secret sauce?
That's our, like, secret sauce right now.
Okay.
You know?
That's, that's fine.
Yeah. No, it's fine. People have asked for less.
That question never happened, okay?
So one thing that's notably missing-
MD Overrated40:01
Yeah
... is the, like, simulation, right?
Yeah.
Molecular dynamics or-
Yeah
... or, or, like, uh, bolts or...
Yeah. I wanna, uh, help you guys pump up your views here. So I, I think molecular dynamics is overrated.
In fact, coming from someone-
Yeah
... who's been-
Yes, yes
... that's like-
That goes in the, that goes in the intro
... doing it.
In the thumbnail, you know? Yeah.
But just-
Yeah. And, and, and DFT is overrated. In fact, DFT may be even more overrated than molecular dynamics. I think these methods-
You mean for materials or for biology or for both?
For materials.
Okay.
And I can explain more about that. Basically, MD and DFT have consumed an enormous number of PhDs and scientific careers-
Yeah
... at the altar of, you know, the beauty of the simulation.
Also, random interjection.
Yeah.
Once I, I did an estimate, I think pre, like, ChatGPT, something like 20% of the world's computing power just went to simulating water.
Oh, my fucking God, water.
Yeah, yeah.
I had to deal with so many water simulations.
Yeah, me and water.
I did, I did DFT simulations of water, and they are so annoying. I used these big computers, uh, from, uh, the Department of Defense, and we s- I spent, like, I don't know, five months... And by the way, this is pre-LLM training days.
Yeah.
Five months of compute is actually a really long time.
This is huge.
Yeah.
I simulated water with quantum, you know, effects, with the groatest mechanism for how a proton hops through water, and it's on YouTube. It's my number one YouTube video, and it represents, like-
Until now.
And, yeah. And it represents, like, I don't know, a million CPU hours of compute. It was, you know, one of the biggest computes that I, uh, probably the biggest one I've done in my life so far. Maybe EtherZero is bigger, but, but it took a lot more work.
Anyway, and, and what's the point?
Yeah, what'd you learn?
You know, what's the point?
What'd you learn from that?
All I learned was, like, what set of hyperparameters reproduce some physical effects of water. But n- none of it was de novo, right? And this is the, this is the issue with, with molecular dynamics and DFT, is that, um, they don't model the world correctly, and so we have to invent little stories we tell ourselves about we're, like, making good inductive biases, and then it models the world more correctly.
Like, in DFT, you simulate water at 330 Kelvin-
Mm-hmm
... when you want room temperature water. Is room temperature 330 Kelvin? No, it's not. That's a little too hot, right? And so this is a... The issue is that, like, people just make up these, these, these things, or, like, I don't know, GGA or, like, BEILIP or B3LYP, all these different, like, methods people invent.
They're clearly empirical, and then they bolt it on to DFT, and they say, "Look, it's a first principles method," right? But actually, you made a whole bunch of choices-
Yeah
... and, you know, you, you, whatever, overfit to the validation data-
Mm-hmm
... to get this to work. And, and that's, I think MD and DFT are like that because if you go look at the catalysts, you know, what catalysts change the world, none of them are single crystal materials that are really well suited for DFT.
They're always, like, they have grain boundaries. They have dopants. They're complicated, right? And, and you'll never capture that with DFT. So I think this is one of the- Fundamental, I don't know, dichotomies of the world is that simulations simulate really boring things really well.
They don't simulate interesting things very well. And so that's why I don't do DFT and MD anymore.
What about somewhere like the machine learning-y stuff like AlphaFold and, and-
Alpha was trained on X-ray crystallography data.
Yeah.
And I think, you know, this is the, this is the story of MD, is that MD was supposed to be the protein folding solution. There is a great counter example. There's a... I don't know what there's a word.
The counterfactual is basically a group called Desres, Dee Shaw Research. They had, you know, similar funding to DeepMind, um, probably more actually. They tested the hypothesis to death that MD could fold proteins.
Yeah.
They built their own silicon. They built their own clusters. They had them taped out all themselves. They burned into the silicon the algorithms to run MD. They ran MD at huge speeds, huge scales.
Yeah.
I remember David Shaw came to a conference once on MD, and he flew in by helicopter and was like- ...to this, this pretty famous guy-
Wow
...kinda rich.
Yeah.
And, um, he, he gave, uh, an amazing presentation about the special computers and special room and out- outside of Times Square, and like what they can do with it. It was-
Yeah
...beautiful, amazing. And I always thought that protein folding would be solved by them, but it would require a special machine. Maybe the government would buy like five of these things, and we could fold, you know, maybe one protein a day or two proteins a day.
Yeah.
And when AlphaFold came out and it's like you can do it in Google Colab, you know, or on a GPU or desktop- ...it was so mind-blowing. I forget, like, that protein folding was solved. I always thought that was inevitable.
But the fact that it was solved and on, like, your desktop, you can do it, was just completely floored, changed everything.
This is like the, the bitter lesson on steroids.
Yeah, I don't even know what it is. But it's like, imagine ChatGPT came out, but instead it was like, "Oh, you can just run it on your phone or locally on your own desktop." Like, that's the level of, like-
Yeah
...shock that came out.
Yeah.
And it, it gets down to this thing that humans are really bad at estimating problems that aren't human-made problems. Protein folding, we all thought was like, would require a huge amount of compute, very challenging problem, the most hardest problem in the world, right?
And it turns out that you can actually do it on, I don't, I think the numbers are now like ten thousand GPU hours, you can train a, a good protein folding model. It's actually turned out to be barely an inconvenience.
Therefore, why not?
Oh, oh, therefore, protein folding was-
Yeah
...highly efficient based on experimental data. It-- They took X-ray crystallography. They-- That's what DeepMind did, is they took X-ray crystallography data.
Yeah.
Desres tried the first principles method.
Yeah, yeah.
And it's like a nice head-to-head comparison-
Yeah, yeah. Okay
...two very well-resourced groups.
Yeah, yeah. Yep, yeah.
They both tried different ideas, and the machine learning on experimental data beat out first principle simulation by, you know-
Yeah
...a very large margin.
And so why isn't, uh, like Bolts or whatever inside of Kosmos? Like, why isn't there a tool that's, that can run Bolts?
Oh, we have Bolts inside of... We have Bolts Gen, Bolts Gen.
Yeah.
Yeah, yeah, we have that inside of Kosmos.
Okay. Oh, okay, it is then. That's it.
I mean, I think in the version that we have, uh, for people to just sign up and use, it's not in there.
Yeah.
But, like, uh, you know, you can imagine that you can just Modal or Lambda or Tamarind or Three Ten, there's all these companies that basically-
Yeah
...wrap a lot of these, these like, um, uh, deep learning protein design tools or chemistry design tools. They wrap them in an API, and you just give that to, to-
Yeah. Right
...give it to Cloud Code if you want.
Yeah.
You can give it to Kosmos, and you can be like, "Hey, you know, if you wanna design a protein for X, use these tools."
Your mechanism, it sounds like, or one of the primary mechanisms that has been successful is, like it, like, enumerate a whole bunch of possibilities and filter.
Yeah.
Right? And so how do you think about serendipity and out of, out of distribution thinking and getting there, and how far have you gotten and what's left to do and whatever?
Safety46:26
Yeah. That's a great question. I think... I guess the, the short answer is that there's very f-- So, so this is the domain of CBORNS, so chemical, biological, radiological, nuclear, um, uh, weapons, or-
Yeah
...I don't know, safety.
Yeah.
This domain has been explored a lot in history by a lot of organizations.
Yeah.
And, um, I would say that there was a, a big question mark for us a few years ago, was like, how much of this stuff is, uh, intellectually bottlenecked?
Yeah.
Like, how often are people like, "Oh, well, I wanna cause harm, um, but I need to know, like, some facts," and-
Yeah
...could LLMs make that easier or go faster or anything like that? I think, you know, the first set of answers in twenty twenty-three, I think, was basically no, is that, like, you know, you can go find the synthesis route for many dangerous compounds on Wikipedia.
Yeah.
People know what are the targets in the human body that, like, are, are, are targeted by most biological weapons. It's, it's not really that much of a mystery. So I, I don't think there was a lot of, like, um, there was a lot of new ground when LLMs first came about.
Then there's a lot of concern about, like, laboratory protocols. Is that could agents or LLMs, uh, reveal some tacit knowledge that, like, maybe people couldn't find on Wikipedia or, like, maybe for making something, there's some technique that req- is required when you scale it up in size or something.
Or maybe there's, like, some way to get around, like, tracking lists by ordering different compounds or something.
Yeah.
So, and that, I think, was really well tested, um, by a few different labs. Not, not me, but there were some groups that spun up that started making, like, tests for this, and it, and labs pay attention to it.
I think it's really been put into process where LLMs will, like, kind of shut down or be filtered in, in those scenarios. But I think that is actually an area where there is, is some risk. Um, and so I think that's something that people pay attention to for open source models, and there's still, I think, some, some discussion there.
But I think to a large extent, it's, it's not really been greatly accelerating in practice, or at least I haven't seen much evidence of it.
Yeah.
Um, and again, I think it comes down to the fact that it's not really available, but, like, you, if you look hard enough, you can find most of the information you would need to, to get up to no good-
Yeah
...in the public domain already. But then I think now i- is, the, the next frontier is, like, uh, can it somehow help you with real-time protocols, troubleshooting, like, more in the loop and-
Yeah
...um, and, and more especially in the computational side of things. There are some scenarios that are now coming into focus that could be more dangerous or more intellectually bottlenecked. And so I think people are trying to pay attention to that.
To some extent, there was like a first wave that we thought this could unlock a lot of stuff, and I don't think it came to pass.
Yeah.
I think there's now an emerging sort of second wave of, like, there are some actually new scenarios that were just too far-fetched to consider two years ago that I think are now realistic. Um, some smart people are paying attention to it, but I don't think it's solved yet.
Yeah.
I don't know. Uh, it's very vague.
Yeah.
No, I mean, so, so I guess, like, one kind of differentiator, there's a lot of talk about AI safety-
Yeah
... and, like, the broader LLM, you know, ASI space. And, you know, there it's jokes about paper- or paperclip maxing robots-
Yeah, yeah
... or something. But, like, the, the core threat here is more, like, a malicious actor using this as a tool to accelerate something dangerous. And, like, kind of the first order hypothesis is that you basically already have to be an expert to effectively create a bioweapon or a chemical weapon.
Oh.
And a non-expert... Or an expert would already know how to do this.
Yeah, I, I think, you know, so, so each of the categories in the CBRN, they're, they're all a little different. But I think to a large extent, it's a, a lot of, like, pushing material around.
Mm-hmm.
You know, the classical example of nuclear is, like, it's a lot of, lot of centrifugation.
Yeah.
A lot of ultra centrifugation, a lot of high pressure or high RPMs.
Yeah.
And so it- it's just, you can maybe get smarter about how to set up, you know, the, the economy of scale to do that with an LLM. But to a large extent, I, I think you can call your, your friend in country X, and they can tell you what are the steps.
It's, it's not a... I don't think it's that much of a secret. It's just a lot of, like, moving material around, and I don't think it's acceler- meaningfully accelerated. Now, that said, there are all kinds of like, you know, dumb dual use things of, like, maybe you wanna call a company that makes centrifuges, and you wanna make sure that they sell you them, and they go through some KYC steps.
And maybe an LLM can get you through the KYC faster. And that's, like, a dumb thing that like, okay, like, yes, like, uh, you know, email makes it so that you can order centrifuges off the internet more easily.
Is email, like, a dual use technology? Like, yeah, to some extent it is. And so I think there's a lot of, like, weird second-order things that we don't pay attention to in AI safety of, like, does it make KYC easier?
Does it make it easier for people to know, like, where, where to order this from? Or, like, what is the expected price? Or, like, what should you order first, right? All those, like, sort of simple logistical things I think are accelerated by AI, just as, like, a, a consequence of AI being an accelerating technology.
Um, but certainly, I mean, shit, guys, there's some scary stuff and- I try not to think about it too much and, uh-
Yeah.
Yeah
... I don't know. I guess I don't wanna get too political, but I do think that right now, um, the, the United States government is maybe taking a, a slower, less intensive look at safety. And, um, but there's definitely people, I think, in other spaces than the US government thinking about it hard.
And do you think it's a thing people need to spend more time on? Mm-hmm.
I do get waves of angst about AI.
Yeah. Yeah.
And I'm sure many people living in San Francisco do get like a little bit of, uh, a little bit of waves of it. And, uh, sometimes I think that there isn't enough work being done on it. And then sometimes I think, "Wow," like, "I need to mellow out," and, like- ...
you know, we have lots of time to think about it. What is my opinion on it then? I don't know. I, I, I think my opinion is, um, not formed fully.
FROs52:13
Yeah. You and Sam have done a lot of thinking about funding science.
Yeah.
And future of science. You have, you've been vocal about the reproducibility crisis and other things. First question, why this focused research organization or for, uh, yeah-
FRO.
Yeah.
Focused Research Organization.
Yeah. What, what does that get you that you don't get from academia or, you know, big lab or whatever?
Um, a nice network of, of people, and I think Edison is, like, a real, uh, of course, I think Edison's gonna do great, but I think it's a mystery of what's gonna happen. Um, so I don't think we've had as much friction there as you might expect.
Um, but yeah, this is all stuff that we, that, that, that Sam and I think about all the time is, like-
Yeah
... how do you balance stuff like this? How do you balance the economics? Um, you know, there are some, there are some venture-backed companies that are s- having cash salaries over a million dollars, and it's, like, insane to me-
Yeah.
Yeah
... that you would use all of your cash from your equity financing, you know, in these insane salaries.
But they can s- in, in terms of, like, total spend on GPUs, it can still be a total, a small fraction of your burn, so sometimes it kinda makes sense.
Yeah. Yeah. That's, that's one way to think about it.
So, so, w- like, you... This is a good, uh, lead-in to y- you are automating science-
Yeah
... in some capacity.
Yeah.
So where does that leave scientists?
Scientists' Future53:34
So I think, um, this is a Jevons paradox we can try here- ... is, uh, um... So, uh, let me start with the contrast here is that, uh, you know, if we automate, um, you know, taxi cab drivers, uh, there's a fi- there's not gonna be an increase in people needing to go places.
Maybe there'll be somewhat an increase, but, like, there is a finite amount of, like, time people will be spending in cars.
Yeah.
And so there's an upper limit. So when you automate that, that's like a scarcity thing, is basically you're displacing jobs when you automate driving.
Yeah.
In science, I don't think there is a finite appetite or a finite capacity for science. I don't think science is, like, a, a scarcity thing. Like, there's, you know, 100 more discoveries left to be made, and then we'll be done, and so, like, we're displacing jobs.
I think instead, actually, if we can, you know, make science go much, much faster, there will be no, there will be no decrease in demand. There will be actually, I think, an increase in demand that will match whatever automation amount we have.
And so my vision for what a scientist would be in the future is that they will be, I don't know, like, uh, agent wranglers or Kosmos wranglers of, like, okay, they're exploring 100 ideas simultaneously, or they're, like, working with systems like ours to, to make 10X the discoveries, 100X discoveries, because I think there's an unlimited amount of scientific discoveries to be made.
And so there's no, like, scarcity set where basically we will displace them all. Now, that's kind of like, you know, this is what I would tell when I go talk to a first-year PhD student.
Yeah.
Like, "Everything's gonna be just fine." You know, but then when it gets into the nuts and bolts, I, I do agree that this is gonna be, like, a really hard thing, where, like, if I am CEO of a company that makes science, like a pharma company or a material science company or something like that-
Yeah
... or a R&D arm at, at IBM, I think while I could spend, you know, a million more dollars on, on compute for the AI scientist, or could hire 10 more people, I might just choose to go with the AI scientist because, you know, to a large extent, like, hiring people is hard, right?
Yeah.
Yeah.
And, and hiring an AI scientist is probably a little bit easier.
Yeah.
And so I think that there could be some, there could be some friction. But another thing is, like, science is in some ways closer to art in the sense that, like, there is a large number of people who just appreciate good science.
Like, if you get published in Nature, it's not because it's really gonna be world-changing. Of course, that's part of it, but it's also because, like, people are like, "Wow, this is really interesting science."
Yeah.
So I think the, the enjoyers of science are also scientists, and so I think that it's kind of hard to imagine a scenario when there's not scientists as the consumers of science. And so I think if they're gonna be consumers of science, they're also gonna be some of the producers who are involved in the, in the process by itself.
Right. Yeah.
I don't know if that makes any sense.
Yeah. You touched on this. My- the question in my mind is just what does a scientist do then?
There's a great short story, um, by, um, Ted Chiang, I think in like 2003 or something.
Okay.
And it's about, like, well, at first, scientists were displaced, and they became, like, the, uh, interpreters of, like, what the AI scientists are doing. Like, the scientists read the AI scientists', like, papers and then, you know, translate them for whatever Popular Science or something.
Yeah, yeah.
And then after that, like, they couldn't read the papers anymore, and so they were left behind, and so they had nothing to do, and they just sat around. And but the problem is that science is like, you know, you, you have to translate science to make any impact.
Like, science cannot exist by itself. I do agree there's, like, engineering can exist by itself. Like, if you give some kind of system a goal of, like, making a new material that it can make a space elevator out of, you could be not participating in the beginning of the process or the middle of the process, and you just come by the end and be like, "Okay, follow this recipe."
But, like, science of, like, what's the origin of life? Or like, is there water on another planets? Or, you know, um, why is some catalyst better than another catalyst? That has to be hitting human eyes and human brains at some point, so I think a human has to be involved in the process.
Don't wanna be contrary but-
Yeah, be contrary
... why, why does a human have to be involved?
Why does a human have to be involved?
Yeah.
Well, a human has to be involved at at least some point to be like, "Yes, this is good science," or, "This is bad science."
Oh, okay. So y- y- it's, it goes back to taste.
Yeah. But I don't know. Maybe you're right. Maybe there is no point for humans. Maybe it will be like, you know, what is it? Sora. Uh, you know, like the AI slop app. But I think in Sora, there's still humans at the end clicking the videos or something.
Yeah.
So.
So, so the bring, the, the Sora analogy kind of brings up a interesting point. Like, is it possible that, like, due to the biases of AI science, if we really go full in science, that, you know, there still is a market for kind of boutique human-
Boutique
... science, like-
Yeah, yeah
... you know, there's still people who wanna, you know, paint things the old-fashioned way. But more to the point, does it become even more important for to have a human who is, uh, actively doing their own exploration because there will be, like, large blind spots and biases due to the models that just you'll never be able to overcome because this is sort of baked in now, um, due to your training data.
Yeah.
And without a human, y- that, you'll always get stuck, and there will be a blind spot that will never...
Bio, which is a company in, in Oakland, um, or in Emeryville, they do really cool stuff with automation. I think they're gonna be testing this theory of like, okay, maybe if that's the bottleneck, we can see evidence of it because they're gonna start doing really well.
Yeah.
Um, it could be true.
Mm-hmm. I, I still though wanna say, all of those I, in my mind, are still sort of scoped in terms of, like, R&D for pharma or bio, but they're not like, none of them are attempting to answer big fundamental questions.
And maybe there's, like, different levels. When I think about the-
Yeah
... you seem to be, um, y- it seems like the Future Hou... the, the focus of Future House and Edison is much more towards, like, you know, sorta R&D and sort of end-run science. But, um, you know, I, I have some background in, you know, fundamental physics.
Yeah, yeah.
Um, you know, it's like, is there any thought about, like, how do you, like, take on, you know, dark matter candidates? And like-
Yeah
... I just, you know, think the data to really give us a complete story is just not there yet.
You know what? Like, uh, I'm sure everybody at every company f- like, is the biggest critic of their own product-
Yeah.
You know?
Yeah.
Yeah.
So-
Yeah
... we think Kosmos is, we think it's great, but there's an, a very large amount of area for improvement and
S- so with Kosmos, can... So th- there's, like, a open, like, sort of access to everybody version.
Yeah.
Uh, do you provide access to other labs that, um, is less open?
Um, we have a version of Kosmos- ... that has, like, um, bigger resources. Like it can-
Yeah
... run for longer. It uses GPUs.
Yeah.
Um, so, like, basically when it does data analysis, it'll have a GPU.
Yeah.
So we use that for things like, um, like machine learning experiments. You know, if we wanna know, like, this question about whether it's better to pre-train first on noisy data or not.
Yeah.
Um, we have, like, pre-release models that, that are coming out, and we try those. But, um, yeah, so I guess-
Yeah
... like, yes, we do. When we do have, like, research partnerships with, with companies where we, like, build something specific for them, and that is something we think about.
Yeah.
But broadly, I would say Kosmos that's on the website is pretty close to-
Yeah
... to what is the best we have internally.
Yeah.
Uh, I have a question. Um, so you, you previously have stated that you think that language is the natural, um, language. Was it-
Language Argument1:00:44
Language of chemistry
... the future of chemistry is language.
Yeah, yeah.
Yeah, yeah. Um, okay. So, uh, I wonder, do you still believe that?
Good question. I think, I, I w- I would say yes. I still believe that, um- That, so, so in that article, that opinion article, my, my point was that, uh, you know, at the time when I wrote that article, which I think maybe three years ago now or something, maybe 2023, um, it, it was that we have models for predicting solubility of compounds.
We have, like, data about very large populations, and we have, like, papers, and we have code. And, and the only way to bridge all that information is natural language. And, and the argument was that, like, humans, like, you know, whenever we can't bridge information, like if I can't talk about my code or I can't talk about some idea to you, I will invent words until I can get the point across, right?
Mm-hmm.
And that humans are always innovating on language to make it represent all known observations, and people innovate on language to represent whatever code pattern they have, right? Like, this is like the, the only shared activity we've been doing for this long is like coming up with words to represent everything we know.
Mm-hmm.
And so I think for that, for that reason, natural language is the only possible way to connect all the different pieces of data we need in biology and medicine, or any domain for that matter. Um, I think there's some caveats to this of like, you know, you can make an argument, like if Yann LeCun were here and he would make an argument about like, you know, world models or like vision or embodiedness, right?
Like the, there's arguments against natural language that like, you know, that maybe there's something more that it, it does, it's not the complete story. Or maybe natural language imposes limitations-
Yeah
... you cannot exceed because you are stuck in this abstract space that was invented by humans and you can't escape it, until you can like touch something.
Yeah. I mean, it, it is an abstraction, right? And like, like scientists basically work exclusively in abstractions to some degree.
Yeah.
Um, I, I just, I find, I found that interesting because it seems like most scientists you're right, like when they explain things, they explain things through language. But, uh, many conversations, maybe most at some point result in people drawing diagrams or something.
Yeah.
Like, you know, chemistry, like biochemistry largely, or, or medicinal chemistry is oftentimes a, it's, it's a language of graphs, right?
Yeah.
Or, you know what I mean, bonds are abstractions, yes, but like they're pretty good abstractions for most ca- for many cases.
Yeah.
Or like, you know, geometry, you know, thinking about, you know, protein as like a, the geometry of a protein.
Yeah.
You know, it's like I think that that's how people will... a lot of scientists like to think about things. And, um, so I find it interesting that like, yeah, that, that you are focusing primarily on language. Like have you thought about essentially a multimodal version of this?
Like where, you know, when it comes along a smiles per string, it doesn't just say, "Oh, this is a smile string," but like, this is a graph, this is a representation of some higher like abstract object.
You're absolutely right. And, and the problem with these, this like, I don't know, Jacob's Ladder or something, whatever you wanna call it, is like yes, you can say that a mol- you can call a molecule by its name.
Mm-hmm.
You can show the graph.
Yeah.
Then if you go to a molecule like ferrocene, well, it doesn't really have bonds-
Mm-hmm
... between like part of it. And so then you're like, well, we need to draw it visually.
Mm-hmm.
And then you go to a molecule like, I don't know, cyclobutane.
Mm-hmm.
Well, there's dihedral angle, right? And so like it's not actually this thing I drew, it's actually an ensemble between this thing and this thing, right?
Mm-hmm.
Then you go to benzene, you're like, well, not only is it like a ensemble of these different conformers, it actually has electron density and you can't really ignore the electron density in benzene. You like need to treat it correctly.
And then it's like, well, you can't actually represent the electron density that way. You actually have to look at the correlation of the electrons individually, right? Because you can't really model benzene with like DFT, right? Or functional, you have to actually look at the, the electron correlation.
Then you see the electron correlation like, well, you know, you can model electron correlation, but you know, actually these things, when they're in a solution, they have like, you know, relativistic effects because it's like there's a whole bunch of stuff around it.
So you really gotta have the relativity in there. And you're like, well, you can have the relativity and you have the electron correlation. You could have the bonds and, and you have the conformers, but you really need to think about the cosmic radiation background because like, you know, it does actually impact everything and there is some, some energy there, right?
Mm-hmm.
And before you know it, you've ran out of, you know, you've ran out of compute or whatever resource you're using to model this. And so I think, um, you have to draw the line somewhere.
Mm-hmm.
Natural language, like I said, is that humans have worked for a long time to make it be the, you know, what's the word? Like the least abstract- ... or the, you know, it's somewhere on the border of like it's still abstract enough that you don't need to know all these details.
Mm-hmm.
But it's still granular enough or con- concretized enough that you actually can make use of it. Um, there may be some other representation like multimodal. Might turn out the video or maybe, I don't know, there's some other like fusion that you can make.
I like natural language because we all work really hard to make it right at that boundary. And I do agree, sometimes, sometimes ideas slip and they can't be in language, and you have to get out the whiteboard. Or ideas slip and you have to wave your hands around, you know?
Or-
Mm-hmm
... maybe then, then you need that, that, uh, degree of freedom to communicate.
Just digging in on this a little bit-
Yeah, yeah
... more. Like, uh, famously quantum mechanics is like undescribable, right? Like the, there's, there's an argument that you cannot understand quantum mechanics with words. It ha- or in, and with our preconceived understanding of the physical world because it doesn't behave like the macroscopic world.
And so the, the only way to understand it is through mathematics, right? Um, and I largely see language as the joint key of science as well, but I wonder if that's not true for many domains, and quantum mechanics is just the one that hits you in the face.
I mean, I don't know. Actually, I think the, what, there's like seven principles of quantum mechanics or five or something like this, that you can actually express pretty concisely in language.
Mm-hmm.
I agree that like you need to actually look at the consequences of them. You need some mathematics. Um, I don't know. I actually, I don't know. This is like a challenge.
Yeah.
I think you could actually describe a lot of quantum mechanics in language.
Sure, sure.
But, but I, I see your point and, um, yeah, I, I guess, uh, uh, I'm a realist. Like I, when I talk to my kids, you know, maybe I, I will be like, "Okay, let me draw it for you."
I don't, I don't make sure in our house everything is described with natural language. Uh, so I, I agree with you there. Um, I think maybe we can be a little, a little flexible with, with natural language and include equations and smiles strings in it-
Mm-hmm
... and I think we can get a little bit farther. Um, so maybe that's okay. Uh, but Some people, I think, like optionality. You know, like the, "Oh, it could be this or it could be that." I'm somebody that like to take, take strong opinions- Yeah ...
and see how much farther they can get me. And I think in, in my career, it's actually been better for me to take strong opinions- Right ... which in my deepest of hearts I know that are maybe not correct or not fully correct.
But once you take these strong opinions, it just, you can sort of move many steps down the road once you take these strong opinions. Right. And like, for example, at Future House, we took the opinion that scientific agents are the future, and that allows, skipped a lot of steps, because a lot of other people were like, "We need to bu- build a foundation model for X."
Yeah. And we just skipped all that, right? Right. And I think if you also were unopinionated and you had optionality, like I can think of a famous example of a different company that, like, liked the optionality, and they wasted a lot of time on foundation models for something, then, then I think you, you get stuck.
So that's one of my strong opinions is that natural language is a, a, is- Yeah ... a way to join all these different domains. Yeah. It may not be a correct opinion, it may not... It may be more subtle or more complicated, but it's allowed me to get very far.
Um- Yeah ... maybe I'll drop it someday and- ... maybe find a new one, but yeah. Not yet, though. That's my meta- Yeah ... opinion- Yeah ... on the matter.
The EtherZero story on your blog I find hilarious and kind of awesome.
EtherZero1:08:16
Yeah.
You know, when I was a kid, I loved the, like, g- genie/monkey paw-
Yes
... like, concept of be careful what you wish for because you just might get it.
Yes.
Maybe just, like, quick story. Can, can you just talk about that? That was, that was just a really fun-
EtherZero was a, a hell of a project because conceptually it was a very short project of like, hey, people have made a lot of progress in verifiable rewards in math and in comput-i-i-in, in code, let's see if we can do it in chemistry.
So chemistry is, like, not a verifiable field, right? Like, of course you can go test it in a lab, but then we like had to think about all these, like, ways that we can make chemistry verifiable. And one of the ones we settled on was like, make a molecule that has like three nitrogens, two oxygens, 10 hydrogens or something.
And we thought that was, like, a pretty verifiable, ver- pretty verifiable question. But every time we would train a model, it would find some new insanely weird trick to generate these molecules. And, and I, I'll just tell you one of the examples was that, um, uh, it would make these molecules, and we would do some checks to make sure, like, it had the right bonds, the right number of electrons, the right number of atoms and stuff like that.
Um, but it would just solve the problem in any way possible, right? Yeah. So, like, it would just put all the nitrogens over here, put all the oxygens over here, just, like, things that don't look good. Yeah. And so we started coming up with these rules of like, oh, let's check to make sure it followed these good practices or these good practices.
And we found ourselves into this, like, you know, it's like the opposite of the bitter lesson, like, I don't know, the boutique lesson- ... where you, like, try to make everything custom. But one of the things it kept doing is it kept putting these nitrogens in a row, and it put like one nitrogen, two nitrogen, three nitrogen all in a chain.
And this is, like... You know, if you have three nitrogens, it's, like, explosive, you know- ... two nitrogens is, like, bad, and, like, four nitrogens you can't make. And it kept telling everyone, like, it would make these, like, six-nitrogen compounds, and they're just, they're just literally impossible, and they're not possible.
And, uh, many of the people on the team were, like, computer scientists, like, on this team, and one of them, like, one day sent me the, like, "This is on the cover of Nature today on Nature's website. Somebody made a six-nitrogen compound."
"And this is, like, somebody's, like, career to deliver this compound because this is the most unstable, like, insane compound you can make." "It's some ridiculous setup, and it, like, the spectroscopy to get that proven was, like, very difficult-" Yeah ...
"and this, this... I don't know how they did it. It was an amazing accomplishment." I'm like, "Look, Andrew, like, it's not actually impossible." And it was so funny to me that, like, our model is sitting here spitting out these six-nitrogen compounds in, like, you know, 2024 or 2025, and, like, the paper just happened to come out that year- Yeah.
... that, like, mankind had finally made a six-nitrogen compound. So do, do you think that those were actually synthesizable even under these extreme circumstances? No, no, our model was just, it was just- It just got- ... reward hacking.
Okay. It, it was just the, the model was so creative in ways to reward hack. Yeah. Like, one of the o- another one we did was, um, you know, we, we wanted it to make sure that the, when it would propose a reaction, like make this compound, tell me how to make this compound, we would try to make it sure that all the reagents were purchasable, like you could purchase them.
They were not, like, made up. Yeah. Um, and, and the reason we came up with that is that originally we would just, like, take the end- ... compound and then, like, remove one atom and be like, "Here's- Buy this ...
buy this." And then put the atom on, and it's like, okay. It's like, well, that's r- I wish it was like that. Um, so they, they had to be purchasable. And then well, we're like, we thought it might be hard if they're all purchasable because sometimes you actually order things custom or, or something.
So we went, "We'll just make sure one purchasable." So the first thing it starts doing is just putting nitrogen in there because nitrogen is purchasable, and it, like, has no participation in the reaction, right? Like, oh my God, okay.
So then like, okay, it has to be purchasable, it has to participate in the reaction. Then it started just putting, like, acid base chemistry. We'll just put an acid here. Acids are purchasable, and it'll move one atom. And then we go, "Okay, fine.
Can't be that. Everything has to be purchasable." Then we find ourselves, and I'm, like, sitting there one day building this, like, ridiculous catalog of purchasable compounds in a bloom filter so it can go fast enough in our training loop- Yeah.
... and I'm like, "Why am I doing this?" How did I get here? How did I get here? And, and I don't know. It was really funny because, um, pre-training or training transformers, you know, o- o- on, on just data, like just supervised training where you just have the inputs and the outputs directly- Yeah ...
very nice, relaxing. Yeah. You know, like, things are always robust. You know, things are, uh- Yeah ... go pretty smoothly. When we do these verifiable rewards where you have to, like, write a, a bulletproof verifier, it is really difficult.
Yeah. And we had so many models trained only to find out they were hacking some other, like, random thing- Yeah ... in our setup. It's really hard, and I, and I, I don't envy the frontier labs that have to do this at a very massive scale- Yeah ...
because we had a lot of adventures in EtherZero. And, and you guys should read the, the blog post. It-
Yeah, definitely read the blog post.
Yeah. It is very fun.
It is a great read.
GRPO? We did make some modifications, um, to GRPO. Yeah. Um, I actually, I used to know all the names of these modifications. Yeah. But, uh, uh, I think it's like, uh, DAPO is one modification, and, like, the clipping we did was special.
I mean, we explored a lot of that stuff. Yeah. Um, and it was, uh, also one of these things where, like, you think the hypers are wrong, the algorithm is wrong, and then you find out it's just because, like, you had somehow sorted the reagents when you made your training data, but in your test data you didn't sort them alphabetically, and the model was just, like, barfing because its whole strategy was to exploit something in the way you sorted things.
So yeah- Yeah ... we, we explored a lot of different methods, and it was, um... I learned a lot about chemistry, a lot about nomenclature. Um, and actually there's a, I learned a lot about medicinal chemistry as well, more than I ever wanted to.
Awesome. If you wanna do some, like, engineering, just check out Edison Scientific, and they have, you know, I think a lot... They're hiring with lots of, like, interesting things, everything from scientists to, you know, infrastructure engineer.
Yeah.
Yeah. Thanks, Andrew, again.
Yeah. Thank you very much for, for joining us.





