LALatent SpaceJan 28, 2026· 1:13:56

🔬 From Red Teaming GPT-4 to Automating Drug Discovery: The Future of AI in Science — Andrew White

Andrew White, co-founder of Future House and Edison Scientific, argues that automating the scientific method with LLM agents is now feasible, explaining how ChemCrow triggered White House briefings, how Kosmos uses a world model to generate and test hypotheses, and why EtherZero's reward hacking revealed the difficulty of verifiable chemistry tasks. He shifts from his academic work on molecular dynamics to building agents that enumerate and filter ideas, claiming scientific taste remains the frontier. White recounts the counterexample of D.E. Shaw Research's MD vs. AlphaFold, asserts that natural language is the universal bridge for scientific data, and predicts that automation will expand rather than eliminate scientific jobs.

  1. 0:00Opening
  2. 2:50Academic Roots
  3. 8:49ChemCrow
  4. 11:44Future House
  5. 15:54Automating Science
  6. 33:05Inside Kosmos
  7. 40:01MD Overrated
  8. 46:26Safety
  9. 52:13FROs
  10. 53:34Scientists' Future
  11. 1:00:44Language Argument
  12. 1:08:16EtherZero

Powered by PodHood

Transcript

Opening0:00

Andrew White0:00

MD was supposed to be the protein folding solution. There is a great counterexample. The counterfactual is basically a group called Desres, DeeSha Research. They had, you know, similar funding to DeepMind, um, probably more actually. They tested the hypothesis to death that MD could fold proteins.

Brandon Anderson0:18

Yeah.

Andrew White0:18

They built their own silicon, they built their own clusters, they had them taped out all themselves. They burned into the silicon the algorithms to run MD. They ran MD at huge speeds, huge scales. I remember David Shaw came to a conference once on MD, and he flew in by helicopter and just like- ...

to this, this pretty famous guy.

Brandon Anderson0:37

Wow.

Andrew White0:37

Kinda rich.

Brandon Anderson0:38

Yeah.

Andrew White0:38

And, um, he, he gave, uh, an amazing presentation about the special computers and special room and out- out, sort of Times Square, and, like, what they can do with it. It was-

Brandon Anderson0:46

Yeah

Andrew White0:46

... beautiful, amazing, and I always thought that protein folding would be solved by them, but it would require a special machine. Maybe the government would buy, like, five of these things, and we could fold, you know, maybe one protein a day or two proteins a day.

And when AlphaFold came out and it's like you can do it in Google Colab, you know, or on a GPU or desktop- ... it was so mind-blowing. I forget, like, that protein folding was solved. I always thought that was inevitable, but the fact that it was solved and on, like, your desktop, you can do it, was just completely floored, changed everything.

Brandon Anderson1:13

This is the first episode of the new AI for Science podcast on the Latent.Space network. I'm Brandon. I work on RNA therapeutics using machine learning at Atomic AI.

Andrew White1:24

My name is RJ Honicky. I'm the co-founder of MiraOmics, where we build spatial transcriptomics AI models.

Brandon Anderson1:30

The point of this podcast is to bring together AI engineers and scientists, or bring together the two communities. These are two communities which have been developed independently for quite some time, but there's been some attempt to combine them, and only now, after, you know, many years, are we starting to see some of the big developments start to play out in the real world and start to solve, you know, key scientific problems.

There's no, like, one size fits all solution. You need domain expertise. You need k- people on both sides of the aisle who can really talk to each other and really work together and understand both the modeling and all of the real subtleties of the system you're actually trying to work on.

We hope that we can connect these communities and that we can provide a starting point for this new era of AI and science to move forward.

Andrew White2:17

So without further ado, let's get started on the first podcast. We're really happy to have in the studio today Andrew White, co-founder of Future House and newly formed startup Edison Scientific. Um, rather than introduce him, I'll let him introduce himself.

Uh, hey, I'm Andrew from San Francisco, former professor, now running two startups, uh, one that's a nonprofit research lab and one that's a for-profit venture backed company, and we're trying to automate science. We're gonna get into all those points.

Academic Roots2:50

Andrew White2:50

Yeah, really happy to be here. Thanks for having me on. I wanna know personally about jump from academia to industry and- Yeah ... or quasi industry. Yeah. So, like, I would love to hear that story. Yes, um, I guess, like, that's the whole story, right?

Uh, so I did my PhD at University of Washington, and, uh, I worked in a group with, um, I think 19 people doing experiments and, like, two people doing simulations, and I was working on a topic, uh, called molecular dynamics, um, which I think is actually suddenly becoming interesting again as everyone's looking for ways to generate data from first principle simulation, and molecular dynamics, you know, covers, uh, basically everything that's molecules moving around in dynamic systems, so, like, biology, things like that.

Of course, the complement in material science is things like density functional theory, where you can model chemical reactions in these, like, solid systems. So I was working on that, and we were working on biomaterials. And so the goal of my, my PhD was trying to find what are called non-fouling materials.

So in biological systems, whenever you put, like, a, um, foreign object into the body, it will trigger a response, and that response is called the foreign body response. Basically, it encapsulates it in, like, this, um, layer of collagen.

And this actually is exploited for some implants. Like, if you get a heart, uh, sorry, pacemaker installed, like, it coats it with this collagen so that if you go to change the battery, you can almost change the battery out, like, without even bleeding because, like, the body has, like, completely encased it, and this is great for pacemakers.

But for, like, a glucose sensor or, like, a, you know, brain- Yeah ... cognitive interface, BCI is what they call it now. Yeah. Yeah. There, it's not so great, and so that's why some of those things have, like, a limited lifetime because eventually your body treats it as, like, a wound and, and heals.

Rejects it. Yeah. It's, it's kind of like-- So some rejection is, like, immune-based. Okay. And so that's where, like, if the body can see anything on it, like- Yeah ... if it can see, like, some, some, um, uh, ligand that it can bind to- Yeah ...

with, with antibodies- Right ... then you get this, like, inflammation, which is like a rejection response that you see- Okay ... in organ transplants. Yeah. But with materials, the body's just like, "Oh, there's just, like, a wound," or, "There's just some-" Oh, it's-- "...

something here," and it just covers it up. It's like a cut. Yeah. I think, you know, the research in that field has gone on a long time since I left my PhD, and there was a lot of theories about it's related to the mechanical properties in the material, like if it's spongy.

Mm-hmm. There's things like if it's trabecular, like, does it have a bunch of little pores in it? We worked on the theory that it had to do with how hydrophilic the material was. Okay. But anyway, so I was only one working on computers in this group.

I couldn't figure out, like, how to connect what's on the computer with what's done in the lab because you can make, like, a simulation of whatever, 10,000 part, 10,000 particles, 10,000 atoms. It's like, well, this is not gonna model a human body- Yeah.

... in an implant. You know, there's a lot more- It's a little too small. ... a lot more atoms involved. So I had a good time. We did some cool stuff, some bioinformatic stuff. I learned a lot, but then when I did my postdoc, I was like, "Okay, we're gonna try to merge experiments and, and simulations."

So I worked on this, this theory called, um, maximum entropy, and it's about, like, how do you take complex simulations and match them to limited observations? Mm-hmm. And it's like the inverse of machine learning. Yeah. Machine learning is, like, you have simple models, you're m- going to a lot of data- Right ...

where I had, like, complicated models I'm trying to fit to very little data. Yeah. It was fine. It was great. We wrote some papers. It was useful, and then I wrote-- Like, I started my research group at University of Rochester on applying these methods to model peptides.

Yeah. I'm always, like, too early for things. We studied peptides for, I don't know, four or five years, and it was a cool niche field, not that popular. Yeah, yeah. Now peptides are, like, the hottest thing ever. Right.

I think there's even, like, a peptide rave I heard about a couple of- Oh, yeah ... weeks ago. I heard about that, too. But, but- I didn't go though ... when I was an assistant professor, nobody cared about peptides.

Um, so we worked a lot on, on different ways to combine them. We looked at, like, uh, different experimental methods that we could do these molecular dynamic simulations of peptides. And then in, in 2019, I was, like, out on a sabbatical at, at UCLA.

Um, they have a place called the Institute of Pure and Applied Mathematics there, which is, like, this institute where people can go and do a sabbatical and learn new methods, and they were happen- to be doing, they happened to be doing, um, machine learning for physics.

Mm. I think the name of it was like some, like, symmetric thing. It's like machine learning for physics and physics of machine learning. Okay. It's kind of a cool- Yeah, yeah ... cool concept. Right. But, like, Jan Lekan was there and, like, um- Nice.

Yeah ... uh, Frank Noe was there, who's a big guy in Europe in, in this field. Uh, I don't know, it's just... Terrence Tao came by. Just a great group. Yeah. And everyone's kind of jamming. It was, like, 2019, so, like, there'd not really been the big hit- Yeah ...

um, especially in, in, in, uh, in non-computer science fields. Right. And then I came back from that and was like, "Well, I gotta teach a class on this." So writing the book about, like, how you can apply these methods in chemistry.

It was very kind of niche field because, uh, every machine learning class that my PhD students could take at the time, this is when I was professor at University of Rochester, it was always would end in like, okay, this is an RNN, and this is, like, what you need to know.

Or, like, this is how you do image classification. Yeah. But in chemistry, it's all about graphs, right? It's all about how do you represent these graph structures. It's all about symmetry and geometry. Yep. And that was, like, not a thing.

It, it was very popular, but you had Max Welling on before- Yes. ... and so I, like, you know, the, the godfather of, of geometric deep learning. Yeah. So I wrote this textbook about, like, this, these methods, and there was a bunch of interesting mathematics to it.

I had a good time and stuff. And then, um, uh, I think I was following, you know, the news in the space and Codex, the original Codex came out. And, um, I had been looking at transformers for a while.

I just tinkered with them. And we started trying them on doing some chemistry tasks, and we were really impressed actually. And we wrote a benchmark, and this is like 2019 or something. We wrote a benchmark of verifiable rewards in 2019.

Wow. Maybe it's 2020 by then, but, like- Yeah, yeah ... here's a- Little ahead of the curve. Little ahead of the curve, yeah. It's like, here's a, here's a function and, um, uh, sorry, there's, like, a, a task which is like I have a body of a function for, like, a Markov chain Monte Carlo simulation.

It's missing some pieces, complete it. And then we had, like, a verifier that would see did it... is it a valid MCMC simulation? Yeah. Yeah. So it's Markov chain Monte Carlo. Right. We wrote this paper. It ended up coming out, I think, in 2021, 2022, um, because it took a long time to bank enough questions.

ChemCrow8:49

Andrew White8:49

But I wrote an opinion piece about how transformers could change, um, how we think about chemistry and things like this and how we teach it. And then OpenAI f- some people there, um, Lama, uh, w- was there. She saw this paper, and they reached out like, "Hey, we're building this new model, and we think it'd be great to red team it to see, like, what could happen with these models if they're applied to chemistry or biology."

And so I was a red teamer for GPT-4, and I was using it, like, nine months or something before release. It was, like, August. Yeah. So GPT-4 came out in March, and I was using it in August. Yeah.

And then, like, the ReAct and Miracle paper came out, um, I think Shen Yu, he wrote that paper, and- Yeah ... I plugged it in with GPT-4, like, in the fall, you know. Right. Right. And I was like, "Wow, there's so much stuff coming out."

With ReAct? Yeah. Yeah. And it was really exciting. And then so when GPT-4 came out, I released this paper called ChemCrow. I worked with Philippe Schwaller in, in, um, Switzerland on this and IBM. So that was, like, ReAct applied to chemistry?

Yeah. And what we had- Yeah ... is we had, like, there was a cloud lab that IBM built in Switzerland. Yeah. So we had, like, GPT-4 operating the cloud lab, and then it was, like, I had written a literature research agent that did, like, agentic RAG.

Again- Yeah ... nobody knew what agentic RAG was at the time. Yeah. Um, uh, a- and I think actually, uh, uh, Harrison Chase had, like, written a blog post about some ideas there. Yeah. And so I, I, I stole some of the ideas.

A really smart guy. Um, and basically, we applied that, and we saw some really cool stuff. It was really exciting. And then I... And we wrote the paper. It set off this crazy storm of, like, everyone was... had a lot of anxiety about AI progress.

Yeah. And, um, I ended up visiting the White House. Um, I guess my, my paper was, like, the only time a preprint or peer-reviewed paper was presented to the president- Whoa ... on, like, their schedule- ... for, like, a 30-minute block.

Wow. And the national security advisor at the time, um, uh, Jake, uh, God, I always confuse them with the- Epper, Epper. Yeah. No, sorry. One of them is a talk show host, and one of them is, like, the national security advisor.

Oh, yeah. I forget which is, which is which. Um- That guy. Yeah, that guy. Yeah. He, like, had a presentation about our s- our, our paper, and they- Yeah ... presented to all the... 'Cause there was a big tech CEO summit- Yeah ...

at this time where they sent out Sam Altman and some other CEOs out there, and they- This is the Future of Chemistry is Language or a different one? This is, uh, the ChemCrow paper. It's named- Oh, ChemCrow ...

after the paper that we just discussed. Right. Okay. Yeah, yeah. Yeah, yeah. Sorry. I probably should- Yeah ... name these things this. Yeah, yeah. And, and so it was crazy, and they had me go out there, and then I, I met a lot of three-letter agencies I didn't really wanna meet and, um- Yeah.

And then, like, you know, there's, like, some- somebody from, uh, one three-letter agency is like, "How does this change explosives?" You know- Yeah ... the three letter agencies, like, "How does it change breakout time for nuclear- Right ...

weapons research of- Yeah, yeah. I was like, "Guys, like, I don't- Yeah. "I'm not really sure." And but it turns out that, like, you know, there was not that many people world experts on, on- Yeah ... on AI and science.

Right. So what's the answer? Uh, yeah. Uh, great. Um, good question. Uh, we'll come back to that. Uh- Okay, yeah. Let's come back to that. In the end, um, I, I, you know, had a lot of energy, a lot of, a lot of excitement about this area, so took a sabbatical from University of Rochester, and it was Sam Rodriguez.

And, um, Sam had been talking to Eric Schmidt and Tom Kalil, who was, um, uh, also a, a National Security Council at, under Obama administration, about how to, uh, uh, like, scale up these ideas. And so Sam had this concept of, like, focused research organizations- Yeah ...

Future House11:44

Andrew White11:59

which is how do you do science, like, not in academia, not in, in a, um, in, like, one of these kinda near monopoly tech companies- Yeah ... these, these big labs, and, uh, wanna try this idea out. Yeah.

And I was like, "Hey, we should do this around agents for science or AI for science." I love Sam. He pushes me to come up, you know, with really Lofty ambitions, so we decided to automate science as the goal instead of, like, see what fun stuff we could do with agents and science.

RJ Honicky12:23

Yeah.

Andrew White12:23

But I think that was maybe the real mission, but of course, automating science is the, the long-term mission.

RJ Honicky12:28

Yes.

Andrew White12:28

And so that was what led to Future House.

RJ Honicky12:30

Okay. Yeah.

Andrew White12:30

And that was a very long-winded-

RJ Honicky12:31

Yeah. No, no, that's great.

Andrew White12:33

Yeah.

RJ Honicky12:33

So but, so you chose to leave a tenure track position-

Andrew White12:36

So I-

RJ Honicky12:37

... to do this

Andrew White12:37

... I was on sabbatical, which is a, a beautiful concept.

RJ Honicky12:40

Yeah.

Andrew White12:41

Um, but then I did resign my tenure position when we co-founded Edison.

RJ Honicky12:44

Yeah. Okay.

Andrew White12:45

And I think-

RJ Honicky12:46

I see

Andrew White12:46

... it was, you know, I, I had been on sabbatical for a very long period of time, and so at a certain point, I just had to resign my tenure. So I resigned tenure in June.

RJ Honicky12:54

Okay.

Andrew White12:55

Um-

RJ Honicky12:55

Oh, so that's only recently.

Andrew White12:56

Yeah, only recently.

RJ Honicky12:56

Yeah, yeah. And you just felt like this is, this is the direction of my career. I, I need to do this.

Andrew White13:00

Yeah, yeah. I mean, I got tenure, and I had these, these early career awards, like the NSF Career Award. It was great, and I think academia is really exciting, but, um, I just thought that right now, these, this kind of area, like AR science is, um, just I think, A, difficult to do in academia, and B, so exciting that I think you can take bigger bets.

RJ Honicky13:22

Yeah.

Andrew White13:23

And, um, I think having a tenured position and writing research grants is maybe not the biggest bet you can take on, on a field.

RJ Honicky13:30

Yeah.

Andrew White13:31

So now we have a venture-backed startup called Edison-

RJ Honicky13:34

Right

Andrew White13:34

... which we spun out of Future House.

RJ Honicky13:36

Right.

Andrew White13:36

And, you know, we took a lot of the ideas, and we're trying to do this at an even bigger scale right now.

RJ Honicky13:40

Yeah.

Brandon Anderson13:41

And, and so Edison was always kind of the plan. Like, going back to Sam's idea of a FRO or a, like, what's it, fundamental research-

RJ Honicky13:48

Yeah

Brandon Anderson13:48

... organization. Like, he, he always had this goal of, like, let's do fundamental research in this, like, tightly scoped nonprofit, which can kind of explore, and then you have that as a natural arm for spinning off-

Andrew White14:00

Yeah

Brandon Anderson14:00

... um, you know, venture-backed, uh-

Andrew White14:03

Yeah, I think that's right. Um, I think some things that, like, make that not as clean these days is how expensive AI research is and how expensive GPUs are.

Brandon Anderson14:12

Mm-hmm.

Andrew White14:12

So I don't think we can repeat it many times from Future House. It might be, like, an N of one thing right now. It just, may- maybe not. I don't know, if, if-

RJ Honicky14:20

Yeah

Andrew White14:20

... venture capital keeps growing, then maybe, I mean-

RJ Honicky14:22

Yeah

Andrew White14:22

... maybe we can. But yeah, I think we took a lot of the ideas of Future House. Another thing, like, I, I think we expected it to be harder to automate science, and actually, it's really hard. Like, I'm, I feel like I'm always miscalibrated in this domain, but it's always hard to predict progress.

RJ Honicky14:37

Yeah.

Andrew White14:37

Um, and I think that I overestimate the speed of things on month scale, and I underestimate the things on year scale. So the two years from, like, 2023 to 2025 was an enormous amount of progress.

RJ Honicky14:49

Yeah.

Andrew White14:50

Um, it always felt like things were not going as fast as I thought, but when you look back on it, wow, like, there's a lot of progress. And so I think the idea that, I think at Future House and our, Sam actually regrets us writing this, but in the, the original marketing or, like, the announcement was like, "It's our 10-year mission to automate science."

And then, like, now it's like, okay, yeah.

RJ Honicky15:07

It's like two year, yeah.

Andrew White15:07

Yeah. So two years later, we, we had Kosmos, and, um, things are going so much faster, and it, also this kind of, this, this thing you, you notice it in, um, in, in San Francisco is, like, it's actually kind of hard to find problems which are, like, so hard that, like, they are a challenge for language models, but not so hard that they're impossible.

And we're in this, like, gray zone, and actually, I feel like that's where we are right now, is that we can actually automate so much of the scientific method because it turns out, especially in a field like biology, which is very empirical limited, you know, the top 1% guesser of, you know, what, what they think will happen in an experiment, the, you know, the top, you know, quintile or quartile, they're about equal.

And so, you know, even if you had an even, we wait 10 years and get even smarter models, I don't think it's gonna really change the fact that we're ready to automate a lot of science with, with existing LLMs.

Automating Science15:54

Brandon Anderson15:54

I mean, what do you mean by automate science? That's, like, a pretty loaded statement. There's lots of things.

Andrew White15:59

Yeah.

Brandon Anderson15:59

Like, there's-

Andrew White16:00

Absolutely

Brandon Anderson16:00

... many ways of thinking about that, so-

Andrew White16:01

So we try to draw a line between, um, what I would call, uh, uh, groups that are trying to model something like the cell or how proteins fold or, like, how antibodies can be designed or, like, maybe the virtual cells, like an example.

If they're trying to use machine learning or AI to model some very specific system-

Brandon Anderson16:17

Mm-hmm

Andrew White16:18

... we're trying to automate the, like, cognitive process of scientific discovery, making hypotheses, choosing experiments to do, analyzing the results from the experiments, and using it to update your hypotheses or your confidence in those hypotheses, and then leading to, like, a world model which of, like, okay, this is how I understand this process to be, and then that, you know, begets new hypotheses or new experiments.

We want to automate that sort of loop. We thought that we would have to build up, like, a whole new organization from the ground up for agents. So it means, like, automated labs. It means, like, putting all the papers in one spot, like getting APIs wrapped around everything.

But, um, over time, like, the model's gotten better and better that we had to, like, you know, stop and rethink. Okay, we don't actually have to hold their hands so much anymore.

Brandon Anderson16:58

Yeah.

Andrew White16:58

Or, like, they don't actually need to necessarily have an automated lab. They can, like, write an email to a CRO or something, or they can, like, tell you what experiment to do, and you can take a video of you doing it and then show it to the model-

RJ Honicky17:08

Yeah

Andrew White17:08

... and it can be like, "Okay, well, this is, you know, what happened."

Brandon Anderson17:10

Mm-hmm.

Andrew White17:11

So it's been a really interesting experience of, like, sometimes, you know, we over-engineer things, and sometimes-

RJ Honicky17:17

Yeah

Andrew White17:17

... uh, uh, well, actually, basically, just mostly over-engineer.

RJ Honicky17:21

So, so I always think about systems, and scientific is a system, like scientific process is a system. I always think about in terms of constraints, right?

Andrew White17:29

Mm-hmm.

RJ Honicky17:29

And, like, what is a bottleneck in the system? So that, so what is your hypothesis about this, right? Like, in my mind, not knowing a ton, but in my mind, the constraint of the scientific process is the work you do in the lab, and that's sort of notably missing from...

Well, not entirely. You said auto- you mentioned automating lab and whatever. So, like, how are you thinking about this?

Andrew White17:50

Yeah, I, I think you're right, is that, uh, basically the, the best model, um, whatever, Opus 7 or GPT-10, like, um-

RJ Honicky17:58

Yeah

Andrew White17:59

... it really can only propose the first experiment, maybe slightly more clever, but at a certain point, you just need information, right? Like, some little calculations you can do that, like, there's more atoms in the brain than you could ever simulate, even if you had all the energy from the sun, right?

Like, I think you could simulate maybe 1,000 brains-

RJ Honicky18:16

Yeah

Andrew White18:16

... in real time with all the energy in the sun 'cause there's too much information.

RJ Honicky18:19

Yeah.

Andrew White18:19

So science really hits these, like, bottlenecks where you just actually have to go measure things.

RJ Honicky18:22

Yeah.

Andrew White18:23

We definitely think about maybe like lab in the loop sort of situations. Like, one of our papers, which was called Robin, is that we, like, had an, uh, one of our agents propose an experiment. We did the experiment, and then we had our agent analyze the experiment, then propose the next experiment.

Brandon Anderson18:36

Yeah.

Andrew White18:36

And that kind of loop, I think, is where you wanna get to.

Brandon Anderson18:39

Yeah.

Andrew White18:39

So what is the bottleneck in that? I don't think it's, like, the intelligence of the first experiment. I think the bottleneck might be something... Like right now, I think the bottleneck is something silly like knowing what's the lead time on all the reagents that you need and what is available in the lab, right?

Like, you know, I-

Brandon Anderson18:54

Yeah, yeah, yeah.

Andrew White18:55

I think whether GPT-5.2 Codex Max or Opus 4.5 is gonna do better is probably doesn't matter.

Brandon Anderson19:02

Yeah.

Andrew White19:02

It's just a matter of, like, which one's gonna have all the information of what's in the lab and how much will it cost, how long will it take.

Brandon Anderson19:07

Right.

Andrew White19:07

Um, and also, I guess the, the kind of frontier that I think about for these models is taste, which is, like, a lot of science, and of course we want to, you know, accelerate technology, we wanna improve the economy, we wanna improve people's life expectancies, we want everyone to be happier.

But a lot of what is done in science is, is based around, like, human preferences. Like, why do people study, I don't know, a particular worm? Well, like, there is a theory that by studying the worm it has led to good medicines or it's led to discovering new genes.

But also people studied it in the past, people's careers depend on that worm- ... and people wanna write papers about that worm.

Brandon Anderson19:42

Yeah.

Andrew White19:42

And, and so there's a human element to some of this.

Brandon Anderson19:44

Yeah.

Andrew White19:44

And I think that, that, uh, models don't capture that so well about knowing what is an exciting result and what is a boring result.

Brandon Anderson19:50

I see.

Andrew White19:51

So I think that's like a, a scientific taste. It's like a broad category of, of all these things.

Brandon Anderson19:55

How do you def- Like, do you try to quantify taste in any way? I mean, I know that I, I have some, like, fun anecdotes about this, but maybe, yeah, I'd just like to hear what you think.

Andrew White20:05

Um, yeah, actually, we sat on this idea, we sat on it, but we, like, argued about it for a long time. Sam and I usually every Monday morning at 8:00 in the morning, Sam and I meet, and we're both, you know, caffeinated and ready, and we argue about stuff like this.

Brandon Anderson20:20

Yeah.

Andrew White20:20

And we had a lot of Mondays where we talked about scientific taste. And in the end we're like, "Okay, let's just, let's just do the dumbest thing," which is to, like, have our agents make hypotheses and put them in front of humans and have them be like, "I like this one," or, "I like that one," right?

So we just did, like, whatever, RLHF on, on, on hypotheses. And, um, we learned a lot about how bad RLHF is with people. Just, like, people pay really attention to, to the tone, to the details, to, like, how many specific facts or figures are on the hypothesis, right?

Like actionability about, like, if the experiment is feasible. But what people didn't really pay attention to is, like, I don't know how to describe this, like, if this hypothesis is true, how does it change the world? If the hypothesis is false, how does it change the world?

This like, um, how much information you gain. It's not really information, but, like-

Brandon Anderson21:04

Yeah

Andrew White21:04

... impact or something.

Brandon Anderson21:05

Impact, yeah.

Andrew White21:06

And that really didn't come through from those things. So then we're like, "Okay, well, this is maybe one strategy," and so we had to go back and think about it more. And, um, we then took a pause from that research, and then we made Kosmos.

And then Kosmos, like, has baked into it taste, right? Like, at the end of the day, there will be some report, and we're working on generalizing this. Basically, at the end of the day, we'll be like, "Okay, I made these discoveries."

Then a person will be like, "Great, I'm gonna download that one," or, "I like that one," right? "I don't like this one." And that rolls up to some hypothesis that came earlier in the process.

Brandon Anderson21:34

Mm-hmm.

Andrew White21:34

And so w- we think we can get to end-to-end on this-

Brandon Anderson21:37

Yeah

Andrew White21:37

... as opposed to human preferences.

Brandon Anderson21:39

So you, you mean the feedback loop is the click?

Andrew White21:42

It could be the click. It could be also, like, the, you know, it... We do an experiment. Sometimes in Kosmos you could ask it to end an experiment, and you can see was the experiment a success or failure-

Brandon Anderson21:51

Yeah

Andrew White21:51

... or something like. But I, I guess, like, we, we brought it out of this kind of like hard to quantify, is this a good hypothesis or a bad hypothesis, and into this, like, you can see some downstream consequences of the hypothesis.

Brandon Anderson22:02

So yeah, humans have, I think, a very, like, strongly well-calibrated nose for science. Like, I mean, maybe you could argue there are sociological effects-

Andrew White22:12

Yeah

Brandon Anderson22:12

... like, um, across the community. But ultimately, oftentimes people really, like, good scientists know right off the bat, like, is this going to be likely to be, um, useful or not? How long, how many attempts did it take before you started to, like, see results that to yourself seemed useful?

Like, you've been working on this for, I guess, two years now. Um-

Andrew White22:35

You know, I think when the Co-Scientist paper came out from Google, I think it was a really interesting idea to do this, like, tournament style or just pairwise ranking of hypotheses, right? So I think Co-Scientist is a very interesting counterexample to what we built.

What we built is something with either lab in the loop or data analysis in the loop or literature research in the loop, where you're, like, iterating on an idea. And then Co-Scientist took this, like, very different approach of, like, let's list all the ideas and then try to come up with a filtration process to come up with the best hypotheses.

Brandon Anderson23:05

Mm.

Andrew White23:05

So Co-Scientist will produce these very long reports of, like, "Oh, we, like, really tested this idea with lots of dialogue," and, and, and it's very interesting stuff. And I was really impressed with the paper that came out. And then we had this Robin paper, and one of the things that came out of the Robin paper is that the hypothesis that people thought was best was not the one that led to success in that paper.

Brandon Anderson23:25

Interesting.

Andrew White23:26

It was in, um, age-related macular degeneration or ocular, uh, age-related macular degeneration. So basically it's like part of the eyes. You're going blind because you have this, like, accumulation of debris in the eye and can't clear it out.

That's one of the major cause of blindness in people over-

Brandon Anderson23:39

Yeah

Andrew White23:40

... uh, 60. Uh, Oli, who works on the hill-

Brandon Anderson23:43

Yeah, yeah.

Andrew White23:44

He'll, he'll cringe when he hears me say that, but-

Brandon Anderson23:45

Something, something like that.

Andrew White23:46

Something like that.

Brandon Anderson23:47

Yeah.

Andrew White23:47

Sorry, Oli. Um, in that one, like we, we went to optometrists or ophthalmologists. Actually get confused on that as well. Sorry, Oli. Um- But essentially, you know, asked them, "What hypotheses do you think are, are, are good hypotheses?

What do you think would lead to, like, um, a good mechanism for treating dry AMD?"

Brandon Anderson24:03

Yeah.

Andrew White24:04

And, um, yeah, there was, you know, they agreed maybe on the top 10.

Brandon Anderson24:07

Yeah.

Andrew White24:07

But, but beyond that, it was, it was kind of noise.

Brandon Anderson24:10

Yeah.

Andrew White24:10

Um, and then, you know, the, the... what we found was ribose dutil was a very good, um, medicine and, and had a, a mechanism that I think is novel, although there was lots of debate on X.

Brandon Anderson24:20

Yeah.

Andrew White24:20

Because in, I think in 2012, there was a master's thesis which proposed this mechanism on like page 38. I actually think it was a typo. I think they meant wet, wet AMD. But anyway, I won't, I won't belabor the point.

I will concede that, that maybe was... there was one reported, uh, example of it in the past. That was a, a really eye-opening experience for me because that was a, the first really serious test where we really went to the lab, and we, you know, spent, like, four weeks on a battery of experiments to see-

RJ Honicky24:48

Yeah

Andrew White24:48

... what is, uh, which hypothesis led to a good mechanism and a good repurposed drug.

RJ Honicky24:53

Right.

Andrew White24:53

And it was not as correlated with human opinions as I expected.

RJ Honicky24:57

Opinions. Yeah, yeah.

Andrew White24:58

And so since then, I think that, uh, I have a lot more faith in these, like, um, verifier-in-the-loop kind of scenarios where you have either data analysis, literature search, or you're running a unit test or whatever. You're, you're, you're going and running the experiment.

Anything like that, I think, is, is gonna give you a higher signal than the sort of vagaries of, like, "Oh, this is a higher opinion," or "That we like this one better."

RJ Honicky25:20

Yeah, Max, Maxwell called it nature's computer.

Andrew White25:22

Yeah.

RJ Honicky25:23

It's like-

Andrew White25:24

There you go. Yeah

RJ Honicky25:24

... it's like you have this computer computational cycle you're running, and the nature is part of that computation cycle.

Andrew White25:29

Yeah, yeah.

Brandon Anderson25:30

I'm curious, so you said that there is a paper which maybe, like, could propose, maybe propose where this, um, molecule came from, but, like, do you have some way of interpreting or, like, understanding where that hypothesis originated in the absence of that?

Like, is there-

Andrew White25:46

Um, oh, from the model?

Brandon Anderson25:46

... a traceable thought train?

Andrew White25:47

Yeah, yeah, yeah. Actually, this is something we pay really close attention to at, at Future House and at Edison, was, um, provenance of, like, information. So our first sort of, like, agent was, was PaperQA. Sorry about the name.

PaperQA sounds like an EL7 that was an agent.

RJ Honicky26:01

It, it really does.

Andrew White26:02

Yeah, yeah. Um, PaperQA was, uh, has, like, every sentence that it, it outputs has a citation to a page, right? So it's like a lot of provenance. And then we basically built along a philosophy for everything. So Robin would, which is the name of this, I don't know, workflow or something you can call it, that led to this result on repes Gl being a good, um, therapeutic for dry AMD.

RJ Honicky26:23

Mm-hmm.

Andrew White26:24

It has, like, data analysis that goes, shows you, like, which line Python code-

RJ Honicky26:29

Mm-hmm

Andrew White26:29

... led to the result here, and then that is like, okay, then it goes to this other model, which says, "Well, based on this literature finding and this result from the data analysis, I believe this is the right thing."

But, you know, where does the original idea come from? Like, going after these ROCK inhibitors, which is the, the mechanism for the, the target, was basically enumeration. And so this is like, if you can't be smarter, you can be, I don't know, you can try more times.

RJ Honicky26:52

Brute force.

Andrew White26:53

Brute force.

RJ Honicky26:53

Yeah, yeah.

Andrew White26:53

And I think that was, like, the theory of, of the Robin paper was that we can put out a whole bunch of hypotheses-

RJ Honicky26:57

Yeah

Andrew White26:57

... and then we can filter them, just like I think is about how CosScientist did, is you go for a filtration process. But the difference is that in CosScientist, their filtration process was other LLMs sort of ranking it with rubrics or, like, personas, and our filtration process was, like, literature search and data analysis.

Um, like, here's some data. Is it consistent with the data? Go see if anyone's discovered it in the paper, in the literature, or if they've disproven it.

RJ Honicky27:19

Yeah.

Andrew White27:19

And I think that's the easy way to succeed in, in AI over humans is, is you can try more ideas faster.

Brandon Anderson27:26

Something I've, I've heard people say, and maybe I've experienced this in my own life, um, like, sometimes hypotheses are kind of cheap, especially in, you know, biology.

Andrew White27:35

Yeah.

Brandon Anderson27:36

It's many ways actually easy to come up with what you think could be happening.

Andrew White27:39

Yeah.

Brandon Anderson27:39

And, um, it seems like to me verifying is oftentimes a big bottleneck and maybe the biggest bottleneck. Like, if you have lots of hypotheses, you know, and it costs, you know, one-one hundredth of your runway to test each one of them or something, you don't have many shots on goal.

Andrew White27:54

Yeah.

Brandon Anderson27:55

Yeah. So how do you make sure that, like, you, you are actually enriching for good hypotheses?

Andrew White28:01

Literature and data analysis, right? You know, like-

Brandon Anderson28:03

You just, you just do.

RJ Honicky28:04

Yeah.

Brandon Anderson28:04

Yeah.

Andrew White28:05

There was a time when we used, um, something called tiling trees, and tiling trees is like a literal brute force method invented by Ed Boyden, Sam's PhD advisor, and basically the idea is, okay, I wanna, like, accomplish X.

Okay, I could try these methods. And then, like, once you pick, "I'm gonna try this method," then you split into, like, two different paths. I'm gonna use this method or not use this method.

Brandon Anderson28:26

Yeah.

Andrew White28:26

If I'm using this method-

Brandon Anderson28:27

Mm-hmm

Andrew White28:27

... I need to have, like, I don't know, some kind of substrate. I'm gonna try this substrate or this substrate or this substrate, right? And you can basically try to really, like, tile space of all the ideas. We tried some early experiments there, and you're right.

You run into this thing where some of the hypotheses that come out just don't make any sense, and, like, you are going to waste a ton of effort if you actually test them all. Nowadays, I actually would argue that if you go to an LLM, and you ask it to evaluate, you know, hypotheses, including some garbage ones, it will probably do as good of a job as an expert in the field in, in filtering them out.

That's not always the case.

Brandon Anderson28:57

Yeah, I've actually seen that myself.

Andrew White28:58

Yeah, but there's a lot of gotchas-

Brandon Anderson29:00

Yeah

Andrew White29:00

... and I think people can miss those.

Brandon Anderson29:01

Yeah.

Andrew White29:01

But I think they're actually pretty good. And so I'm not as worried about hypotheses that can fail fast by an expert looking at them.

Brandon Anderson29:08

Mm-hmm.

Andrew White29:08

I think now the filtration process really happens in, in, in literature, and I think the filtration process happens in looking at, like, um, you know, like, uh, a biobank data or, like, you know, what do we know from GWAS or something.

Brandon Anderson29:20

Yeah.

Andrew White29:20

You know, you know, other sources of, of existing data as much as you can draw upon.

Brandon Anderson29:24

Yeah. So yeah, and with regards to existing data, um, another, like, maybe contrarian, uh, take is that, like, oftentimes the, the hardest part is just understanding the, like, the context of data and, like, where it comes from and how do you interpret it.

Like, I can also think from my own life multiple cases where, you know, the, the data in some sense was, like, there, and you had two people who were both experts and very smart people who looked at it and drew very different interpretations.

And in fact, like, when we were interviewing Heather Culic, she, she had some fun stories about using LLMs and, um, she would find that there would be raw data in a paper which wouldn't agree with the conclusions of the actual paper, and it's straight from the paper.

It's not even, like, cross-paper talk or something.

Andrew White30:09

Man, I'm gonna be a really boring interviewer and be like, "Yes, you're right." You know, like, this is, this is a hard question.

Brandon Anderson30:16

These things are hard. Yeah.

RJ Honicky30:16

Yeah.

Andrew White30:16

Um, I think, you know, to put, give you something concrete, um, we have a, a, a bioinformatics benchmark we call BigBench.

Brandon Anderson30:24

Mm-hmm.

Andrew White30:24

BigBench is like we put it out. We've updated a few times. It's in, it's in some frontier, um, uh, LLMs when they release their system card, they'll mention BigBench as, like, one of the things they test on.

Brandon Anderson30:33

Yeah.

Andrew White30:33

And, um, you know, we're getting to 60%, 70% correctness on BigBench.

Brandon Anderson30:39

Mm-hmm.

Andrew White30:40

And, um, we found that actually we're at the point where humans disagree at this level. Like, humans only agree 70% of the analysis.

RJ Honicky30:50

Yeah.

Andrew White30:50

And so it's true that, like, when it comes to analyzing data, like humans do not agree 100% of the time. There is a certain amount of, like, choice that goes into it. And, you know, we, we try to, um...

So Edison is a, is a for-profit company. We like trying to sell, uh, some of this stuff to, to companies, and we'll go to some companies that like, "Oh, well, we never impute data. Imputing data is bad." Like, or, you know, whatever.

And then like, okay, well, we'll have to change our agents so we don't impute data for them. And then some of the companies are like, "Oh, yeah, we impute data. It makes everything easier," right? Or, uh... And, and, and, you know, you wonder what the real modern dark arts are, the, like, AI-resistant area of the world is, like, uh, medicinal chemistry.

That is, like, the spot where, like, you know- ... there's so much superstition.

RJ Honicky31:27

Oh, yeah, yeah.

Andrew White31:28

Yeah.

RJ Honicky31:28

Everyone, yeah.

Andrew White31:29

It is, it's-

RJ Honicky31:29

Everyone is, like, pseudo-religious.

Andrew White31:31

Yeah, exactly. You have to be to survive.

RJ Honicky31:32

Yeah.

Andrew White31:33

And I feel otherwise you get burnt out.

RJ Honicky31:34

But the religions never agree, too.

Andrew White31:36

They... Yeah, exactly, yeah.

RJ Honicky31:36

Two medicinal chemists will have completely different viewpoints about, like, a functional group.

Andrew White31:40

Yes, exactly.

RJ Honicky31:41

Yeah.

Andrew White31:41

And I remember this is, uh... Talking to somebody who works with a CRO, and they're like, "Oh, whenever, like, company X orders anything, we never put boron on any of the compounds- ... because they hate boron," because there was one program-

RJ Honicky31:50

Yeah

Andrew White31:50

... that was killed because there was a boron, you know, somewhere in the core, and it led to some s- toxic side effects. So no boron for this company. This company, they, like, love things to be fluorinated or something-

RJ Honicky31:59

Yeah

Andrew White31:59

... because they love... think it's great for, um, the admin properties, right? And so there's, like, all this, this, this stuff where you reach the point where you're at, I don't know, human bias level or human disagreement level, and I think we're getting to that point in data analysis.

And so, of course, you will see then that if I take the raw data from a paper and I analyze it myself, I will get a different conclusion. One of the cool tricks you can do is... This is back to this brute force thing, is that I can go to our agent and I can run it 100 times, and I can take the consensus, like, analysis.

Or I can say, even if you make these three different choices in your data analysis, you get the same conclusion, right? Or this conclusion is somehow s- sensitive to those choices. And then you can... There's even, like, words.

It's like epistemic versus aleatoric-

RJ Honicky32:38

Oh, yeah, yeah

Andrew White32:39

... uncertainty, right? It's like this is aleatoric, which means, like, I think it's noise from the data, or this is epistemic uncertainty, which means, like, I think there's some choices that are being made. There's some model differences that lead to the, the disagreement.

Anyway, there's, like, a, there's, like, a Donald Rumsfeld formulation of this as well. Like- ... the known unknowns-

RJ Honicky32:57

Yes, yeah

Andrew White32:57

... and yeah, it is, it's aleatoric, epistemic, uh, uh, uh, debate there.

RJ Honicky33:01

Interesting. This kind of digging into your Kosmos model-

Andrew White33:05

Yeah, yeah

Inside Kosmos33:05

RJ Honicky33:05

... a little bit. So I, I glanced at the paper, and one of the things that jumps out is that there were a certain class of problems for which, uh, it was only 50-some percent accurate. And-

Andrew White33:17

Oh, yeah, yeah.

RJ Honicky33:17

And can you talk a, a little bit about that and how that... Like, okay, so if I'm just raw getting 50% accurate answers, then, and then I'm going to the wet lab and being like, "Okay, try this," and then it's like, "Ah."

Andrew White33:28

Yeah.

RJ Honicky33:28

Like, "The, the stupid thing did- told me to do a dumb thing."

Andrew White33:30

Well-

RJ Honicky33:31

How do you-

Andrew White33:32

I would say, first of all-

RJ Honicky33:33

Yeah, yeah

Andrew White33:33

... that 50%, it's actually pretty good because it's rare that experiments in the lab are actually coin tosses, right? There are usually a lot more outcomes than, than, you know, than, than binary. Um-

RJ Honicky33:42

Yeah, yeah. Sure. Okay, yeah.

Andrew White33:43

But, but that particular number was, uh, human agreement in the interpretation of the results.

RJ Honicky33:49

Okay.

Andrew White33:49

And so, uh, we asked people to evaluate different aspects of Kosmos. We had them evaluate, like, the data analysis decisions. We had people ask to evaluate the literature. Like, is this, are these, is... Do you agree with its finding in the literature?

That number that was 50%, that came from Kosmos' interpretation of, uh, some of the analysis.

RJ Honicky34:06

Yeah.

Andrew White34:07

So, like, it might go in literature and find this result, and it would say, "Wow, this is super exciting. This is amazing."

RJ Honicky34:13

Yeah.

Andrew White34:13

Or it might do data analysis, be like, "This is a novel discovery. We're really excited about it."

RJ Honicky34:16

Yeah.

Andrew White34:16

And then people would disagree. "That's actually not interesting," or like-

RJ Honicky34:19

Yeah

Andrew White34:19

... "I don't agree with the interpretation of it."

RJ Honicky34:21

So it's like picking bad problems maybe.

Andrew White34:23

Yeah.

RJ Honicky34:23

In the, the, the negative-

Andrew White34:25

Yeah

RJ Honicky34:25

... class. Yeah.

Andrew White34:25

And so I think it's like that, that 52 or 55, whatever it is, that's, um, interpretation. And so I agree. I think that's where, like I was saying, I think the frontier right now-

RJ Honicky34:33

Yeah, this kind of taste

Andrew White34:33

... is scientific taste.

RJ Honicky34:34

Yeah, yeah. Taste.

Andrew White34:35

Yeah. And so that's what we're working on right now, is-

RJ Honicky34:37

I see. Yeah, yeah

Andrew White34:37

... how do you get the interpretation to, to match? Um.

RJ Honicky34:40

Interesting.

Brandon Anderson34:40

Could you just step back and just introduce Kosmos-

Andrew White34:42

Yeah, yeah

Brandon Anderson34:42

... from a high level? Yeah, yeah. Um, it's, it's-

RJ Honicky34:44

It-- I would actually be in-

Brandon Anderson34:46

Yeah

RJ Honicky34:46

... even curious to hear starting from, like, ChemCrow and, uh, you know, you have, uh, PaperQA, Avery, Ether 零 .

Andrew White34:54

Zero. Yeah, yeah.

RJ Honicky34:55

I'd like to hear a little bit of the, the lineage and how those different decisions were made. What were the key learnings, and how did you get to where-

Andrew White35:03

Yeah

RJ Honicky35:03

... you are now?

Andrew White35:04

Yeah.

RJ Honicky35:04

Yeah.

Andrew White35:04

So I could retcon and tell a really great story about how we arrived at Kosmos. But I will say that, like, to a large extent, we just try a lot of stuff, and sometimes-

RJ Honicky35:11

Yeah

Andrew White35:11

... it works, and sometimes it doesn't.

RJ Honicky35:13

Yeah. Okay.

Andrew White35:13

You know, I'll, I'll, I'll say that we're very... I'm, I'm a, I'm a builder. Like, I like to, like, build things piece by piece. I'm, uh, probably some fancy word for it, but I'm like a Lego guy or something.

RJ Honicky35:24

Yeah.

Andrew White35:24

My vision was that we would make an agent that does this part of the scientific process, an agent that does this part of the scientific process, whatever. And so we had, like, um, ChemCrow, which is gonna help us with setting up our medicinal chemistry work.

We had ProteinCrow, which we haven't released. I don't know if we will ever release, but ProteinCrow was, like, bu- designing proteins we might need for, for some part of our workflows. Um, or we had a data analysis agent.

RJ Honicky35:45

Is that LLM, or that's a-

Andrew White35:46

It's an agent, so an LLM plus tools.

RJ Honicky35:47

Okay.

Andrew White35:48

Or we had Ether 零 . It was like, okay, we noticed that the frontier models can't work with molecules very well, so let's make a, a model with intuition for medicinal chemistry, and that was what led to Ether 零 .

But then, uh, Sam actually really pushed on us to like, "Let's just even do the whole thing. You know, let's, let's just try to build an AI scientist. Let's just try the whole thing." And, and that was what led to Robin.

And, um, Robin was like, "Let's just take these agents we already have, and we'll just put them in, like, a, a, a workf- basically it's like you could express it in a concise Python file of like, you know, try a whole bunch of ideas, then go see if they all filter through literature or if they've been disproven, and then go, like, uh, come up with experiments that you could do in a wet lab."

RJ Honicky36:24

Yeah.

Andrew White36:25

"And this is our inventory list. And then go analyze all the data, then go back and repeat the process." Right? So that's, like, what Robin was. And, um, and we, like, we came acro- a- across Kosmos. We were trying to, like, understand what is the process that Robin is, is automating?

And it came from this idea of, like, a world model, which is that when we first started Edison, we were thinking, like, what do we, what do we want to change about this? Like, what is new here?

Brandon Anderson36:46

Yeah.

Andrew White36:46

And so we spent some time thinking about, well, the scientific process, like, what is actually going on in, like, my brain, which is that I have some understanding of-

Brandon Anderson36:53

Yeah

Andrew White36:53

... of the world or the-

Brandon Anderson36:54

Yeah

Andrew White36:54

... phenomena I'm studying, and that's my world model. And then a lot of the actions I take are about trying to update that world model. And it's something that changes over time, and so this is, like, this ability to change over time, but it's also something that is practical.

Like, I can use it to make predictions about, I know from this experiment, this will happen. That's why it's like a model and not just, like, you know, a, a memory or like-

Brandon Anderson37:14

Yeah

Andrew White37:14

... a bunch of, like, papers or something like that. It's, like, uh, it's supposed to operate. In Kosmos, we tried this idea out, and actually, um, uh, uh, Ludo, uh, who's the, the, uh, first author on the paper, we tried a whole bunch of ideas around world models.

And, uh, we kind of thought they weren't really appropriate. Like while we tried a lot of different ways to do this, we tried, you know, method A, method B, method C, and, and they were okay. And so we all just had to take a break.

Ludo, like, his project didn't work on trying to do this world model stuff, but he's like, "I'm gonna keep trying it." Ludo's a very stubborn person. So he tried it-

Brandon Anderson37:45

Yeah

Andrew White37:45

... for, like, I don't know, a week or two weeks, then he-

Brandon Anderson37:47

Yeah

Andrew White37:48

... just kind of, like, quietly was like, "Hey, can you guys t- come take a look at this?"

Brandon Anderson37:51

Yeah.

Andrew White37:51

And, and we're like, "Wow, this is actually really cool," and then we, like, started building on it a- and jamming, really. And I think what Ludo figured out is that you have to get this, like, experiment loop thing.

You have to be a- let it... Then the data analysis agent is what got us in the loop.

Brandon Anderson38:03

Yeah.

Andrew White38:04

So if you put that in the loop of, like, it can really update this world model because the... We were trying to build it around literature before. And when you build it around literature, there's just, like, not really experiments you can do and then see the results for.

That was, like, our surrogate was, was literature.

Brandon Anderson38:15

Yeah.

Andrew White38:15

It just wasn't working.

Brandon Anderson38:16

Right.

Andrew White38:16

But data analysis actually really lets you explore ideas. And so that was what led to Kosmos, and so in Kosmos, we basically, we had all the pieces sitting around. We've been working on world models, we've been working on a data analysis agent, working on a literature agent, um, and then we're working on, you know, we built a platform for scientific agents.

So we had things that can write a LaTeX report. We had things that can make nice plots.

Brandon Anderson38:35

Yeah.

Andrew White38:35

Then we put that all together, and, like, a world model was, like, sort of the, the glue that allowed it to, to fit together.

Brandon Anderson38:41

Yeah.

Andrew White38:41

An analogy is, like, um, in coding agents, like, GitHub is sort of the glue. Like, there's some shared repo, and everyone works on the repo, and, like, software engineers have spent whatever, lots of brain cycles thinking about what's the way to coordinate, you know, and organize working on code together-

Brandon Anderson38:56

Yeah

Andrew White38:56

... for a long time.

Brandon Anderson38:57

So the world model is actually like a memory system kind of sort of?

Andrew White38:59

Yeah, you can think of it as a memory system.

Brandon Anderson39:01

Oh.

Andrew White39:01

Um, we, we think about it as, as a model. So, like, it actually-

Brandon Anderson39:04

Yeah

Andrew White39:04

... you can put in-

Brandon Anderson39:04

Yeah, sure

Andrew White39:05

... input, and it will output predictions.

Brandon Anderson39:06

I see. Okay.

Andrew White39:07

And we, we think about calibration.

Brandon Anderson39:08

Yeah, yeah.

Andrew White39:08

But, like, really, it is a set of, like, a big bundle of information-

Brandon Anderson39:13

Right

Andrew White39:13

... that we accumulate over time-

Brandon Anderson39:14

Yeah, yeah

Andrew White39:14

... that's distilled in some way, and, and that is, like, uh, what allows us to do this. And I think you can think about, like, um, a GitHub repo as, like, it's a distillation, right? Like, really, there's a long graph of commits that lead up to it-

Brandon Anderson39:25

Yeah

Andrew White39:25

... and, like, the current file system in that GitHub repo-

Brandon Anderson39:28

Yeah

Andrew White39:28

... or, I keep saying GitHub. I'm-

Brandon Anderson39:29

Yeah

Andrew White39:29

... such a corporate- ... uh, shill here. Git-

Brandon Anderson39:32

Yeah

Andrew White39:32

... your Git repo is like a, a distillation of all of the work that people put in into the PRs, into the, into the, uh, the commits. And so I think there's a nice analogy, uh, between a Git repo and, and what a world model is.

Brandon Anderson39:43

I see. Yeah.

Andrew White39:44

And I think that's just sort of what allows us to automate scientific discovery so well.

Brandon Anderson39:49

Can you talk about, like, kind of how you implement a, a world model, or is that sort of like secret sauce?

Andrew White39:53

That's our, like, secret sauce right now.

Brandon Anderson39:56

Okay.

Andrew White39:56

You know?

Brandon Anderson39:56

That's, that's fine.

Andrew White39:56

Yeah. No, it's fine. People have asked for less.

Brandon Anderson39:59

That question never happened, okay?

RJ Honicky39:59

So one thing that's notably missing-

MD Overrated40:01

Andrew White40:01

Yeah

RJ Honicky40:01

... is the, like, simulation, right?

Andrew White40:04

Yeah.

RJ Honicky40:05

Molecular dynamics or-

Andrew White40:06

Yeah

RJ Honicky40:06

... or, or, like, uh, bolts or...

Andrew White40:10

Yeah. I wanna, uh, help you guys pump up your views here. So I, I think molecular dynamics is overrated.

Brandon Anderson40:15

In fact, coming from someone-

RJ Honicky40:17

Yeah

Brandon Anderson40:17

... who's been-

Andrew White40:17

Yes, yes

Brandon Anderson40:18

... that's like-

RJ Honicky40:19

That goes in the, that goes in the intro

Brandon Anderson40:20

... doing it.

Andrew White40:20

In the thumbnail, you know? Yeah.

RJ Honicky40:22

But just-

Andrew White40:22

Yeah. And, and, and DFT is overrated. In fact, DFT may be even more overrated than molecular dynamics. I think these methods-

Brandon Anderson40:28

You mean for materials or for biology or for both?

Andrew White40:30

For materials.

Brandon Anderson40:31

Okay.

Andrew White40:32

And I can explain more about that. Basically, MD and DFT have consumed an enormous number of PhDs and scientific careers-

Brandon Anderson40:40

Yeah

Andrew White40:40

... at the altar of, you know, the beauty of the simulation.

Brandon Anderson40:43

Also, random interjection.

Andrew White40:45

Yeah.

Brandon Anderson40:45

Once I, I did an estimate, I think pre, like, ChatGPT, something like 20% of the world's computing power just went to simulating water.

Andrew White40:52

Oh, my fucking God, water.

Brandon Anderson40:54

Yeah, yeah.

Andrew White40:55

I had to deal with so many water simulations.

RJ Honicky40:57

Yeah, me and water.

Andrew White40:57

I did, I did DFT simulations of water, and they are so annoying. I used these big computers, uh, from, uh, the Department of Defense, and we s- I spent, like, I don't know, five months... And by the way, this is pre-LLM training days.

Brandon Anderson41:11

Yeah.

Andrew White41:11

Five months of compute is actually a really long time.

Brandon Anderson41:13

This is huge.

RJ Honicky41:13

Yeah.

Andrew White41:13

I simulated water with quantum, you know, effects, with the groatest mechanism for how a proton hops through water, and it's on YouTube. It's my number one YouTube video, and it represents, like-

Brandon Anderson41:24

Until now.

Andrew White41:25

And, yeah. And it represents, like, I don't know, a million CPU hours of compute. It was, you know, one of the biggest computes that I, uh, probably the biggest one I've done in my life so far. Maybe EtherZero is bigger, but, but it took a lot more work.

Anyway, and, and what's the point?

Brandon Anderson41:39

Yeah, what'd you learn?

Andrew White41:40

You know, what's the point?

Brandon Anderson41:40

What'd you learn from that?

Andrew White41:40

All I learned was, like, what set of hyperparameters reproduce some physical effects of water. But n- none of it was de novo, right? And this is the, this is the issue with, with molecular dynamics and DFT, is that, um, they don't model the world correctly, and so we have to invent little stories we tell ourselves about we're, like, making good inductive biases, and then it models the world more correctly.

Like, in DFT, you simulate water at 330 Kelvin-

Brandon Anderson42:04

Mm-hmm

Andrew White42:05

... when you want room temperature water. Is room temperature 330 Kelvin? No, it's not. That's a little too hot, right? And so this is a... The issue is that, like, people just make up these, these, these things, or, like, I don't know, GGA or, like, BEILIP or B3LYP, all these different, like, methods people invent.

They're clearly empirical, and then they bolt it on to DFT, and they say, "Look, it's a first principles method," right? But actually, you made a whole bunch of choices-

Brandon Anderson42:30

Yeah

Andrew White42:30

... and, you know, you, you, whatever, overfit to the validation data-

Brandon Anderson42:33

Mm-hmm

Andrew White42:33

... to get this to work. And, and that's, I think MD and DFT are like that because if you go look at the catalysts, you know, what catalysts change the world, none of them are single crystal materials that are really well suited for DFT.

They're always, like, they have grain boundaries. They have dopants. They're complicated, right? And, and you'll never capture that with DFT. So I think this is one of the- Fundamental, I don't know, dichotomies of the world is that simulations simulate really boring things really well.

They don't simulate interesting things very well. And so that's why I don't do DFT and MD anymore.

RJ Honicky43:07

What about somewhere like the machine learning-y stuff like AlphaFold and, and-

Andrew White43:13

Alpha was trained on X-ray crystallography data.

RJ Honicky43:15

Yeah.

Andrew White43:15

And I think, you know, this is the, this is the story of MD, is that MD was supposed to be the protein folding solution. There is a great counter example. There's a... I don't know what there's a word.

The counterfactual is basically a group called Desres, Dee Shaw Research. They had, you know, similar funding to DeepMind, um, probably more actually. They tested the hypothesis to death that MD could fold proteins.

RJ Honicky43:40

Yeah.

Andrew White43:40

They built their own silicon. They built their own clusters. They had them taped out all themselves. They burned into the silicon the algorithms to run MD. They ran MD at huge speeds, huge scales.

RJ Honicky43:52

Yeah.

Andrew White43:52

I remember David Shaw came to a conference once on MD, and he flew in by helicopter and was like- ...to this, this pretty famous guy-

RJ Honicky44:00

Wow

Andrew White44:00

...kinda rich.

RJ Honicky44:01

Yeah.

Andrew White44:01

And, um, he, he gave, uh, an amazing presentation about the special computers and special room and out- outside of Times Square, and like what they can do with it. It was-

RJ Honicky44:09

Yeah

Andrew White44:09

...beautiful, amazing. And I always thought that protein folding would be solved by them, but it would require a special machine. Maybe the government would buy like five of these things, and we could fold, you know, maybe one protein a day or two proteins a day.

RJ Honicky44:21

Yeah.

Andrew White44:21

And when AlphaFold came out and it's like you can do it in Google Colab, you know, or on a GPU or desktop- ...it was so mind-blowing. I forget, like, that protein folding was solved. I always thought that was inevitable.

But the fact that it was solved and on, like, your desktop, you can do it, was just completely floored, changed everything.

RJ Honicky44:36

This is like the, the bitter lesson on steroids.

Andrew White44:39

Yeah, I don't even know what it is. But it's like, imagine ChatGPT came out, but instead it was like, "Oh, you can just run it on your phone or locally on your own desktop." Like, that's the level of, like-

RJ Honicky44:47

Yeah

Andrew White44:47

...shock that came out.

RJ Honicky44:49

Yeah.

Andrew White44:49

And it, it gets down to this thing that humans are really bad at estimating problems that aren't human-made problems. Protein folding, we all thought was like, would require a huge amount of compute, very challenging problem, the most hardest problem in the world, right?

And it turns out that you can actually do it on, I don't, I think the numbers are now like ten thousand GPU hours, you can train a, a good protein folding model. It's actually turned out to be barely an inconvenience.

RJ Honicky45:07

Therefore, why not?

Andrew White45:09

Oh, oh, therefore, protein folding was-

RJ Honicky45:10

Yeah

Andrew White45:11

...highly efficient based on experimental data. It-- They took X-ray crystallography. They-- That's what DeepMind did, is they took X-ray crystallography data.

RJ Honicky45:17

Yeah.

Andrew White45:17

Desres tried the first principles method.

RJ Honicky45:18

Yeah, yeah.

Andrew White45:19

And it's like a nice head-to-head comparison-

RJ Honicky45:20

Yeah, yeah. Okay

Andrew White45:21

...two very well-resourced groups.

RJ Honicky45:22

Yeah, yeah. Yep, yeah.

Andrew White45:23

They both tried different ideas, and the machine learning on experimental data beat out first principle simulation by, you know-

RJ Honicky45:31

Yeah

Andrew White45:32

...a very large margin.

RJ Honicky45:33

And so why isn't, uh, like Bolts or whatever inside of Kosmos? Like, why isn't there a tool that's, that can run Bolts?

Andrew White45:39

Oh, we have Bolts inside of... We have Bolts Gen, Bolts Gen.

RJ Honicky45:42

Yeah.

Andrew White45:42

Yeah, yeah, we have that inside of Kosmos.

RJ Honicky45:43

Okay. Oh, okay, it is then. That's it.

Andrew White45:45

I mean, I think in the version that we have, uh, for people to just sign up and use, it's not in there.

RJ Honicky45:50

Yeah.

Andrew White45:50

But, like, uh, you know, you can imagine that you can just Modal or Lambda or Tamarind or Three Ten, there's all these companies that basically-

RJ Honicky45:57

Yeah

Andrew White45:57

...wrap a lot of these, these like, um, uh, deep learning protein design tools or chemistry design tools. They wrap them in an API, and you just give that to, to-

RJ Honicky46:04

Yeah. Right

Andrew White46:04

...give it to Cloud Code if you want.

RJ Honicky46:06

Yeah.

Andrew White46:06

You can give it to Kosmos, and you can be like, "Hey, you know, if you wanna design a protein for X, use these tools."

RJ Honicky46:11

Your mechanism, it sounds like, or one of the primary mechanisms that has been successful is, like it, like, enumerate a whole bunch of possibilities and filter.

Andrew White46:20

Yeah.

RJ Honicky46:20

Right? And so how do you think about serendipity and out of, out of distribution thinking and getting there, and how far have you gotten and what's left to do and whatever?

Safety46:26

Andrew White46:26

Yeah. That's a great question. I think... I guess the, the short answer is that there's very f-- So, so this is the domain of CBORNS, so chemical, biological, radiological, nuclear, um, uh, weapons, or-

RJ Honicky46:37

Yeah

Andrew White46:37

...I don't know, safety.

RJ Honicky46:38

Yeah.

Andrew White46:39

This domain has been explored a lot in history by a lot of organizations.

RJ Honicky46:43

Yeah.

Andrew White46:44

And, um, I would say that there was a, a big question mark for us a few years ago, was like, how much of this stuff is, uh, intellectually bottlenecked?

RJ Honicky46:53

Yeah.

Andrew White46:54

Like, how often are people like, "Oh, well, I wanna cause harm, um, but I need to know, like, some facts," and-

RJ Honicky46:59

Yeah

Andrew White47:00

...could LLMs make that easier or go faster or anything like that? I think, you know, the first set of answers in twenty twenty-three, I think, was basically no, is that, like, you know, you can go find the synthesis route for many dangerous compounds on Wikipedia.

RJ Honicky47:16

Yeah.

Andrew White47:16

People know what are the targets in the human body that, like, are, are, are targeted by most biological weapons. It's, it's not really that much of a mystery. So I, I don't think there was a lot of, like, um, there was a lot of new ground when LLMs first came about.

Then there's a lot of concern about, like, laboratory protocols. Is that could agents or LLMs, uh, reveal some tacit knowledge that, like, maybe people couldn't find on Wikipedia or, like, maybe for making something, there's some technique that req- is required when you scale it up in size or something.

Or maybe there's, like, some way to get around, like, tracking lists by ordering different compounds or something.

RJ Honicky47:50

Yeah.

Andrew White47:50

So, and that, I think, was really well tested, um, by a few different labs. Not, not me, but there were some groups that spun up that started making, like, tests for this, and it, and labs pay attention to it.

I think it's really been put into process where LLMs will, like, kind of shut down or be filtered in, in those scenarios. But I think that is actually an area where there is, is some risk. Um, and so I think that's something that people pay attention to for open source models, and there's still, I think, some, some discussion there.

But I think to a large extent, it's, it's not really been greatly accelerating in practice, or at least I haven't seen much evidence of it.

RJ Honicky48:22

Yeah.

Andrew White48:23

Um, and again, I think it comes down to the fact that it's not really available, but, like, you, if you look hard enough, you can find most of the information you would need to, to get up to no good-

RJ Honicky48:33

Yeah

Andrew White48:34

...in the public domain already. But then I think now i- is, the, the next frontier is, like, uh, can it somehow help you with real-time protocols, troubleshooting, like, more in the loop and-

RJ Honicky48:45

Yeah

Andrew White48:45

...um, and, and more especially in the computational side of things. There are some scenarios that are now coming into focus that could be more dangerous or more intellectually bottlenecked. And so I think people are trying to pay attention to that.

To some extent, there was like a first wave that we thought this could unlock a lot of stuff, and I don't think it came to pass.

RJ Honicky49:02

Yeah.

Andrew White49:02

I think there's now an emerging sort of second wave of, like, there are some actually new scenarios that were just too far-fetched to consider two years ago that I think are now realistic. Um, some smart people are paying attention to it, but I don't think it's solved yet.

RJ Honicky49:16

Yeah.

Andrew White49:16

I don't know. Uh, it's very vague.

RJ Honicky49:18

Yeah.

Brandon Anderson49:19

No, I mean, so, so I guess, like, one kind of differentiator, there's a lot of talk about AI safety-

RJ Honicky49:23

Yeah

Brandon Anderson49:23

... and, like, the broader LLM, you know, ASI space. And, you know, there it's jokes about paper- or paperclip maxing robots-

Andrew White49:31

Yeah, yeah

Brandon Anderson49:31

... or something. But, like, the, the core threat here is more, like, a malicious actor using this as a tool to accelerate something dangerous. And, like, kind of the first order hypothesis is that you basically already have to be an expert to effectively create a bioweapon or a chemical weapon.

Andrew White49:48

Oh.

Brandon Anderson49:48

And a non-expert... Or an expert would already know how to do this.

Andrew White49:53

Yeah, I, I think, you know, so, so each of the categories in the CBRN, they're, they're all a little different. But I think to a large extent, it's a, a lot of, like, pushing material around.

Brandon Anderson50:02

Mm-hmm.

Andrew White50:02

You know, the classical example of nuclear is, like, it's a lot of, lot of centrifugation.

Brandon Anderson50:06

Yeah.

Andrew White50:06

A lot of ultra centrifugation, a lot of high pressure or high RPMs.

Brandon Anderson50:10

Yeah.

Andrew White50:10

And so it- it's just, you can maybe get smarter about how to set up, you know, the, the economy of scale to do that with an LLM. But to a large extent, I, I think you can call your, your friend in country X, and they can tell you what are the steps.

It's, it's not a... I don't think it's that much of a secret. It's just a lot of, like, moving material around, and I don't think it's acceler- meaningfully accelerated. Now, that said, there are all kinds of like, you know, dumb dual use things of, like, maybe you wanna call a company that makes centrifuges, and you wanna make sure that they sell you them, and they go through some KYC steps.

And maybe an LLM can get you through the KYC faster. And that's, like, a dumb thing that like, okay, like, yes, like, uh, you know, email makes it so that you can order centrifuges off the internet more easily.

Is email, like, a dual use technology? Like, yeah, to some extent it is. And so I think there's a lot of, like, weird second-order things that we don't pay attention to in AI safety of, like, does it make KYC easier?

Does it make it easier for people to know, like, where, where to order this from? Or, like, what is the expected price? Or, like, what should you order first, right? All those, like, sort of simple logistical things I think are accelerated by AI, just as, like, a, a consequence of AI being an accelerating technology.

Um, but certainly, I mean, shit, guys, there's some scary stuff and- I try not to think about it too much and, uh-

RJ Honicky51:27

Yeah.

Brandon Anderson51:27

Yeah

Andrew White51:27

... I don't know. I guess I don't wanna get too political, but I do think that right now, um, the, the United States government is maybe taking a, a slower, less intensive look at safety. And, um, but there's definitely people, I think, in other spaces than the US government thinking about it hard.

Brandon Anderson51:46

And do you think it's a thing people need to spend more time on? Mm-hmm.

Andrew White51:49

I do get waves of angst about AI.

RJ Honicky51:52

Yeah. Yeah.

Andrew White51:52

And I'm sure many people living in San Francisco do get like a little bit of, uh, a little bit of waves of it. And, uh, sometimes I think that there isn't enough work being done on it. And then sometimes I think, "Wow," like, "I need to mellow out," and, like- ...

you know, we have lots of time to think about it. What is my opinion on it then? I don't know. I, I, I think my opinion is, um, not formed fully.

FROs52:13

RJ Honicky52:13

Yeah. You and Sam have done a lot of thinking about funding science.

Andrew White52:18

Yeah.

RJ Honicky52:18

And future of science. You have, you've been vocal about the reproducibility crisis and other things. First question, why this focused research organization or for, uh, yeah-

Andrew White52:30

FRO.

RJ Honicky52:31

Yeah.

Andrew White52:31

Focused Research Organization.

RJ Honicky52:32

Yeah. What, what does that get you that you don't get from academia or, you know, big lab or whatever?

Andrew White52:38

Um, a nice network of, of people, and I think Edison is, like, a real, uh, of course, I think Edison's gonna do great, but I think it's a mystery of what's gonna happen. Um, so I don't think we've had as much friction there as you might expect.

Um, but yeah, this is all stuff that we, that, that, that Sam and I think about all the time is, like-

RJ Honicky52:54

Yeah

Andrew White52:54

... how do you balance stuff like this? How do you balance the economics? Um, you know, there are some, there are some venture-backed companies that are s- having cash salaries over a million dollars, and it's, like, insane to me-

RJ Honicky53:07

Yeah.

Brandon Anderson53:07

Yeah

Andrew White53:07

... that you would use all of your cash from your equity financing, you know, in these insane salaries.

Brandon Anderson53:14

But they can s- in, in terms of, like, total spend on GPUs, it can still be a total, a small fraction of your burn, so sometimes it kinda makes sense.

Andrew White53:21

Yeah. Yeah. That's, that's one way to think about it.

RJ Honicky53:24

So, so, w- like, you... This is a good, uh, lead-in to y- you are automating science-

Andrew White53:32

Yeah

RJ Honicky53:32

... in some capacity.

Andrew White53:34

Yeah.

RJ Honicky53:34

So where does that leave scientists?

Scientists' Future53:34

Andrew White53:36

So I think, um, this is a Jevons paradox we can try here- ... is, uh, um... So, uh, let me start with the contrast here is that, uh, you know, if we automate, um, you know, taxi cab drivers, uh, there's a fi- there's not gonna be an increase in people needing to go places.

Maybe there'll be somewhat an increase, but, like, there is a finite amount of, like, time people will be spending in cars.

RJ Honicky53:58

Yeah.

Andrew White53:58

And so there's an upper limit. So when you automate that, that's like a scarcity thing, is basically you're displacing jobs when you automate driving.

RJ Honicky54:05

Yeah.

Andrew White54:05

In science, I don't think there is a finite appetite or a finite capacity for science. I don't think science is, like, a, a scarcity thing. Like, there's, you know, 100 more discoveries left to be made, and then we'll be done, and so, like, we're displacing jobs.

I think instead, actually, if we can, you know, make science go much, much faster, there will be no, there will be no decrease in demand. There will be actually, I think, an increase in demand that will match whatever automation amount we have.

And so my vision for what a scientist would be in the future is that they will be, I don't know, like, uh, agent wranglers or Kosmos wranglers of, like, okay, they're exploring 100 ideas simultaneously, or they're, like, working with systems like ours to, to make 10X the discoveries, 100X discoveries, because I think there's an unlimited amount of scientific discoveries to be made.

And so there's no, like, scarcity set where basically we will displace them all. Now, that's kind of like, you know, this is what I would tell when I go talk to a first-year PhD student.

RJ Honicky55:02

Yeah.

Andrew White55:02

Like, "Everything's gonna be just fine." You know, but then when it gets into the nuts and bolts, I, I do agree that this is gonna be, like, a really hard thing, where, like, if I am CEO of a company that makes science, like a pharma company or a material science company or something like that-

RJ Honicky55:14

Yeah

Andrew White55:14

... or a R&D arm at, at IBM, I think while I could spend, you know, a million more dollars on, on compute for the AI scientist, or could hire 10 more people, I might just choose to go with the AI scientist because, you know, to a large extent, like, hiring people is hard, right?

RJ Honicky55:30

Yeah.

Brandon Anderson55:31

Yeah.

Andrew White55:31

And, and hiring an AI scientist is probably a little bit easier.

RJ Honicky55:34

Yeah.

Andrew White55:35

And so I think that there could be some, there could be some friction. But another thing is, like, science is in some ways closer to art in the sense that, like, there is a large number of people who just appreciate good science.

Like, if you get published in Nature, it's not because it's really gonna be world-changing. Of course, that's part of it, but it's also because, like, people are like, "Wow, this is really interesting science."

RJ Honicky55:56

Yeah.

Andrew White55:56

So I think the, the enjoyers of science are also scientists, and so I think that it's kind of hard to imagine a scenario when there's not scientists as the consumers of science. And so I think if they're gonna be consumers of science, they're also gonna be some of the producers who are involved in the, in the process by itself.

RJ Honicky56:10

Right. Yeah.

Andrew White56:11

I don't know if that makes any sense.

RJ Honicky56:12

Yeah. You touched on this. My- the question in my mind is just what does a scientist do then?

Andrew White56:16

There's a great short story, um, by, um, Ted Chiang, I think in like 2003 or something.

RJ Honicky56:21

Okay.

Andrew White56:21

And it's about, like, well, at first, scientists were displaced, and they became, like, the, uh, interpreters of, like, what the AI scientists are doing. Like, the scientists read the AI scientists', like, papers and then, you know, translate them for whatever Popular Science or something.

RJ Honicky56:34

Yeah, yeah.

Andrew White56:34

And then after that, like, they couldn't read the papers anymore, and so they were left behind, and so they had nothing to do, and they just sat around. And but the problem is that science is like, you know, you, you have to translate science to make any impact.

Like, science cannot exist by itself. I do agree there's, like, engineering can exist by itself. Like, if you give some kind of system a goal of, like, making a new material that it can make a space elevator out of, you could be not participating in the beginning of the process or the middle of the process, and you just come by the end and be like, "Okay, follow this recipe."

But, like, science of, like, what's the origin of life? Or like, is there water on another planets? Or, you know, um, why is some catalyst better than another catalyst? That has to be hitting human eyes and human brains at some point, so I think a human has to be involved in the process.

RJ Honicky57:21

Don't wanna be contrary but-

Andrew White57:22

Yeah, be contrary

RJ Honicky57:23

... why, why does a human have to be involved?

Andrew White57:25

Why does a human have to be involved?

RJ Honicky57:25

Yeah.

Andrew White57:25

Well, a human has to be involved at at least some point to be like, "Yes, this is good science," or, "This is bad science."

RJ Honicky57:29

Oh, okay. So y- y- it's, it goes back to taste.

Andrew White57:32

Yeah. But I don't know. Maybe you're right. Maybe there is no point for humans. Maybe it will be like, you know, what is it? Sora. Uh, you know, like the AI slop app. But I think in Sora, there's still humans at the end clicking the videos or something.

RJ Honicky57:43

Yeah.

Andrew White57:43

So.

Brandon Anderson57:44

So, so the bring, the, the Sora analogy kind of brings up a interesting point. Like, is it possible that, like, due to the biases of AI science, if we really go full in science, that, you know, there still is a market for kind of boutique human-

RJ Honicky57:59

Boutique

Brandon Anderson57:59

... science, like-

Andrew White58:00

Yeah, yeah

Brandon Anderson58:00

... you know, there's still people who wanna, you know, paint things the old-fashioned way. But more to the point, does it become even more important for to have a human who is, uh, actively doing their own exploration because there will be, like, large blind spots and biases due to the models that just you'll never be able to overcome because this is sort of baked in now, um, due to your training data.

Andrew White58:23

Yeah.

Brandon Anderson58:23

And without a human, y- that, you'll always get stuck, and there will be a blind spot that will never...

Andrew White58:29

Bio, which is a company in, in Oakland, um, or in Emeryville, they do really cool stuff with automation. I think they're gonna be testing this theory of like, okay, maybe if that's the bottleneck, we can see evidence of it because they're gonna start doing really well.

Brandon Anderson58:40

Yeah.

Andrew White58:40

Um, it could be true.

Brandon Anderson58:42

Mm-hmm. I, I still though wanna say, all of those I, in my mind, are still sort of scoped in terms of, like, R&D for pharma or bio, but they're not like, none of them are attempting to answer big fundamental questions.

And maybe there's, like, different levels. When I think about the-

Andrew White58:58

Yeah

Brandon Anderson58:58

... you seem to be, um, y- it seems like the Future Hou... the, the focus of Future House and Edison is much more towards, like, you know, sorta R&D and sort of end-run science. But, um, you know, I, I have some background in, you know, fundamental physics.

Andrew White59:14

Yeah, yeah.

Brandon Anderson59:14

Um, you know, it's like, is there any thought about, like, how do you, like, take on, you know, dark matter candidates? And like-

Andrew White59:22

Yeah

Brandon Anderson59:22

... I just, you know, think the data to really give us a complete story is just not there yet.

Andrew White59:27

You know what? Like, uh, I'm sure everybody at every company f- like, is the biggest critic of their own product-

RJ Honicky59:32

Yeah.

Andrew White59:32

You know?

Brandon Anderson59:32

Yeah.

RJ Honicky59:33

Yeah.

Andrew White59:33

So-

RJ Honicky59:33

Yeah

Andrew White59:33

... we think Kosmos is, we think it's great, but there's an, a very large amount of area for improvement and

RJ Honicky59:39

S- so with Kosmos, can... So th- there's, like, a open, like, sort of access to everybody version.

Andrew White59:45

Yeah.

RJ Honicky59:46

Uh, do you provide access to other labs that, um, is less open?

Andrew White59:52

Um, we have a version of Kosmos- ... that has, like, um, bigger resources. Like it can-

RJ Honicky59:57

Yeah

Andrew White59:57

... run for longer. It uses GPUs.

RJ Honicky59:59

Yeah.

Andrew White59:59

Um, so, like, basically when it does data analysis, it'll have a GPU.

RJ Honicky1:00:02

Yeah.

Andrew White1:00:02

So we use that for things like, um, like machine learning experiments. You know, if we wanna know, like, this question about whether it's better to pre-train first on noisy data or not.

RJ Honicky1:00:11

Yeah.

Andrew White1:00:12

Um, we have, like, pre-release models that, that are coming out, and we try those. But, um, yeah, so I guess-

RJ Honicky1:00:21

Yeah

Andrew White1:00:21

... like, yes, we do. When we do have, like, research partnerships with, with companies where we, like, build something specific for them, and that is something we think about.

RJ Honicky1:00:28

Yeah.

Andrew White1:00:29

But broadly, I would say Kosmos that's on the website is pretty close to-

RJ Honicky1:00:32

Yeah

Andrew White1:00:32

... to what is the best we have internally.

RJ Honicky1:00:35

Yeah.

Brandon Anderson1:00:35

Uh, I have a question. Um, so you, you previously have stated that you think that language is the natural, um, language. Was it-

Language Argument1:00:44

RJ Honicky1:00:45

Language of chemistry

Brandon Anderson1:00:46

... the future of chemistry is language.

Andrew White1:00:47

Yeah, yeah.

Brandon Anderson1:00:48

Yeah, yeah. Um, okay. So, uh, I wonder, do you still believe that?

Andrew White1:00:53

Good question. I think, I, I w- I would say yes. I still believe that, um- That, so, so in that article, that opinion article, my, my point was that, uh, you know, at the time when I wrote that article, which I think maybe three years ago now or something, maybe 2023, um, it, it was that we have models for predicting solubility of compounds.

We have, like, data about very large populations, and we have, like, papers, and we have code. And, and the only way to bridge all that information is natural language. And, and the argument was that, like, humans, like, you know, whenever we can't bridge information, like if I can't talk about my code or I can't talk about some idea to you, I will invent words until I can get the point across, right?

Brandon Anderson1:01:36

Mm-hmm.

Andrew White1:01:36

And that humans are always innovating on language to make it represent all known observations, and people innovate on language to represent whatever code pattern they have, right? Like, this is like the, the only shared activity we've been doing for this long is like coming up with words to represent everything we know.

Brandon Anderson1:01:51

Mm-hmm.

Andrew White1:01:52

And so I think for that, for that reason, natural language is the only possible way to connect all the different pieces of data we need in biology and medicine, or any domain for that matter. Um, I think there's some caveats to this of like, you know, you can make an argument, like if Yann LeCun were here and he would make an argument about like, you know, world models or like vision or embodiedness, right?

Like the, there's arguments against natural language that like, you know, that maybe there's something more that it, it does, it's not the complete story. Or maybe natural language imposes limitations-

Brandon Anderson1:02:21

Yeah

Andrew White1:02:22

... you cannot exceed because you are stuck in this abstract space that was invented by humans and you can't escape it, until you can like touch something.

Brandon Anderson1:02:28

Yeah. I mean, it, it is an abstraction, right? And like, like scientists basically work exclusively in abstractions to some degree.

Andrew White1:02:34

Yeah.

Brandon Anderson1:02:34

Um, I, I just, I find, I found that interesting because it seems like most scientists you're right, like when they explain things, they explain things through language. But, uh, many conversations, maybe most at some point result in people drawing diagrams or something.

Andrew White1:02:49

Yeah.

Brandon Anderson1:02:49

Like, you know, chemistry, like biochemistry largely, or, or medicinal chemistry is oftentimes a, it's, it's a language of graphs, right?

Andrew White1:02:57

Yeah.

Brandon Anderson1:02:57

Or, you know what I mean, bonds are abstractions, yes, but like they're pretty good abstractions for most ca- for many cases.

Andrew White1:03:05

Yeah.

Brandon Anderson1:03:05

Or like, you know, geometry, you know, thinking about, you know, protein as like a, the geometry of a protein.

Andrew White1:03:10

Yeah.

Brandon Anderson1:03:11

You know, it's like I think that that's how people will... a lot of scientists like to think about things. And, um, so I find it interesting that like, yeah, that, that you are focusing primarily on language. Like have you thought about essentially a multimodal version of this?

Like where, you know, when it comes along a smiles per string, it doesn't just say, "Oh, this is a smile string," but like, this is a graph, this is a representation of some higher like abstract object.

Andrew White1:03:36

You're absolutely right. And, and the problem with these, this like, I don't know, Jacob's Ladder or something, whatever you wanna call it, is like yes, you can say that a mol- you can call a molecule by its name.

Brandon Anderson1:03:45

Mm-hmm.

Andrew White1:03:45

You can show the graph.

Brandon Anderson1:03:47

Yeah.

Andrew White1:03:47

Then if you go to a molecule like ferrocene, well, it doesn't really have bonds-

Brandon Anderson1:03:50

Mm-hmm

Andrew White1:03:50

... between like part of it. And so then you're like, well, we need to draw it visually.

Brandon Anderson1:03:54

Mm-hmm.

Andrew White1:03:54

And then you go to a molecule like, I don't know, cyclobutane.

Brandon Anderson1:03:57

Mm-hmm.

Andrew White1:03:58

Well, there's dihedral angle, right? And so like it's not actually this thing I drew, it's actually an ensemble between this thing and this thing, right?

Brandon Anderson1:04:04

Mm-hmm.

Andrew White1:04:05

Then you go to benzene, you're like, well, not only is it like a ensemble of these different conformers, it actually has electron density and you can't really ignore the electron density in benzene. You like need to treat it correctly.

And then it's like, well, you can't actually represent the electron density that way. You actually have to look at the correlation of the electrons individually, right? Because you can't really model benzene with like DFT, right? Or functional, you have to actually look at the, the electron correlation.

Then you see the electron correlation like, well, you know, you can model electron correlation, but you know, actually these things, when they're in a solution, they have like, you know, relativistic effects because it's like there's a whole bunch of stuff around it.

So you really gotta have the relativity in there. And you're like, well, you can have the relativity and you have the electron correlation. You could have the bonds and, and you have the conformers, but you really need to think about the cosmic radiation background because like, you know, it does actually impact everything and there is some, some energy there, right?

Brandon Anderson1:04:47

Mm-hmm.

Andrew White1:04:47

And before you know it, you've ran out of, you know, you've ran out of compute or whatever resource you're using to model this. And so I think, um, you have to draw the line somewhere.

Brandon Anderson1:04:58

Mm-hmm.

Andrew White1:04:58

Natural language, like I said, is that humans have worked for a long time to make it be the, you know, what's the word? Like the least abstract- ... or the, you know, it's somewhere on the border of like it's still abstract enough that you don't need to know all these details.

Brandon Anderson1:05:14

Mm-hmm.

Andrew White1:05:14

But it's still granular enough or con- concretized enough that you actually can make use of it. Um, there may be some other representation like multimodal. Might turn out the video or maybe, I don't know, there's some other like fusion that you can make.

I like natural language because we all work really hard to make it right at that boundary. And I do agree, sometimes, sometimes ideas slip and they can't be in language, and you have to get out the whiteboard. Or ideas slip and you have to wave your hands around, you know?

Or-

Brandon Anderson1:05:37

Mm-hmm

Andrew White1:05:37

... maybe then, then you need that, that, uh, degree of freedom to communicate.

Brandon Anderson1:05:43

Just digging in on this a little bit-

Andrew White1:05:44

Yeah, yeah

Brandon Anderson1:05:44

... more. Like, uh, famously quantum mechanics is like undescribable, right? Like the, there's, there's an argument that you cannot understand quantum mechanics with words. It ha- or in, and with our preconceived understanding of the physical world because it doesn't behave like the macroscopic world.

And so the, the only way to understand it is through mathematics, right? Um, and I largely see language as the joint key of science as well, but I wonder if that's not true for many domains, and quantum mechanics is just the one that hits you in the face.

Andrew White1:06:18

I mean, I don't know. Actually, I think the, what, there's like seven principles of quantum mechanics or five or something like this, that you can actually express pretty concisely in language.

Brandon Anderson1:06:27

Mm-hmm.

Andrew White1:06:28

I agree that like you need to actually look at the consequences of them. You need some mathematics. Um, I don't know. I actually, I don't know. This is like a challenge.

Brandon Anderson1:06:36

Yeah.

Andrew White1:06:36

I think you could actually describe a lot of quantum mechanics in language.

Brandon Anderson1:06:39

Sure, sure.

Andrew White1:06:40

But, but I, I see your point and, um, yeah, I, I guess, uh, uh, I'm a realist. Like I, when I talk to my kids, you know, maybe I, I will be like, "Okay, let me draw it for you."

I don't, I don't make sure in our house everything is described with natural language. Uh, so I, I agree with you there. Um, I think maybe we can be a little, a little flexible with, with natural language and include equations and smiles strings in it-

Brandon Anderson1:07:03

Mm-hmm

Andrew White1:07:03

... and I think we can get a little bit farther. Um, so maybe that's okay. Uh, but Some people, I think, like optionality. You know, like the, "Oh, it could be this or it could be that." I'm somebody that like to take, take strong opinions- Yeah ...

and see how much farther they can get me. And I think in, in my career, it's actually been better for me to take strong opinions- Right ... which in my deepest of hearts I know that are maybe not correct or not fully correct.

But once you take these strong opinions, it just, you can sort of move many steps down the road once you take these strong opinions. Right. And like, for example, at Future House, we took the opinion that scientific agents are the future, and that allows, skipped a lot of steps, because a lot of other people were like, "We need to bu- build a foundation model for X."

Yeah. And we just skipped all that, right? Right. And I think if you also were unopinionated and you had optionality, like I can think of a famous example of a different company that, like, liked the optionality, and they wasted a lot of time on foundation models for something, then, then I think you, you get stuck.

So that's one of my strong opinions is that natural language is a, a, is- Yeah ... a way to join all these different domains. Yeah. It may not be a correct opinion, it may not... It may be more subtle or more complicated, but it's allowed me to get very far.

Um- Yeah ... maybe I'll drop it someday and- ... maybe find a new one, but yeah. Not yet, though. That's my meta- Yeah ... opinion- Yeah ... on the matter.

Brandon Anderson1:08:16

The EtherZero story on your blog I find hilarious and kind of awesome.

EtherZero1:08:16

Andrew White1:08:21

Yeah.

Brandon Anderson1:08:21

You know, when I was a kid, I loved the, like, g- genie/monkey paw-

Andrew White1:08:26

Yes

Brandon Anderson1:08:26

... like, concept of be careful what you wish for because you just might get it.

Andrew White1:08:30

Yes.

Brandon Anderson1:08:30

Maybe just, like, quick story. Can, can you just talk about that? That was, that was just a really fun-

Andrew White1:08:37

EtherZero was a, a hell of a project because conceptually it was a very short project of like, hey, people have made a lot of progress in verifiable rewards in math and in comput-i-i-in, in code, let's see if we can do it in chemistry.

So chemistry is, like, not a verifiable field, right? Like, of course you can go test it in a lab, but then we like had to think about all these, like, ways that we can make chemistry verifiable. And one of the ones we settled on was like, make a molecule that has like three nitrogens, two oxygens, 10 hydrogens or something.

And we thought that was, like, a pretty verifiable, ver- pretty verifiable question. But every time we would train a model, it would find some new insanely weird trick to generate these molecules. And, and I, I'll just tell you one of the examples was that, um, uh, it would make these molecules, and we would do some checks to make sure, like, it had the right bonds, the right number of electrons, the right number of atoms and stuff like that.

Um, but it would just solve the problem in any way possible, right? Yeah. So, like, it would just put all the nitrogens over here, put all the oxygens over here, just, like, things that don't look good. Yeah. And so we started coming up with these rules of like, oh, let's check to make sure it followed these good practices or these good practices.

And we found ourselves into this, like, you know, it's like the opposite of the bitter lesson, like, I don't know, the boutique lesson- ... where you, like, try to make everything custom. But one of the things it kept doing is it kept putting these nitrogens in a row, and it put like one nitrogen, two nitrogen, three nitrogen all in a chain.

And this is, like... You know, if you have three nitrogens, it's, like, explosive, you know- ... two nitrogens is, like, bad, and, like, four nitrogens you can't make. And it kept telling everyone, like, it would make these, like, six-nitrogen compounds, and they're just, they're just literally impossible, and they're not possible.

And, uh, many of the people on the team were, like, computer scientists, like, on this team, and one of them, like, one day sent me the, like, "This is on the cover of Nature today on Nature's website. Somebody made a six-nitrogen compound."

"And this is, like, somebody's, like, career to deliver this compound because this is the most unstable, like, insane compound you can make." "It's some ridiculous setup, and it, like, the spectroscopy to get that proven was, like, very difficult-" Yeah ...

"and this, this... I don't know how they did it. It was an amazing accomplishment." I'm like, "Look, Andrew, like, it's not actually impossible." And it was so funny to me that, like, our model is sitting here spitting out these six-nitrogen compounds in, like, you know, 2024 or 2025, and, like, the paper just happened to come out that year- Yeah.

... that, like, mankind had finally made a six-nitrogen compound. So do, do you think that those were actually synthesizable even under these extreme circumstances? No, no, our model was just, it was just- It just got- ... reward hacking.

Okay. It, it was just the, the model was so creative in ways to reward hack. Yeah. Like, one of the o- another one we did was, um, you know, we, we wanted it to make sure that the, when it would propose a reaction, like make this compound, tell me how to make this compound, we would try to make it sure that all the reagents were purchasable, like you could purchase them.

They were not, like, made up. Yeah. Um, and, and the reason we came up with that is that originally we would just, like, take the end- ... compound and then, like, remove one atom and be like, "Here's- Buy this ...

buy this." And then put the atom on, and it's like, okay. It's like, well, that's r- I wish it was like that. Um, so they, they had to be purchasable. And then well, we're like, we thought it might be hard if they're all purchasable because sometimes you actually order things custom or, or something.

So we went, "We'll just make sure one purchasable." So the first thing it starts doing is just putting nitrogen in there because nitrogen is purchasable, and it, like, has no participation in the reaction, right? Like, oh my God, okay.

So then like, okay, it has to be purchasable, it has to participate in the reaction. Then it started just putting, like, acid base chemistry. We'll just put an acid here. Acids are purchasable, and it'll move one atom. And then we go, "Okay, fine.

Can't be that. Everything has to be purchasable." Then we find ourselves, and I'm, like, sitting there one day building this, like, ridiculous catalog of purchasable compounds in a bloom filter so it can go fast enough in our training loop- Yeah.

... and I'm like, "Why am I doing this?" How did I get here? How did I get here? And, and I don't know. It was really funny because, um, pre-training or training transformers, you know, o- o- on, on just data, like just supervised training where you just have the inputs and the outputs directly- Yeah ...

very nice, relaxing. Yeah. You know, like, things are always robust. You know, things are, uh- Yeah ... go pretty smoothly. When we do these verifiable rewards where you have to, like, write a, a bulletproof verifier, it is really difficult.

Yeah. And we had so many models trained only to find out they were hacking some other, like, random thing- Yeah ... in our setup. It's really hard, and I, and I, I don't envy the frontier labs that have to do this at a very massive scale- Yeah ...

because we had a lot of adventures in EtherZero. And, and you guys should read the, the blog post. It-

Brandon Anderson1:12:41

Yeah, definitely read the blog post.

Andrew White1:12:42

Yeah. It is very fun.

Brandon Anderson1:12:43

It is a great read.

Andrew White1:12:43

GRPO? We did make some modifications, um, to GRPO. Yeah. Um, I actually, I used to know all the names of these modifications. Yeah. But, uh, uh, I think it's like, uh, DAPO is one modification, and, like, the clipping we did was special.

I mean, we explored a lot of that stuff. Yeah. Um, and it was, uh, also one of these things where, like, you think the hypers are wrong, the algorithm is wrong, and then you find out it's just because, like, you had somehow sorted the reagents when you made your training data, but in your test data you didn't sort them alphabetically, and the model was just, like, barfing because its whole strategy was to exploit something in the way you sorted things.

So yeah- Yeah ... we, we explored a lot of different methods, and it was, um... I learned a lot about chemistry, a lot about nomenclature. Um, and actually there's a, I learned a lot about medicinal chemistry as well, more than I ever wanted to.

Brandon Anderson1:13:32

Awesome. If you wanna do some, like, engineering, just check out Edison Scientific, and they have, you know, I think a lot... They're hiring with lots of, like, interesting things, everything from scientists to, you know, infrastructure engineer.

Andrew White1:13:45

Yeah.

Brandon Anderson1:13:45

Yeah. Thanks, Andrew, again.

Andrew White1:13:46

Yeah. Thank you very much for, for joining us.