Solid Gold Magikarp0:00
All right. We are here in the remote studio for a Lightning Pod with Dr. Jessica Rumbelow herself. Welcome to the Lightning Pod. W- welcome to LeanSpace.
Thank you so much. Thanks for having me.
You are very, very famously known as the discoverer of the solid gold Magikarp. I think that will go with you to your grave.
Yeah. Never-- I'm never gonna get rid of that bloody fish.
Well, it's a, it's a cute fish. It, uh-
Yeah
... it woke people up to tokenization and-
It's a sweet-
... the mysteries of Reddit.
It was fun work. It was really fun. I'm glad I did it.
It got you interested in interpretability, but since then you've, you've sort of, uh, gone into scientific discovery, right?
The Journey0:39
Yeah.
With, with Leap Labs. And so basically I just want a, a clean slate. How did you find your way to what Leap Labs is today?
Sure.
And, and then we'll, we'll launch into what Leap Las- Labs is today. I just wanted the journey a little bit.
Yeah, of course. So actually, my interest in interpretability for scientific discovery predates solid gold Magikarp by some time. This was, um... I was, I was working in academia. Um, I was building neural networks for different tasks in histopathology.
So doing things like, "Let's build an image classifier to diagnose different cancers," or, "Let's build a segmentation model to, uh, to pick out different types of immune cells from a biopsy." Interesting work. Um, and during that kind of time, I, I ended up building, uh, a, a classification system to identify different types of immune cells, but did this-- it did this in a way that human pathologists didn't understand and couldn't replicate.
It was from a very, very cheap biopsy stain, and you can identify this, this particular kind of immune cell, or rather, like, you the scientist, you can't. You don't know how to do this. The model I trained could do this.
And in one sense, that's pretty cool. You have a new capability. You can now automatically identify these immune cells. That's very useful for, like, personalized immunotherapy, stuff like that. But it, it occurred to me that, like, the capability isn't really the interesting part here.
What's interesting is, how is the model doing this? Like, what, what signals has it found in the data that enable this capability? Like, what does it know that we don't? And that kind of sparked my, my interest in interpretability, because it's a very, like, in a Chris Olah kind of sense, if you've got really good interp, you can start to reframe neural networks, not as just a tool for automating things that we already know how to do, but as a tool for, for discovery, as, as like a lens through which you can see patterns in data that would otherwise be hidden from you.
Because humans are very bad at seeing patterns in data. Um, neural networks are really great at it, and that's what they do. And, um, several years later, worked in interp for a while, did a bunch of research. Thought, "What shall I do now?"
And I had this thing in m- in the back of my mind, I was like, "Okay, there are loads of applications of interpretability." There's this big one that is this notion of scientific discovery, train a neural network to do something that you don't know how to do, and then-
Mm
... interpret that neural network to learn a new thing. And nobody else was doing it, so I thought, "Well, I guess that, that better be it." And initially, Leap, as, as you know, so it was, was very much like we're building interpretability tools, so, like, anybody can interpret their neural network.
And this is like... it's useful. You know, you can identify bias, like, is your model sexist or racist or some other terrible thing, and you wanna make sure that it's not before you deploy it so that you don't get embarrassed later.
You might wanna do things with interpretability, like predict failure modes be- before, before you deploy the model. That's really important as well. You, you might wanna understand the model better so that you can intervene on it in various ways, right?
And make it, you know, make it more poetic or more, I don't know, what- whatever other personality traits are important to you, stuff like that. Lots of applications of interpretability, but yeah, ultimately-
Yeah
... ab- about a year into Leap, we decided that the most interesting one was science.
Yeah. Uh, just to comment on the interpretability stuff before we go on. You know, I think, like, let's say Goodfire and, and Anthropic have, have picked up where, where you left off. Um, I always think of it as...
I, I explain it to people who, who don't really get it. Imagine if you had a knob that is not just temperature to turn up and down-
Mm-hmm
... and you could actually turn it up and down, uh, you know, different va- various qualities. And I think the most interesting one, uh, I've been theory- theorizing about this, is if Anthropic ever found the activations for reasoning and could break down the different kinds of reasoning and tune, tune them properly, I think that's, that's the thing that takes Anthropic to leapfrog OpenAI.
Huh. Do you imagine that there is a reasoning vector or, like, a vector for different types? I would imagine that it's kind of like suffused in the whole model, no? I don't know.
Well, if you haven't tried, you don't know.
Uh, yeah, that's true.
I think, I think it's more likely that reasoning actually decomposes to, like, probably 100 different things, and if you just figured out which two of the 100 different things mattered, you could just focus on those.
I guess it's gonna depend a lot on the granularity of that decomposition, right? Because maybe it decomposes to, you know, when this token and that token in this order, then it's a French verb. And it's just made up of, like, stacks of this stuff.
Yeah, spurious.
Yeah, yeah, yeah.
Like, on, on the whole, it's gonna-
I mean, you need some hierarchy, right?
Yeah, yeah.
But I don't know if you-
Okay. Not, not to digress too far. I wanted to, to obviously go on to Discovery Engine.
Discovery Engine5:40
Mm-hmm.
So I have your slides here. Why don't you tell us a little bit about what we're looking at?
Yeah. So this, this is a, a, a screenshot of Discovery Engine.
There we go.
What we've done-
I think it freeze
... is, is, like, ta- take this insight that is something like- Deep neural networks find patterns in data that humans miss, and quite often those patterns are, are novel. They are new to us. We can use this to learn a new thing about the world, and we have automated it.
So Discovery Engine is a end-to-end system, takes in arbitrary scientific dataset, automatically trains a bunch of neural networks on it, and then we systematically, with our interpretability methods, which is the real secret, um, extract the patterns that have been learned by those models, and then we contextualize them with existing literature.
We, we rank them by novelty and prevalence in the data, stuff like that, and we make them human possible, and, uh, we made a bunch of novel scientific discoveries doing this.
Yeah. Um, what are some examples or do we, do we go to the demo now?
Go- Okay. So for, for the novel discoveries-
Yeah
... we've-
One of these, right?
... we've done a few publications.
Yeah.
On the slides there is ... Oh, yeah, here's, here's one of them. Yeah, so th- so this was, like, novel, novel markers associated with ... Well, you can see the T-cell receptors in physiochemical signatures. I want to caveat all of this by saying I'm an AI researcher.
Well, no, I'm the CEO of Leap Laboratories now, but I am absolutely not, like, an immunologist or a computational- ... bioinformatician. For the domain-specific context, you'll have to read the paper and look at the, the wonderful things that our scientific collaborators, who are the domain experts who have used these tools, have written about the discoveries that we made.
As you can see, identifying T-cell receptors is really important. We want to understand which ... like, why certain T-cells are, are reactive to tumors, why they, you know, like, get going and, and help fight the tumor and why some are not.
Understanding this, predicting this is non-trivial. It's, uh, very much an ongoing field of research, and we found basically novel markers that, that are quite predictive of this, and these novel markers were not what we expected them to be.
Yeah.
Right. Is there a different one that you wanted to take a look at? What, what, what about the other paper that you picked?
Yeah. So ... Oh, yeah. So th- this is the one that ... This was actually our first ever case study that we did-
Nice
... with, like, real, live, scientific data that, um, was, was not, like, an open source dataset or something. So we worked with a, a plant biologist called Matt, um, and Matt is very interested in understanding, like, what, what are the factors that determine the root structure of, um, crops, of plants?
And this sounds kind of esoteric and weird, but it's actually really important if, like, you're trying to grow crops that are robust to very drought-prone areas or areas that have lots of floods. Like, the structure of the roots will determine whether your crop succeeds or fails.
So having, having kind of levers you can pull in terms of I'm gonna put this nutrient in the soil or I'm gonna plant this particular mutation, stuff like that, is, is crucial for, like, food security in these regions.
Um, but Matt only had a small dataset. He had, like, 700 samples. He'd already spent a couple of months doing his own data analysis, which he hates. Um, and yeah, we w- we weren't really expecting to find anything world-shaking here, but we thought, "Hey, we'll validate the Discovery Engine.
You know, it'll, it'll be a nice study. We'll confirm what he knows." And, and we did. We found, like, a bunch of patterns that he was already very aware of, and that was fine, but we also found this new thing.
We found a bunch of new things, actually. This, this is one of them, which is a particular, like, synergistic effect between, um, manganese content in the soil and a, a particular genotype of this, uh, this test crop, which actually has a really profound effect on the root architecture.
Um, so, so this was great. This is very exciting. The other link that I gave you is the meteorological study, which is-
Ah.
Yeah. This, this is the one. Uh, our spectacular, uh, research scientist, Jack, became a, a moonlighting meteorologist and did so much work on this.
He taught you about wind speed gradient?
Yeah, and he knows all about it now. Yeah. It's, it's actually really cool. So meteorological modeling, right, which underpins so much stuff-
Weather, yeah
... like, like, like where do you put offshore wind farms, like, can you predict the path of a hurricane accurately. Like, all of, all of this modeling, um, relies on a, a pretty foundational assumption, um, which is, uh, most ...
It's like the surface layer theory, and it basically posits that various fluxes are constant within, I think it's, like, I don't know, two, 200 meters, um, above the, the land or water surface. This, as it turns out, as we found, as Discovery Engine found, is not true.
In, like, 20% of ... uh, 20% of samples that we worked with, mostly in, like, coastal locations, this assumption does not hold at all. And so these scientists, the wonderful guys, Patrick and Steve, that we, we worked with at the National Center For Atmospheric Research doing this study, like, this is huge for them.
They would never have found this otherwise. They weren't even thinking to question this assumption because it's so foundational. Yeah, and it turns out a chunk of the time it's not true. This says ev- improving meteorological modeling by even a couple of percent is valued at billions of dollars.
So this is probably my favorite.
That's crazy.
Yeah.
Yeah.
Favorite of these.
So when, when do you get called into stuff like this? This one's a client, so you got called in pretty early.
These, these are all academic collaborations that we've done. Yeah. So basically through our network.
Claude & Discovery12:14
Yeah. Okay. So I mean, I, I think, like, this is all great. I, I think, like, if people are in those particular fields, they can already make analogies to, like, what else they might wanna do with you. Um, for something, um, maybe for, like, one of the last details that we prepped was the Claude Meta Discovery Engine.
Yeah.
Yeah.
So this, this is something new that we've been thinking about a lot, and which I'd love to get your take on it, Swyx, as well. So language models are just that, right? They're models of language. They are not particularly well-suited for understanding arbitrary, you know, like big numeric data sets, so that's fine.
Maybe, like, you can get them to use tools and stuff, but still, even if they're, you know, messing around in, in Python and doing Matplotlib and Pandas and, like, and whatnot, they're still basically doing the same kind of thing that a human would do.
Like, this very path-dependent, hypothesis-driven, like, I have an idea, let me go and check that. Oh, like, oh, maybe it's true, maybe it's not, maybe I find something else out. This very, like, iterative process which imports so many assumptions and biases, you know?
It's all hypothesis-driven, so we're only really looking for stuff that we think is gonna be there to begin with, and this massively, like, limits the possible discoveries you might make because you're exploring only, like, a tiny fraction of this-
Things you think are possible, yeah.
Yeah, exactly. So obviously we're doing something quite different. We're relatively hypothesis-free. Data goes in, we find all of the patterns. Like, to the neural network, it's all just numbers. They're not bringing in... I mean, we can argue about, like, imposing priors through architecture choice and all, all kinds of stuff like that, but, um, broadly speaking, it's, it's just numbers.
So, so we find things that humans wouldn't think to look for through this method. And as we see, like, increasingly language models and agents built atop language models being applied to scientific discovery, they're really missing quite an important part of the toolkit, which is this ability to, to approach scientific data analysis in this much more systematic, unbiased way.
Anyway, so we did an experiment. We got Claude, lovely Claude. Claude, Claude Opus in, in thinking mode.
4.1.
4.1.
So pretty recent.
And, and we gave it a dataset, open source dataset, and asked it to analyze this dataset, like find in... to find things that would be interesting to a materials scientist, um, who's, who's, like, trying to increase the efficiency of their catalysts, right?
Trying to understand which elements of the composition are, are more or less important, and, like, which-
Oh, God.
Yeah. Why, "Oh, God"? This is cool.
This is, uh, giving me flashbacks to chemistry, but yes, it is cool.
Yeah, indeed. Yeah, so, so Claude, lovely Claude. I'm sorry, Claude, but you did a terrible job. It hallucinated some stuff. It made some, like, big sweeping over, overgeneralizations. It, like, over-indexed a few outliers. Like, it's fine. It's not Claude's fault.
Like, Cl- Claude is, Claude is just not made for this. It's, it's okay. Like, Claude, we still love you. It's fine. And then we did, we did the exact same thing, but we gave Claude access to Discovery Engine.
Yeah.
And it-
Oh my God, it's so different
... it was amazing. Yeah, like, bec- because Claude... and language models, right? They're exceptional at synthesis, at, like, pulling-
Pulling things
... things together, at, you know... And they've got a really good sense of what's important to us, like, what's important to humans. Giving them access to, to data-driven discovery tools like ours really feels like a step change in terms of their ability to contribute to frontier science, you know?
Because you can't have them writing papers because they're just spectacularly good at saying things that are very plausible but not true, and if you're doing frontier science, like, definitionally, you, you don't have a ground truth. So it's very difficult to tell when the model is making stuff up or not, because, like, you don't know.
You have to go and, like, replicate the whole thing, which is very costly. I'm actually really worried about this because I think we're already seeing arXiv and other online repositories, and indeed-
Shutting down
... journal submissions too, full of these very, very plausible papers-
Yeah
... that may or may not be true. And then, like, at that point, what good is the, is our scientific literature if you don't know if it's true or not? It's a problem.
It's a test.
It's a te- for who? For the reader?
For us. Yeah.
Yeah. I don't know. This-- Like, I don't wish to be too pessimistic here, because I think language models are spectacular technology. I love them, I use them every day, but there are some things for which they are fundamentally ill-suited, and I really, I'm really not sure that scientific discovery, or, or, like, parts of it at least, were within the wheelhouse without specialized tools.
And I, I don't know. Like, you'll have a good opinion on this, Swyx, because you speak to a lot of people. How do you feel about tool use-
How do I feel about tool use?
... for large language models and agents?
I think, like, the newer models, like the GPT-5s of the world, are optimizing for tool use and kind of freezing everything else. Because I think, like, the understanding is that thinking with tools is, is what Ben Highlick calls it on, on the Lane Space blog, is better than pretty much any other form of context engineering you wanna do before it.
Yeah.
Uh, but I think more, more prominently, if you have, like, quote-unquote, "verifiable domains" in science, then you should give them a tool to do that.
Yeah, absolutely. This is exciting, right? I'm very excited about this, actually.
Audience & Demo17:55
You're bursting with joy. So what are your customers? Like, who, who do you work with? Like, uh, who's the audience for this?
Well, first, anybody with data. Like, we're very science-focused, but the tool is domain-agnostic. Like, it doesn't care if it's- ... agricultural data or cellular data or any other kind of data indeed. You can, you can use it. We're, we're launching a, um, basically f- free dashboard service web app platform.
Free for academics, or free for anybody who will, uh, publish their data, basically free for open source data. Um-
Mm-hmm
... yeah, so any- anybody will be able to use this.
Is this Disco?
For our, um, Discovery Engine, yeah.
Yeah, yeah. Okay. So I will show the page, 'cause you sent it over. I did a little bit of a p- trial run, so it's, uh-
Oh, you did?
Yeah. It's just, uh, you choose a dataset, you ch- you select some preset schemas, and then it runs, and the end result looks like this. Um-
This terrifyingly complicated dashboard. Let me, let me-
It's beautiful
... I'll tell you what you're looking at. Do you like it?
Yeah.
Yeah.
Yeah.
Yeah. So this is the Discovery Engine dashboard. This, this is... Yeah, this is the concrete compressive strength dataset, which is a pretty old dataset. It's open source. Been analyzed about a bajillion times, and as you can see in the bottom right-hand corner, you'll notice that we found, yeah, in that little table, we found a bunch of patterns.
We found looks like 17 patterns. None of them are novel. That's because this dataset has been analyzed to death. And it's only little to begin with. Two of them are speculative, 15 of them are validated and confirmatory. Validated means that we can point to a subset of the data that-
Mm-hmm
... evidences this pattern. So this isn't, like, speculative hypotheses. This is, here is a real pattern in the data, and we can prove it. And we can see the features, a big list of them, and how many patterns they're present in.
If so, it's on the left. Um, there's, so there's a bunch of stuff under the features, like we have mutual information and correlation. This is useful context, um, going into it. The thing that we care about is this feature here, CCS, in the top left.
That stands for concrete compressive strength, and that's the thing that we want. We really wanna go up, right? We want concrete that's really strong when it's compressed. That's the aim of the game.
Hmm.
Yeah. Um, and so maybe to start with, we can look at the correlation matrix, and we can look at the mutual information matrix on the table. I think, if I recall correctly, there's some correlations that's maybe useful to look at.
A correlation matrix?
Yeah. Yeah. So this is pretty trivial stuff, right? We can see these things are linearly, linearly correlated, and some of them more positively than others. Um, and similarly with the mutual information matrix, there's, there's, there's some stuff that you might be interested in.
It's not super-duper interesting, but the really interesting thing happens when you look at the patterns, which again, is on the left-hand menu. Yeah, there we go. Okay. So it's a lot of info. You added all graphs. Right. So now we can see the graphs of all of, all of the patterns that we found.
Um, and you can see that there's a bunch-
I see. This is, this is kind of EDA, but, uh, sort of guided and sort of pre-suggested.
It's not exploratory. It is systematic, my friend. These are all of the patterns.
Right.
Yeah.
Right, right, right.
So I think-
Okay
... um, pattern number two is quite good for, um, yeah, for demonstration purposes, 'cause it's a nice example of the kind of patterns that Discovery Engine is really great at finding, that humans are not so good at finding.
So what we're looking at here, on the far left, you can see the overall distribution of concrete compressive strength, and then the violins on the right... Yeah, that's the overall one. The violins on the right, you can see the distribution of concrete compressive strength under different conditions.
And so for each condition, you've got, like, some, some, some, like, range or some category that it's part of. Here we've got the ash to cement ratio needs to fall within particular bounds. The aggregate-cement ratio, the same. Likewise, the age.
What we can see that's really interesting, on the far right, you'll notice that combining these conditions, so, like, there's a combinatorial effect here. There's some synergistic effect here that is way more powerful when you take all of these conditions together than when you look at each one individually.
So correlation could give, could give you these. You might see, like, oh, well, you know, um, I see that concrete compressive strength is positively correlated with age, which is true, 'cause concrete tends to get stronger as it cures.
Or you could be like-
That's interesting
... ah, you know, I see that it's, you know, negatively correlated with the ash to cement ratio. What's really hard to find as a human is this combination of patterns, like this combination of conditions, that when you put them all together in aggregate, look at that.
You increased your concrete compressive strength on average by... I can't read that small. What does it say?
What kind of average? Mean or median?
Yeah. So this is the mean of 53, um, compared to the overall mean, which is-
35
... 35. Not so shabby. And then underneath, you can see we've usefully contextualized this for you. This is not a novel finding. People know about these things. Yeah.
Mm-hmm.
And we can evidence that, so here are the citations, et cetera. If we had found something new, you'd see citations to any literature that's relevant, any literature that's inf- informative, and you'd have a little exclamation mark on your pattern so that you could identify that you'd made a new discovery.
Mm.
Okay. Yeah, I think I should probably log in for that.
And, like, scientists, if you wanna use this, you can request a pilot, and soon it will be freely available and self-serve.
Beautiful.
Mm-hmm.
Roadmap24:18
Okay. Uh, so, you know, I think, I think we've, we've got a nice overview of the platform and the sort of different applications. What's next you're thinking? Like, what's, um, what's on the roadmap?
Oh, God. Okay. There's... Yeah. We're, we're very excited about lots of different things. So the fir- the first one is this notion of tool loose, use, like giving agents this kind of lens on data, we think is super power for all, you know, like science agents of which, you know, we're seeing them pop up a lot now.
So that's one thing. Second thing is more modalities. Like at, at, at its core, the, the... like all, all of these processes are not opinionated about data modality or about domain for that matter. So we're really excited about expanding to multimodal data sets like Vision-
Mm
... and Tabular together. Basically because it's, it's largely impossible to do good data analysis on multimodal data of this kind unless you have really fine-grained labeling of your images, for example, which is just very, very burdensome. But obviously deep learning, we can, we can let the model figure out its own features and then with the interp we can, we can pull them out.
So that's really exciting. We're, we're starting industry pilots, so if you're working in big R&D somewhere and, uh, you wanna massively accelerate your scientific discovery, you can get in touch and use the tool. We're about 100 times faster than manual analysis.
Well, if it's simulated, right?
Simulated?
I don't know. It's, uh... that's usually the, the answer. I mean, it's like partially, uh, you're 100 times faster because you simulate part of it.
No.
'Cause you don't have to go through the... Oh. Then what's the source of the 100?
The source of the 100 is that instead of doing this very like hypothesis driven, path dependent, I'll try one thing, I'll try another thing, I'll-
I see. It's a process improvement. Yeah, yeah.
Massively. Yeah, yeah. Like we're going from months of this iterative processing time to systematically extracting all of the valuable meaning in the data automatically in a couple hours.
Yeah.
Yeah.
That's crazy.
Yeah.
Outro26:28
Okay. Lovely. I think that's, that's a beautiful place to, to end.
All right.
If people want to reach out and, uh, find out more, where do they go?
Uh, they go to our website, leaplabs.com, leap-
Leaplabs.com
... hyphen labs.com. I regret-
Yeah
... that hyphen every day 'cause it's hard to say, right?
Well-
You can put the link, put the link in the show notes and save me having to say.
We'll put it in the show notes. Exactly.
Yeah. Cool.
Yeah. Beautiful. Well, Jessica, it's a, it's a pleasure to catch up with you. I think like there's a lot of applications here that I think like for the right people, they are really gonna light up and use this.
For me, I almost want to think about just general applications of these patterns spotting that you've, you've created. But obviously, you know, advancing science is, is a worth- worthwhile cause. Thanks for your time. This is, this is really great.
You're very welcome. Thanks for having me.





