Intro0:00
Light in space light twenty twenty-five. Recap. Light in space light twenty twenty-five.
Okay, we are here live at NeurIPS with two good folks. We want to cover the State of MechInterp, basically, and, uh, you guys are very passionate. Mark, I had you on for AIE before. Welcome back.
Thank you.
Uh, and Jack, you're new, but, uh, also part of Goodfire. Yeah. Uh, how, how would you describe, uh, what do you guys do and your path into MechInterp? Maybe Jack, you wanna go first? So we do interpretability research primarily.
Uh, we're a company focused on, like, making models interpretable and, and robust and safe. And I guess I would describe, like, the type of work we do as, like, the science of deep learning, which is trying to make models, not just black boxes, like, that we-- things that we can actually trust and, like, deploy in, in high-stakes, uh, industries.
And I guess my own path into interpretability, um, was I was a PhD student, um, from twenty twenty to twenty twenty-five. I graduated in May. Uh, and I started working on... I was working on language models. I was working on grounding in language models, which is basically the idea that you need more than text data to represent meaning in the world.
And it was basically right away as I started, like, at some point, you know, GPT-3 had come out, and I had a moment of like: Oh, this is actually like, you know, just really good at, like, understanding the world and like...
It was kind of a slow transition, but somewhere along the way in grad school, I switched to fully doing interpretability. Cool.
Yeah. So like Jack mentioned, Goodfire, um, is an AI research company, uh, focused on building a platform for interpreting models of all kinds across lots of different modalities and domains. My path looks very different, uh, than Jack. So prior to Goodfire, I was at Palantir as an engineer on our healthcare team, and then, uh, joined Goodfire back in March.
And I guess one cool thing about kind of the, the state of interpretability as a field and also how Goodfire is set up, um, is there's a lot of foundational research still to be done that folks like Jack are, are working on in our team.
But at Goodfire, I'm more focused on sort of like our applied real-world use cases for, for interps or building out a platform that can help in use cases like scientific discovery or for kind of like inference time monitoring of models, um, for deployment in enterprises, uh, and things like that.
So I think there's a lot of exciting, just totally new theoretical research to be done. But, uh, especially over the past year, I think we're starting to see interpretability have practical use cases be actually deployed in like especially situations where models are being used for high-stakes industries, um, and problems, which is, which is super exciting.
Paint Demo2:56
Yeah. I, I think a lot of people are ignorant of the fact that we're actually at this point where people could actually apply these things for, for real-life use cases. I saw the platform directly. I-- You guys had, like, a launch party in your office for the diffusion thing.
Maybe you wanna recap what that was so the people can go play around with it again.
Yeah, yeah. So this is, uh, kind of a, a research preview that we put out. Um, it's at paint.goodfire.ai. It's live. That was a, a use case of interpretability for, um, sort of like the creative domain. I think just giving one hint at, like, interpretability gives you a set of-- I think of it almost as like power user tools for accessing models and doing things with them that you might not have realized you could.
Uh, so that example is you take Stable Diffusion XL Turbo, so a model that typically you use text input to prompt the model, and you get an image out. But using interpretability techniques, you can sort of like plug directly into the mind of the model, and you get a 2D canvas where you can basically, like, paint directly into its mental map of the image.
And so we used unsupervised techniques to basically figure out all the concepts that the model, like, has internally, so animals, backgrounds, um, scenes and stuff like that. And so you can select any of those concepts, which again, were found totally unsupervised, and you can just, like, paint a lion over here and then drag it and move it over here.
It's just a totally new way. Like, this is a model that only takes text prompts as input, but when you plug into the brain, you can do these cool things.
Yeah. And then highlighting other work because I think people-- That one went relatively viral. Uh, Jack, I don't know if you wanna take a turn at, like, other highlights of your year in terms of stuff that you shared.
You start with the memorization. Okay. Sure. Yeah. So there's a lot of work and, like, going back years on, like... So models memorize a lot of their training data. That's a privacy concern, but it's also, I think, just scientifically under- interesting to understand.
Memorization Spectrum4:37
And there's sort of like an unclear picture of, like, how should we think about how it's represented, like, within a model or when it's... Is it, is it like a, a computer with like, like a file system that we can kind of tap into where maybe files are redundantly stored, like, you know, memorized training sequence are kind of spread throughout?
Or should we think about it, like, some other way? And we kind of proposed a, a kinda like different lens on it. I think a, a exciting question around the work was, um, whether we can, like, disentangle, like, core cognitive capacities in a language model from- From knowledge, right?
From knowledge. I always think, uh, like GPQA zero- Mm-hmm ... and, uh, you know, AFV one hundred. Yeah. Exactly. Right. Right. So this is a paper we put out recently in Oc-October, um, and, uh, we showed, um, there's this like nice spectrum of, like, capabilities.
And so there's another-- I think another new lens on, on memorization, which is that in language models, it's not so black and white, like what is memorized and not. So, like, memorizing that the, the capital of France is Paris is, you know, memorized in some sense, but it's, it's a lot different than just, like, memorizing like a single page, like license agreement document that shows up a thousand times in the pre-training data or something like that.
You can actually see like a-- the, the way that we like disentangle memorization, you can kind of see this, uh, like gradient of memorization, uh, in between both mechanistically and, and behaviorally with like logical reasoning tasks being quite distinct from rote memorization and then like factual recall is kind of somewhere in the middle.
What's your, uh, follow-on work af-after you've done this?
So I think like understanding, um, like the reasoning capabilities is, is very interesting. I think like understanding how post-training affects those capabilities in general is super interesting, and I think we're gonna keep pushing on that.
Can you imbue forgetting so actively?
We show... And so yeah, so there's, there's a lot of work on like unlearning. Uh, I would, I would describe it more as not unlearning, but maybe suppression. Um, I think there's like really like, I guess guarantees that you've fully removed information from a model is, is I don't think it's been convincingly showed, uh, anywhere yet.
And I, I think that it's-- I think maybe some of the point that we wanna make in our work is that when we look, when we look at memorization this way, it, it makes a, a point about how hard or, or, you know, possibly like intractable that is.
I don't know if this is a related problem or it's two different, but instead of unlearning but updating.
Mm-hmm.
So I move the capital of France to Marseille.
Yeah.
But there's so much training data saying the capital of France is Paris.
Yeah.
I need like a date. I need to be able to tell the model, "Hey, like this is now out of date." But like short of tagging a date on everything, I don't know when that makes sense. You know, that's like a weird-
Yeah, yeah. A really big, uh, like paper from a couple of years ago on this exact thing on fact editing was the, the Rome paper. It's Rank One Model Edit. And, um-
Nice acronym.
Yeah, yeah. Here we go. Uh, and they look at exactly like, exactly what you're describing. Um, it's a really nice, nice paper. I don't think it's, it's being used like in deployment anywhere, but I think that's like... And it's, it's, it's really an interpretability paper and, um, I think that's somewhere where like, you know, from a basic science perspective, like interpretability has like so much to, to offer is understanding how we can do something like that, uh, which is currently just very difficult.
Production Use8:13
So call you over to Mark a little bit for industry stuff. What can people do at the end of twenty-twenty five that they could not do at the start of twenty-twenty five?
So I'll, I'll preface with saying we, we still have a long way to go, but, but we are excited about seeing like the seeds of, you know, things actually being deployed in practice. I, I would say that we're just continuing to make progress on understanding that models have a lot of latent capacity that you can't get at by treating them entirely as, as black boxes.
So we're seeing interpretability show up in like model cards now for, um, evals that are pe- people are running, um, for various like red teaming exercises and stuff. Yeah, I know, uh, there's some stuff in Gemini 3, um, Claude, I think Claude 4, and it's definitely Claude-
There's a big topic to beat.
Yeah, for sure. I mean, yeah, the interp team there is, is phenomenal and I think works across some of their other teams. But-- And then from Goodfire's perspective, you know, we-- so like one of our partners, Rakuten, is deploying an interpretability-based tool in production with one of their language agents.
Uh, this is a really cool use case where if you-- what they needed to do was take chats between, uh, their customers and their agent, find instances where personally identifiable information is mentioned, so names, emails, phone numbers, things like that, and scrub it out.
And interpretability turned out to be the, the best way to do this at scale and in a cost-effective way. It was both more effective and cheaper than alternative techniques.
And you do that by having like a PII, uh, feature?
Uh, yeah, exactly. So what you do is you-- the, uh, customer is talking with an agent, but then separately, you're putting that transcript through what we call like a sidecar model.
Okay.
And you could-- What's really interesting is if you ask that model, try to like use it as an LLM as a judge, it's not very good. But if you probe its mind and you sort of detect when the features related to personally identifiable information are firing, that gets you the highest recall of anything.
It's, it's the equivalent of using GPT-5 as a judge, but it's, you know, like five hundred times cheaper. So, so they're, they're deploying that. And then I, you know, am personally very excited about the use cases for scientific discovery.
So some of our partners in the life sciences and in materials, uh, there, there are these narrowly superhuman models for things like genomics, for medical imaging, for proteomics, for material science, and those are, um, especially uninterpretable because they are working in domains that, yeah, we can't-- Well, they're, they're huge, and like we can't native-- I don't speak genome.
You know, like it's literally base pairs in, base pairs out, but they're superhuman at ver-- tasks that are very interesting. So we have some, some, you know, exciting early results in, in finding novel biomarkers of disease with some of our partners that, um, we're excited to share soon.
And yeah, I mean, I think AI for science is, you know, becoming a very hot topic, and I think for good reason.
Yeah. We're, we're literally starting AI for Science pod. Like we're spinning out a, a separate in space AI for Science pod.
I love to hear that. Yeah.
Circuit Tracing11:17
Yeah. Featuring other work, uh, done notable work this year, I feel like I have to mention Entopic's circuit tracing paper. I don't know if there's much discussion internally for you guys. I'll let you riff. Like what do you think?
Yeah. Jack, I know, has done a lot of circuit tracing. You wanna-
Yeah. When that-- So when that came out, we, we did a-- So that was, um, back in like-
March
... March. Yeah, March or April.
If you could, uh, let's say people know about the SAE work.
Yeah.
What is the difference? Because I, I struggle summarizing it. My, my summary, you, you can correct me, it's like take an indivi- like have a full access to a model. Take an individual layer and train a replacement circuit that simplifies what it, uh, does.
Yeah. So yeah, that's right. And a good way to maybe put it is that, um, you take a representation in the middle, which is this like weird, dense, uninterpretable like web of concepts and, you know, features in a representation that- The SAE like decomposes it into like primitives, like, um, concept spying for like mentions of coffee shops or mentions of like, uh, New York City or something like that.
And those are, those are mu-much more interpretable and that it basically decompo- it shatters the representation to like many, many pieces. The, uh, big change with the, the cross-layer transcoder, so the circuit tracing work, was, A, to, to really scale up the models to incorporate like every layer.
Uh, so like cross-layer, uh... So the, the model's called a cross-layer transcoder, so it's incorporating like-- it's, um, tying features across different layers. Uh, and then the, the tracing part is a, a method for creating like an, it's called an attribution graph through those features which are interpretable to like describe how the model is like producing one output, like through every layer and through like a bunch of token positions.
I could talk a bit more about it. So when, when that came out in like March or April, we, um... So it's still pretty early for us. There were like eight to 10 of us at the time, I think.
So we, we tried to-- we like, uh, went on- had to replicate some of their findings. We were really curious about, you know, what would look like training one, like, uh, what it's like to use it, and basic scientific questions as well, which is like could it rediscover, um, some like rich representations from like previously understood circuits?
So we, we put out a post, I don't remember when exactly that went up, but over- I think over the summer, like, uh, May or June, on, um, our replication effort on that and like, um, how that works.
Is there, is there an obvious next step? Like, what is, what is this all leading up to? Because to me it's like basically just always scaling it up, always making it more unsupervised. Uh, I don't know if there are other trends that you can see like, yeah, these are the core principles that we're just exploring.
That's like a good point to bring up interpretability as, as an alignment science versus as like a science for understanding models like more broadly, which is like if you-
They're one and the same kind of
Oh, I think it depends on what your motivations are. Um, so if you are focused on reading a model's mind, um, and understanding like maybe what's going on internally to make sure it's not having bad thoughts or like it's not misaligned, it makes a lot-
Bad thoughts.
Yeah. It, uh, it makes a lot of sense to, to like do exactly what they're doing. So I think-- and that there-- I think they're just not there... Uh, I don't know exactly what they're up to, but, uh, applying like these techniques to like read out what's going on inside these models' mind, so it's like very nice detection like framework.
But, um, if we want like really robust control, uh, like really robu-robust control of models, like what they learn during training, I think like that's where there's so much more work to be done. I think, you know, uh, many different teams are all working on, on these problems.
But in terms of next steps, that's, uh, yeah, like some other directions to go.
Yeah, no, I think to, to Jack's point, I think the, the circuit tracing work is super cool, and I think it's, um, it's sort of one arrow in a, in a larger quiver of techniques that are useful. And, you know, w-what technique is useful depends on the task at hand that you're interested in.
So to Jack's point, like the techniques that are useful for alignment science and evaluations, you know, that's one use case, but there's a whole world of other things that you might wish to apply other techniques for, and maybe, maybe you don't need circuits for things that can be uncovered using probes or using other techniques to understand how, you know, post-training changed a model.
So this like model diffing use case and maybe, uh, the things that you wanna use as like inference time sort of guarantees for, for models look a little bit different. So I think it's super cool, but yeah, to your point, kinda like there's, there's a variety of, um, a variety of techniques depending on the, the final use case that you're interested in.
Pragmatic Interp15:44
Last thing that a lot of people here in NeurIPS, I've had like two conversations about it already, is what's happening with Neil Nanda? I don't like celebrity culture, but i-it's hard to avoid Neil's impact, and he basically said he's pivoting his team at DeepMind, uh, to no longer focus on whatever, and now it's like pragmatic interpretability.
What's that conversation about? Like, what, what do, what do insiders think?
You know, I, I read, uh, Neil's post and I think like I also share the, the hope that of, of a pragmatic future for, for interpretability, for sure. Um, I think a lot of the like-
Why didn't we think of that earlier? We should think pragmatic. Like, oh, God.
I, well, I also think, I also think a lot of the, um, the response to that, I'm not sure that folks actually kind of read it all the way, all the way through. I think there's a lot of reading of that, for some reason folks thinking, oh, like interpretability is dead or something.
If anything, that says interpretability is alive and well. Like there are use cases for not blackbox te- black box techniques that, that can be brought to bear in real world use cases. So I, you know, I, I think we're, um, we, we probably agree with Neil on that.
And then, um, you know, I think also good for like different companies to have different agendas that they're pursuing. Uh, I think it's dangerous, you know, a field will stagnate if everyone sort of is just converging on the exact same, same approaches of things.
Not sure if you have other-
To the point about like is, um, interpretability like dead or something, I think is like a, a gro-like gross, like misattribution of like what that post is about.
It's about your point.
Yeah.
Nobody's... I'm not saying it. Nobody, nobody I talked about saying it.
Oh, okay.
It's, it's more like the existing approaches for him were not scaling to some extent.
Ah.
And he was-- The way I put it is kind of like, it's like, like, okay, let's forget about complete understanding and let's manage by outcome, and like let's measure ourselves by our ability to steer outcomes and forget if we know precisely what's going on.
Yeah, yeah. I totally... So in terms of like should we just be focusing on like reverse engineering models, like from the ground up, so that's maybe where the big change is. I think like we, we probably agree that like interpretability should be useful and we're trying to get it like, like used right now, and we are getting it used right now.
I think like maybe where, like where we kind of align on this is like we're really focused on like use case inspired, like basic research. So yes, like we're very pragmatic, but there's also like a, like a lot of like very deep foundational science to be done on understanding models, which is, um- Yes, driven by outcomes, but we also want to, like, develop, like, the understanding of, of what goes on in models.
Because if we're not doing that, then every other, uh, like, lab producing models is just a black box- ... AI lab and, you know, that's not how things should be done.
Jack snuck in a, a little, um, reference there that I think was, was good. So there's this concept that we like to talk about internally called Pasteur's Quadrant. So Louis Pasteur, like-
Is this like a statistical thing with like the, the same distribution but four different-
Oh.
That's Anscombe. Yeah.
Yeah. No, this is more of like a conceptual kind of framework, but the idea is that you can have ... So the two axes of the quadrant are-- I've heard it described as, like, discovery and invention or, like, pure research versus applied research, like something in, in that kinda domain.
But the idea is that you can have just pure basic research. And so the classic example there is, like, Niels Bohr, like understanding the atom, understanding the electron just for sort of like theoretical physics, uh, interest. And then you have the, uh, corner that is just, like, purely applied research.
Purely you have an end in insight, which is like Edisonian research. So Thomas Edison says, "I wanna make a light bulb. I don't really ... Like, I'll learn the chemistry that I need to, but it's just this goal is, is the goal."
And then there's Louis Pasteur who, like, spans both of these. And so the idea is that it's not just this linear thing between basic research and applied research. You can kinda have a combination of those two. And so Pasteur, you know, sort of pioneer of like germ theory, but also engineered some of the first vaccines.
And so this idea that you should sort of be bouncing back and forth between kinda open-ended foundational research, but then also having a goal in sight and be able to really flip back and forth between the two. I think there's a lot of cases where you see this being a really productive way to make progress in a field.
So he's a company, like, hero mascot.
I think it's Tom McGrath, uh, one of the co-fire- uh, co-founders of Goodfire who, um, he started the Interp team, uh, and DeepMind back in the day, and he, uh, he has a lot of these, like, great references from other domains of science that are just, uh, I think really good kind of grounding points for, for how we operate.
Hiring20:26
Amazing. Let's get a quick call to action in. Uh, you guys are hiring, what are you hiring for? What's, what's hard to find?
Definitely actively hiring. Uh, I hope the AI engineer community, um, should, should definitely go to, um, goodfire.ai, check out our careers page. Actively hiring for researchers and, uh, engineers and-
And human beings.
Yeah. Yeah. Exactly. Exactly. There's a combination. Um, uh, so, so yeah, we're, you know, we're doing foundational interpretability research, but we're building out a platform to apply this in real world use cases. You know, customers across life sciences, materials, government, um, financial services.
So if you're, uh, someone with an MLE background, no interp experience required whatsoever, you can see we have different backgrounds. I'm kinda coming from industry, from engineering. Jack's coming from a PhD. Other folks at Goodfire have backgrounds in like quant trading firms or frontier labs, um, all sorts of places.
But if you like training big models, uh, building agents, engineering systems, um, we're hiring across, uh, a variety of roles and really looking to fill some of those engineering gaps.
Excellent. I think that's it.
Yeah. Love it. Thanks for having us on.
Thanks for coming.
Yeah.


