Intro & Release0:00
All right. We are actually going to record this as a intro to the main episode, but here we have my trusty co-host, guest host, I guess, Vibhu, um, as well as Emmanuel from Anthropic. We're gonna talk about the circuit tracing stuff and all the interpretability work.
But Emmanuel, maybe you wanna do a quick self-intro b- because before we get into it.
Yeah, sure. I'm Emmanuel. I work on the, uh, interpretability team here at Anthropic, more specifically on the circuits team. So we recently released a pair of papers about sort of like the work that we've been doing over the last months.
And even more recently, we released some, some code in partnership with, with the Anthropic Fellows Program. It was, it was mostly built by Anthropic fellows that lets people play with the research, basically. And so happy to talk about that and, and we also hope to kinda like keep releasing more things and, and, and partner with other groups that are working on, on similar stuff.
Yeah, amazing. Uh, we'll get deeper into like the behind the scenes on, on the main podcast, but let's maybe just dive right in into what you released because that's the most topical thing. This is like literally just launched it like yesterday and we'll probably release it in, uh, release this episode in a few days.
So yeah, like what can people do or what do you recommend people try?
Totally. So like at a really high level, you know, the, the sorta like idea of the research itself is to try to explain sort of like some of the computation that a model did when, when it predicted a given token.
And so in our paper we kinda like show how to do this and then we show examples of us doing this on internal private models and then the, the release this week sort of lets anyone do it for a set of open source models.
So notably maybe the most easy one here is like Gemma 2-2B. So you can sort of like think of, of some prompt and you kind of like can explain any, any like token that the model sample samples. And explains here means just like basically blow up the internal state of the model and like show all of the sort of like intermediate things that the model was thinking about before it got to like the final token that it predicted.
Yeah. So what some of the things that you guys put out is kind of in the circuit tracing you have a few core examples, right? So like we can see how these models have internal reasoning states and there's like multi-hop reasoning and some of the stuff that we talked about on the podcast was how can people that are interested in how models work kind of do anything, right?
So what are open questions? How can people contribute? And it seems like, you know, the follow-up is, okay, it's been a few weeks. Here's a huge library. So you know, I guess before we even get into it, what are some open questions that you would expect people to like kinda play around with?
You know, what, what are people like gonna do? What, why should we probe Gemma, Llama? What are interesting things we can do and any tips on using it?
Yeah. I, I think there's maybe like two to three like categories of things that people could do. So I'll go from sort of like the most basic kind of low effort to, you know, hey if you wanna dedicate like a month of your life you could do that.
Um, the, the sort of like most basic thing is, you know, um, so it's a Gemma 2 and Llama 1b are like smaller models, uh, but they can still do a bunch of stuff and so and, and for most of the things that they can do we still kind of like don't really know or have a good mental model of how it is that they do the things that they do.
So to give you an example, like one of the things in the paper is this sort of like multi-hop reasoning where you know we ask you know like uh Claude 3.5 Haiku like oh the capital of the state where Dallas is is Austin.
It turns out that like Gemma can do this also and so as part of the release we have you know a notebook where Michael Hanna one of the Anthropic fellows kinda like walks through a bunch of examples including this one and it's really cool 'cause you can see that actually the way the circuit looks in Gemma like a really small model is extremely similar to the way that it looks in like a huge model which that in its uh, in itself is, is I think like a pretty novel discovery.
It's like oh even these models are like super different you know if you look at like their evals or if you just try to use them they're like just very clearly different but for this one task for this one thing actually the way that this do th- that, that they do this multi-step reasoning is like the same way they actually do the multi-step reasoning.
In the notebook there's both other examples of kind of like fun things that we looked at that I think can sort of spike your interest if you're new to thinking about this stuff and at the end of the notebook that's, that's linked in the in the README there's like three examples of like random sort of like cases that we haven't solved or, or, or we haven't labeled that you know have like a graph pre-computed for you and you can just look at it and try to like figure out what's happening.
And by figure out what's happening what we mean is you know we might do like a quick demo here but it's kinda like look at these representations try to understand like okay like what is the computation the model is doing and then part of the release also lets you like run, run experiments to like verify that you're right.
So if you think that like you know ah the model like first like thinks about Texas in this case you can also just like stop it from thinking about Texas and see if like that di- damages it and so like the tools to do that are available and so I would say that's like the first thing is just I think the hope is there are a lot of behaviors that models do way more than like any single group has time to explore and so the hope is like hey pick a behavior that you think is interesting and try to understand like what's happening and try to ground it out and it's like the, the sort of like baseline thing and maybe like the f- the thing that I'm most excited about with this release but
then the other thing I, I do wanna mention to like parts two and three are just we also hope that like other groups and, and kinda like interested researchers can just use this to like extend the method like if you have an idea about how to like do this better you know the whole code to like make this graph is, is open source so you can like take a look at it and just like try to play with it try to find different ways to like create these graphs and also extend it to other models right?
Like there are many different models and so you know part of, part of making this work on any models you have to like train the sort of like replacement model which again there, there is code for it and there's other groups working on and so like that's also something that if you're excited about you could say like okay cool well I want this to work on like another open model and you could sort of like add it if you're like you know more interested in like maybe like the engineering and the ML engineering side of things.
Uh yeah we, we actually get into a little bit about how you how you guys do the extra data viz stuff that makes your your blog posts pop so much. Should we, should we um share the screen a little bit and, and dive in?
I, I think uh you guys prepped some examples.
Totally. Yeah, yeah
It's just like there's nothing better than the creator of the tool walking through the tool, and we might as well capture that so that, you know, people who actually want to do this can follow along.
Tools & Demos6:11
Yeah. Uh, that makes total sense. Let me just, like, actually share my screen.
My, uh, one little experiment, I basically cloned the repo, threw it into Cloud Code, and was like, you know, "Deal with this. Let's, let's try it end to end." So would recommend, you know, Cloud Code is very good at using this.
Um, basically also if you're, if you're just trying to get started, the Circuit Tracing Tutorial notebook, very good. That kind of goes over all the high level and then, you know, shout out Cloud Code, try it out. It, it works very well on this.
That, that's awesome to hear. Actually, I might just open the notebook first, just like quickly walk through the illustrations. But yeah, you're the second person to tell me that they just had Cloud Code sort of like dig in initially on it.
So I'm glad, I'm glad that's, that's working. The, the tur- tutorial here is like linked at the top of the repo, and maybe we can link it fr- from the podcast, but essentially sort of like walks you through kind of like how to think about graphs.
And so it links to these circuits. So here, this is the two-step reasoning that we were talking about. This is kind of like a schematic of it, where it's like the capital of the state of Dallas, and it's like, ah, it has to think of Texas and Austin.
But the notebook links you to all of these circuits here, and this is kind of the thing that you can play with. So this is the UI on, you know, Neuronpedia that, that hosts this and that, that lets you, like, create a new circuit.
So here, you know, we could explore this circuit, and if you open the notebook, you could explore it.
I'm realizing that I switched tabs. Maybe I'm not sharing it. Okay, there we go. Can you see the circuit now? I think so. Okay, cool. Um, but you can make a new graph super easily and quickly. And so maybe this is like the most fun thing.
When I was playing with, uh, right before, uh, you know, joining this call is like, uh, it turns out that podcast guests are very formulaic, and so if you say like, "Thanks for having me on the whatever," so like Gemma seems to have pretty consistently guessed that you're like on a podcast, uh, which makes sense, right?
Like, why would you say, "Thanks for having me on the blah"? Uh, and so here we can try to say like, oh, okay, like how does Gemma know to like complete the sentence with, "Thanks for having me on the Latent Space podcast"?
And so here the way you generate a graph, right, is you type a sentence where the next word is the thing that you're interested in, and then you kind of try to explain like how the model got to the next word.
So here you can
give it a name, and then you can mostly just not worry about any of these parameters, I think, if you're just playing with it, and you can click Start Generation.
This-
And all this generates, like something important for people to know is that these are trained on base models, right? So they're not chat models. So basically when you train these models, they're just trained to predict next token, and they don't have that user assistant chatbot flow.
So they're prompted in a way such that, you know, the output should basically just be the next word.
Yeah. You kind of want to think about it as like maybe the, the, the prompt or the text you're giving is like the text of like a book or an article rather than a conversation where it's like, you know, what is a sentence where if you were to read it in a book, like the next word would be sort of like the interesting one.
Um, yeah.
You know, you can click on it. Uh, sort of like takes a little bit of time to load because there's just a bunch of data. So what we're gonna show you here is basically like almost every single feature that activates in the model, the features are these intermediate representations, and at the bottom there's the prompt.
So here it's like, "Thanks for having me on the Latent Space." And at the top you can see like what the model sort of like, uh, output. So it's most likely output is it's pretty confident that we're talking about a podcast and, you know, it has like some random stop tokens, blog, show, and then some stuff that I think like makes less sense, but also like these are small models and so sometimes they say random stuff.
Um, and so the way that you could then explore this would be like, okay, so the model says podcast. So like why does it say podcast? Uh, so you can click on this output and say like what are the features, again, these like intermediate representations that have an input to this.
So it seems like there's features at, here this is the layer, like layer eighteen that already like are about podcast episodes. You can know this because the features have a label, but also if you want, you can look at the feature itself here, and here you can see that like this shows you like other text of the feature is active o- over and it's just like text about podcasts.
So that's like a way that you can also like understand what the features are. Um, and then you can keep going back. So it's like, oh, okay, so it said podcast here because of this podcast feature. Where, where did that come from?
And it's like, oh, it comes from like words related to podcasts, words associated with podcasts, as well as like an interview feature, um, and also just the word on. So there's like a bias, like if you're saying like, "Blah, blah, having me on," that sort of like slightly increases the chance that you're talking about a podcast at all.
Um, and you can sort of like keep going back, um, and, and kind of like explore the graph interactively. I would say that like the way to do it, and we talked about this on the like kind of like longer version of the podcast, but it's like, you know, kind of like chasing from the interesting outputs back or from the interesting input forward.
There are many nodes on these. I like wouldn't recommend looking at all of them. You can also sort of like prune them a little more aggressively here if this is too busy and kind of look-- Th- this shows you like only the most important ones, and you can sort of like be pretty extreme with it if you want.
Or you can show the whole thing and then be like super overwhelmed. Once, once you kind of do this, you can then kind of like group your nodes into similar ones to kind of like make a graph. I actually made this little summary earlier, so I can just share that tab.
So this is the exact same graph, but just before the, before, before hopping on, I kind of like did, did a few groups. So this is like same thing, podcast, and it's like, oh, there's like a bunch of nodes that are like podcast episodes.
There's a bunch of things like discussing podcast. There's a note about expressing gratitude that amplifies that you're like on an interview or a podcast. Like one fun experiment you could do here, right, is like, oh, like what happens if I mess with this?
Like if, if I don't-- like if I mess with the like, oh, this person is like grateful to be on and just is like this person is on, does it think you're on something else? Like maybe there are things that, you know, you could be on that you're not grateful for, like, "Oh, you're having me on trial," or something.
I don't know. Like that could be one, one interesting sort of like experiment to see what the causal effect of this is. And again, you could sort of like label it more and explore it more. And this UI, the whole point is for it to be like snappy and quick, so you can just like generate a bunch of graphs pretty easily, right?
Like maybe this wasn't exactly what you wanted, so you're like, "I'm super unhappy to be on the latent space." And then you can see what it completes for that or whatever, and you can sort of like just continuously play with it and, and get a better sense for your hypotheses.
Oftentimes you kind of like want different prompts, you know, different examples that are similar to kind of get a sense for it. And then if you're really curious and you want to dig in more, that's when I would recommend going back to like the code base and some of the notebooks.
User Experiments13:01
Maybe one thing-- one last thing I'll say on, on that is that the notebooks themselves, they can all be run on Google Colab. And all of the code, as far as we can tell, we've like tested all the notebooks, just like runs on Colab, and so that means that like you don't need, on the free tier to be clear, like you don't need like an expensive GPU.
You can just run this and kind of like run your interventions and play with it. So in this notebook in particular, the intro one, we show you how to do these interventions, and here we're like, "What happens if we turn this node off?
And what happens if we turn that one off? And what happens if we turn this one off? And what happens if, you know, we inject one from one prompt into another one?" And so I think that's the sort of like deeper dive, trying to understand the mechanism better.
But if you're just trying to even like get a, a sense at all of like how does the model does-- do X, you can just generate a graph and, and take a look at it.
Incredible.
Very cool.
Is there, uh... When I look at the graph, is there's a thought in my mind about maybe this is too easy, too perfect. Um, and one version of this is there's supposed to be superposition, and here there's no superposition kind of.
Well, there is superposition, and like we're, we're sort of like-- So maybe I can share, I can share the, the like graph again and like answer your question, which I think is like, what are we hiding here? Where are the skeletons?
Yeah. It's-- This is like, it's too per- it's too clean. I'm like-
So-
Yeah.
So okay. So, so maybe like a good example is, and we're gonna like make this slightly less overwhelming here, is like, okay, so you, so you look at this graph and you say like, yeah, like we don't actually understand how models work fully, so like what are you hiding here?
And, and the thing that's like important to note here is, you know, I, I didn't say this explicitly, but like the layers are like arranged here, and so let's just look at like one layer. So for this layer, what we're saying is like the only thing that, that is happening or that's like important enough is this one feature, which is just like one small direction in the model space, right?
Like one, one dimension we've pulled out of superposition, uh, or, or, um, let's say that for now. But then also there's these diamonds, and these diamonds are errors. We talked about them on the, on the longer podcast, but they're just like when you train these replacement models to replace some of the model computation, you successfully replace some of it, and then some of it you fail to replace.
And so this is like everything that we don't understand. And so that means that like sometimes if you look at an input, like this guy's input, you'll see a bunch of errors here as the input. And so essentially, you know, there's some graphs and some examples where like if you have, if most of the stuff that you see is these, these errors, basically that's, that just means like, hey, for this prompt, we were not able to sort of like explain, you know, that part of the computation.
And so at least that part is like a, an explicit sort of like we show it in your face, where it's like, here's, here's what we don't understand, and so, so you can sort of like see-
Okay
... what we don't understand. There's also, I will say like one more thing. There's like a bunch more stuff that can get you, uh, and that's like in the paper. But like one example here that I'll just say is like these are just MLPs, so the model has both attention heads and, uh, multi-layer perceptions, MLPs.
We don't just do it. Like, we completely ignore attention. Or like, we don't, we don't try to decompose it at all. So there are some prompts where like all of the interesting stuff is attention, and here you're just not, you're just not seeing it at all.
The way that it's materialized is like you have an edge from here to here, and like some attention head did a bunch of stuff. You don't know what it is. And so that's also the part that we're sort of like not explaining.
So there's, there's definitely, yeah. Th- I don't want to make the claim that we explain everything. I think the correct way to think about this is like if you look at a prompt and you can, by tracing through these, not hit any errors, hit nodes that make sense, and build up a reasonable hypothesis, and then when you test it with interventions, it works, you've at least understood some, and presumably like a reasonable proportion of the computation.
If your interventions are working, that means that it's like the thing you found is not just like a side thing. It's part of the main thing that the model is doing. And so, you know, then the question is like how often does that happen versus how often do you just hit these errors or you're like confused.
And I think that's, that's just sort of like what works and what doesn't, uh, summary here.
Yeah. Crazy. I mean, uh, congrats on this work. I, I know you f- you're low on sleep 'cause you worked really hard on, uh, shipping it, and you're a perfectionist. I, I just think like-
Well-
Yeah. Go ahead.
Sorry. I'll just say that like the actual brunt of the work here is like-
Your fellows
... the Anthropic fellows.
Yeah.
Yeah, yeah.
Yeah.
They... You know, I, I mostly just like- ... coordinated things left and right. But, but they sort of like did all of the implementation as well as, you know, folks on the like Neuronpedia decode research side also did, you know, the, the lion's share of the work here to actually have the, the front-end UI.
I'll just say, like, you know, uh, Vibhu and I were at the Good Fire meetup yesterday, where there were a lot of, uh, interpretability folks. I was shocked at, um, honestly, how young most interpretability people and work are.
Yeah.
And it-- this is a very young field, e-exactly like you say in the podcast. There's a lot of fresh green grass here to, to tread, and it's just really inspiring. Vibhu, do you have any other, like, final thoughts or comments?
Yeah, no. I think there's just a lot of open work to be done, you know, and we talk about this in the podcast too. And just to reiterate like how good the tooling that you guys put out is, like even the fact that, you know, without diving into any code, you can enter in a prompt and start to play through these circuits in like minutes, it's, it's pretty incredible.
Like I, I could share another one, actually. So I was doing this with Pomsky, and I finally got it to work. So our guest host of the episode is Mochi, my little dog. She's our distilled husky. So she's on the podcast later and, you know, I basically put in like I, I had to guide it quite a bit, but my, my prompt is: A Pomsky is a small dog that's a breed of a misc- of a husky and a...
And then, you know, I'm expecting it to put out Pomeranian or Pom. Uh, let me, let me share my screen real quick. And then we can kind of dig through. This is me, like-
By the way, her tagline... Yeah, while, while you put it up, her tagline is officially Mochi the Interpretability Husky.
For today.
For today.
Yes.
We're gonna change our tagline every episode, but yeah.
It feels a little weird, you know? We're digging deep into what Mochi is. But basically this is me, like, no background, like two minutes in, just put in a phrase, and now I get to play around with features, right?
So this is also called Please with four S's because, you know, I tried a few prompts. It's okay, it's okay. We struggle. It only took a few minutes though, so you know. Pomsky is a small dog breed that a, that's a mix of a husky and a...
And then the, the most out- uh, probable outputs, you know, now it says pom. So okay, let's dig into what some of these are. I'm basically just going, like, fresh, haven't done this before, but you know, words related to animals, their emotions, their health.
We have a feature for dog, golden lab, mentioned dog breeds, especially high maintenance. You know, this is basically like AGI. It knows Pomskies are high maintenance. It's, it's figured it out. But realistically, you know, um, as I dig through these features, I can start to pin them, uh, layer them through.
Mentions of garbage and waste. No, that's not nice. That's not nice.
Yeah, interesting.
But basically, you know... And this is already me pruning out most of the features. Uh, as I open it up, you know, it talks about different things like dog breeding. What else? Related to animal welfare, so like... And then you can dig through all this.
There's just so many things that, like, you know, this is in a matter of minutes. I basically made a graph, put in a sentence, and now I have an output, and I can traverse through what are different things.
Okay, animal science, right? So this breed is relatively new. It's not that common that big Huskies and little Pomeranians naturally have offspring, but you know, let's, let's like dig through animal science versions of this, and then we have, like, interesting little features.
So it's, it's very easy for people to kind of get a different understanding of what goes on throughout layers in models, you know? But that's just my fun little experiment of getting it to work.
Oh, yeah. And I, and I think like, you know, one, one thing that I, I would do if you were curious, or maybe I'm just gonna try to bait some listeners into doing it is like, could be like, okay, like, let's try to, like, trace why it said Pomeranian here, and like maybe there's like some of it is about like, like dog breeds, and some of it is about like specific characteristics of a Husky, and then you can ask the same question, but instead of Husky you like try some other dog breed, and then try to see if you can like...
If you understood it, the circuit well, and if you identified where it's thinking about Huskies or where it's thinking about like kinda like breeding two different breeds, then you should be able to like swap these in and out and get it to kinda like say whatever you want.
Um, and, and if you didn't, then maybe there's something complicated going on. But, but yeah, like very cool that you got this going on so quick. That's, that's, that's the whole goal. That's super exciting.
Yeah. A- and like, you know, no disclosure, this was like five minutes of just playing around and like there's, there's stuff to learn there, right? Like, okay, what happens with dog breeding? What are traits of these dogs? And then, you know, the next step for me would basically be let's try clamping some of these features up or down.
Let's, let's do different breeds and see if it makes sense, right? So if I have Husky traits and a different mix, and then, you know, can I, can I get out what's going on? But it also shows internally that there's more than just token completion of, you know, this plus this equals this.
No, it has some under- understanding of characteristics, right? Like this is a pretty stubborn dog. It, it, it has a stubborn feature pretty high up that activates. So very, very cool stuff. I think it'll be cool when we apply this to more like serious topics.
Like right now when it comes to LEM evals, right, we have like... We have pretty straightforward evals, right? Like how good is, does it do on math? Can it write code? Does stuff compile? But we don't have like vibes-based heuristic evals, right?
So like does it understand different queries should be concise? Should they be verbose? Can we kind of trace through how it gives responses to this stuff? And then like the other part is, you know, as we go past base models, how does this happen for different phases of models, right?
So if I have a base Gemma and I have a chat model, what are differences in their attributions, right? What happens kind of in that diff of training? So that's, that's kind of one of my little interests in, in MechInterp.
What happens as we do more training? What are we really changing?
Totally, yeah. You can think about sort of like comparing different models and for me different models are either like Gemma versus some other model or like early Gemma versus late Gemma in pre-training or like fine-tuned versus not fine-tuned.
I think there's also a sense in which like somebody yesterday was telling me like, "Oh, it's fun. I've been playing with it on like the like sort of like weird riddles that the models get wrong." Like it, like you, you're not limited to studying what the model can do, right?
Like if the model's failing at something like, you know, counting the number of letters in strawberry or whatever, um, you could just try that and try to figure out the circuit for like, well, it's getting this wrong, like why?
Like it... Maybe you can see in, in its representation that it's like thinking about something obviously incorrect, right? Um, and so I think, I think that that's also like a fun thing to, to, to play with.
I think that's it for our, uh, s- little intro chat and coverage of the open sourcing. Let's dive right into the episode next. But Emmanuel, uh, your, uh, amazing work and I'm so inspired and also just like I think it, this puts a human face on the, uh, the interpretability work.
I think it's very important and we'd love to keep doing this. Whatever you got next coming up.
Well, yeah. Thanks for having me. Again, I should say cool to put a face on it, but definitely wanna call out this is like a huge-
Of course
... team of people with me. I'm just a talking head here. Um, a- a- and-
And, and paper, paper lead, you know. You did the work. You, you know-
But-
... take credit
... I think that like, yeah, happy to talk about more interp things and also like feel free to, you know, reach out to me. I'm like findable if you're listening to this podcast and you have like questions about stuff that's broken or if this brings up like experiment ideas.
Uh, definitely want more people playing with this. So, so yeah. Thanks for having me. Hope, hope that inspires some folks.
Mech Interp Start24:19
All right. We are back in the studio with a couple special guests. One, Vibhu, our guest co-host for a couple of times now, as well as Mochi the Distilled Husky is-
Mochi
... in the studio with us-
Mochi
... to ask some very pressing questions.
Ooh.
Uh, as well as Emmanuel. I didn't get your last name. Amiesen?
Yep.
Is that Dutch? Is that-
It's actually German.
German?
Yeah. Yeah.
You are the lead author of a fair number of the recent, uh, MechInterp work from Anthropic that I've been basically calling Transformer Circuits, 'cause that's the name of the publication.
Yeah. Well, to be clear, Transformer Circuits is the whole publication. I'm the author on one of the recent papers, Circuit Tracing.
Yes, and people are very excited about that. The other name for it is, like, Tracing the Thoughts of LLMs. There's like three different names for this work. It's like
Yeah.
But it's all MechInterp.
It's all MechInterp. There's two papers. One is Circuit Tracing, it's the methods. One is, like, the biology, which is kind of what we found in the model, and then Tracing the Thoughts is confusingly just the name of the blog post-
Yeah
... poorly announced it.
It's for different audiences, and I think the- when you produce the, like, two-minute polished video that you guys did, that's meant for, like, a very wide audience, you know?
Yeah, that's right. There's sort of, like, very many levels of granularity at which you can go, and I think for MechInterp in particular, 'cause it's kind of complicated going, you know, from, like, top to bottom, most, like, high level to sort of the nitty-gritty details works pretty well.
Yeah. Cool. Um, we can get started. Basically, we have two paths that you can, you can choose. Like, either your personal journey into MechInterp or the brief history of MechInterp just generally, and maybe that might coincide a little bit.
I think my-- Okay, I could just give you my personal journey very quickly 'cause, uh, then we can just do the second path. My personal journey is that I was working at Anthropic for a while. I'd been, like many people, just following MechInterp as sort of like an interesting field with fascinating, often beautiful papers.
Uh, and I was at the time working on fine-tuning, so, like, actually fine-tuning, uh, production models for Anthropic, and eventually I got both, like, sort of like my fascination reached a sufficient level that I decided I wanted to work on it, and also I got more excited about just as our models got better and better understanding how they worked.
So that's the s- simple journey. I've got like a background in ML, kinda like did a lot of applied ML stuff before, and now I'm doing more research stuff.
Yeah, you have a book with O'Reilly.
Oh, yeah.
You're head of AI at Insight Data Science. Anything else to plug?
Yeah. Uh, I actually, I, I want to, like, plug the paper and unplug the book.
Okay.
Uh, I think the book is good. I think the, the advice stands the test of time, but it's very much like, "Hey, you're building, like, AI products. Which do you focus on?" It's, like, very different, I guess, is all I s- I'll say from, from the stuff that we're talking about today.
Today is, like, research. Some of, some of the sort of like deepest, weirdest things about how, like, models work and this book is you wanna ship a random forest to do fraud classification. Like, here are- ... the top five mistakes to avoid.
Yeah.
Um.
The good old days of ML.
I know. It was simple back then.
You also transitioned into research, and I think you also did the management trans-- Like, I feel like there's this monolith of, like, people assume you need a PhD for research. Maybe can you give, like, that perspective of, like, how do people get into research?
How did you get into research? Maybe that gives audience a, a, a insight into Vibhu as well, your background.
Yeah. My background was in, like, economics, data science. I thought LLMs were pretty interesting. I started out with some basic ML stuff, and then I saw LLMs were starting to be a thing, so I just went out there and did it.
And same thing with AI engineering, right? You just kind of build stuff, you work on interesting things, and, like, now it's more accessible than ever. Like, back when I got into the field five, six years ago, like, pre-training was still pretty new.
GPT-3 hadn't really launched, so it was still very early days, and it was a lot less competitive. But yeah, without any specific background, no PhD, there just weren't as many people working on it. But you made the transition a little bit more recently, right?
So what's your experience been like?
Yeah. I think, I think it has maybe never been easier in some ways because a lot of the field is, like, pretty empirical right now. So I think the bitter lesson is, like, this lesson that, you know, you can just sort of like a lot of times scale up compute and data and get better results than, like, thinking.
Than if you sort of like thought extremely hard about a really good, like, prior inspired by the human brain to, to train your model better. And so in terms of definitely, like, research for pre-training and fine-tuning, I think it's just sort of like a lot of the bottlenecks are extremely good engineering and systems engineering, and a lot even of the research execution is just about sort of like engineering and scaling up and things like that.
I think for Interp in particular, there's, like, another thing that makes it easier to transition t-to, which is maybe two things. One, you can just do it without huge access to compute. Like, there are open source models. You can look at them.
A lot of Interp papers, you know, coming out of programs like MATS are on models that are open source that you can sort of like d- dissect without having a cluster of like, you know, a hundred GPUs. You can just even sometimes you can load them, like, on your CPU, on your MacBook.
And it's also a relatively new field, and so, you know, there's, as I'm sure we'll talk about, there's like some conceptual burdens and, and concepts that you just want to, like, understand before you contribute, but it's not, you know, physics.
It's relatively recent, and so the number of abstractions that you have to, like, ramp up on is just not that high compared to other fields, which I think makes that transition somewhat easier. For Interp, if you understand, we'll talk about all these, I'm sure, but, like, what features are and what dictionary learning is, you're, like, a long part of the way there.
I think it's also interesting just on a careers point of view, research seems a lot more valuable than engineering. So I wonder, and you don't have to answer this if it's, like, a tricky thing, but, like, how hard is it for a, for a research engineer in Anthropic to jump the wall into research?
People seem to move around a lot, and I'm like, there, that cannot be so easy. Like, in no other industry that I know of people-- you can do that. Do you know what I mean?
Yeah. I think I'd actually like a-- I'd push back on the sort of like research being more valuable than engineering a little bit-
Okay
... because I think a lot of times, like, having the research idea is not the hardest part. Don't get me wrong, there's some ideas that are, like, brilliant and hard to find, but what's, what's hard, certainly on fine-tuning and to a certain extent on Interp, is executing on your research idea in terms of, like, making an experiment successfully, like, having your experiment run, interpreting it correctly.
What that means, though, is that, like- They're not separate skill sets. So, like, if you have a cool idea, there's kind of not many people in the world, I think, where they can just, like, have a cool idea and then they have a, you know, like, a little minion they'll deputize, being like, "Here's my idea."
You know, go off for three months and, like, run this whole, like, build this model and train it for, you know, hundreds of hours and then report back on what happened. A lot of the time, like, the people that are the most productive, they have an idea, but they're also extremely quick at checking their idea, finding sort of like the shortest path to, to checking their idea.
And, and a lot of, like, that shortest path is engineering skills, essentially. It's just, like, getting stuff done. And so I think that's why you see sort of like people move around is, like, proportionate to, to, to your interest.
If you're just able to quickly execute on the ideas you have and, and get results, then that's really the 90% of the value. And so you see a lot of transferable skills, actually, I think from, from people like, I've certainly seen at Anthropic that are just, like, really good at that inner loop.
They can apply it in one team and then move to a completely different domain and apply that inner loop just as well.
History & Concepts31:52
Yeah. Very cracked, as the kids say. Uh, shall we move to the history of MechInterp?
Yeah.
All I know is that everyone starts at Chris Olah's blog. Is that right?
Yeah, I think that's the correct answer. Chris Olah's blog and then, you know, uh, Distill.pub, uh-
Yeah
... is the sort of natural next step, and then I would say, you know, now there's, for Anthropic, there's Transformer Circuits, which you talked about, um, but there's also just a lot of MechInterp research out there from, you know, I think, like, the, yeah, like maths is a group that, like, regularly has a lot of research, but there's just, like, many different labs that, that put research out there, and I think that's also just, like, hammer home the point.
That's because all you need is, like, a model and then a willingness to kind of investigate it to be able to contribute to it. So, so now it's sort of like there's been a bit of a Cambrian explosion of MechInterp, which is cool.
I guess the history of it is just computational, like, models that are not decision trees, uh, models that are either CNNs or let's say transformers, have just this re-really, like, strange property that they don't give you interpretable intermediate states by default.
You know, again, to go back to if you were training, like, a, a, a decision tree on, like, fraud data for an old school, like, bank or something, then you can just look at your decision tree and be like, "Oh, it's learned that, like, if you make, uh, I don't know, if this transaction is more than $10,000 and it's for, like, perfume, then maybe it's fraud or something."
Uh, you can look at it and say, like, "Cool, like that makes sense. I'm willing to ship that model." But for, for things like, like CNNs and, like, transformers, we, we don't have that, right? What we have at the end of training is just a massive amount of weights that are- ...
connected somehow, uh, or, or activations are connected by some weights, and who knows what these weights mean or what the intermediate activations mean. And so the quest is to understand that. Initially it was done, a, a lot of it was done on vision models, where you sort of have the emergence of a lot of these ideas, like what are features, what are circuits.
And then more recently it's been mostly, or not most, yeah, mostly applied to NLP models, but also, you know, still there's work in vision, and there's work in, like, uh, bio and, and other domains.
Yeah. I'm on Chris Olah's blog, and he has, like, the feature visualization stuff. I think the, for me, the clearest was, like, the vision work where you could have, like, this layer detects edges, this layer detects textures, whatever.
That seemed very clear to me, but the transition to language models seemed like a big leap.
I think one, one of the bigger changes from vision to, to, like, language models has to do with, uh, the superposition hypothesis, which-
Yeah
... maybe is like-
That's the first in, like, toy models post, right?
Exactly. And this is sort of like, it turns out that if you look at just the neurons of a lot of vision models, you can see neurons that are curve detectors or that are edge detectors or that are high/low frequency detectors.
And so you can sort of, like, make sense of the neurons mostly. But if you look at neurons in language models, most of them don't make sense. It's kind of, like, unclear why, or it was unclear why that would be.
And one main, like, hypothesis here is the superposition hypothesis. So what does that mean? That means that, like, language models pack a lot more in less space than vision models. So maybe like a, a, a kind of, like, really hand-wavy analogy, right?
Is like, well, if you want curve detectors, like you don't need that many curve detectors. You know, if each, each curve detector is gonna detect like a, a, a quarter or a twelfth of a circle, like, okay, well, you have your, all your curve detectors.
But think about all of the concepts that, like, Claude or even GPT-2 need to, to know. Like, just in terms of it needs to know about, like, all of the different colors, all the different hours of every day, all of the different cities in the world, all of the different streets on every city.
If you just enumerate all of the facts that, like, a model knows, you're gonna get, like, a very, very long list, and that list is gonna be way bigger than, like, the number of neurons or even the, like, size of the residual stream, which is where, like, the models process information.
And so there's this sense in which, like, oh, there's more information than there's, like, dimensions to represent it, and that is much more true for language models than for vision models. And so because of that, when you look at a part of it, it just seems like it's like there's, it's got all this stuff crammed into it.
Whereas if you look at the vision models, oftentimes you could just, like, be like, "Cool, this is a curve detector."
Yeah. Yeah, Vibhu, you have, like, some fun ways of explaining the toy models or superposition concept.
Yeah, I mean, basically, like, you know, if you have two neurons and they can represent five features, like a lot of the early MechInterp work says that, you know, there are more features than we have neurons, right? So I guess my kind of question on this is, for those interested in getting into the field, what are, like, the key terms that they should know?
What are, like, the few pieces that they should follow, right? Like from the Anthropic side, we had a toy transformer model. We had sparse au- we first had autoencoders. What is this-
That was the second paper, um, right?
Yeah.
Uh, mono-monosemanticity.
Yeah.
Yeah.
What is sparsity in autoencoders? What are transcoders? Like, what is linear probing? What are these kind of, like, key points that we had in MechInterp? And just kind of how would people get a quick, you know, zero to, like, eighty percent of the field?
Okay, so zero to eighty percent. And now I realize I really, like, set myself up for, for failure because I was like, "Yeah, it's easy. There's not that much to know." So, okay, so then, then we should be able to cover it all.
Um, so superposition is the first thing you should know, right? This idea that, like, there's a bunch of stuff crammed in a few dimensions, as you said. Maybe you have, like, two neurons, and you want to represent five things.
So if that's true, and if you want to understand how the model represents, you know, I don't know, the concept of red, let's say, then you need some way to, like, f-find out essentially in which direction the model stores it.
So after the, the sort of like supervision hypothesis, you can think of, like, ah, we also think that, like, basically the model represents these, like, individual concepts, we're gonna call them features, as, like, directions. So if you have two neurons, you can think of it as, like, it's like the 2D plane, and it's like, ah, you can have, like, five directions, and maybe you would, like, arrange them like the spokes of a wheel, so they're sort of like maximally separate.
It could mean that, like, you have one concept this way and one concept that's, like, not fully perpendicular to it but, like, pretty, pretty, like, far from it. And then that would, like, allow the model to represent more concepts than it has dimensions.
And so if that's true, then what you want is you want, like, a model that can extract these independent concepts, and ideally you want to do this, like, automatically. Like, can we just, you know, have a model that tells us like, "Oh, like, this direction is red.
If you go that way, actually, it's like, I don't know, chicken. And if you go that way- it's like the Declaration of Independence," you know? Um, and so that's what sparse autoencoders are.
It's almost like the, the self-supervised learning insight version. Like in, in pre-training, you have self-supervised learning.
Yeah.
And here now it's so self-super-supervised interpretability.
Yeah, exactly. Exactly. It's like an unsupervised method.
Yeah.
And so unsupervised methods often still have, like, labels in the end. Or so, so sometimes I feel like the, the transformer labels-
You form labels by masking.
Yeah, like for, for pre-training, right? It's like the next token. So i-in that sense, you have a supervision signal, and here the supervision signal is simply you take the, like, neurons, and then you learn a model that's gonna, like, expand them into, like, the actual number of concepts that you think there are in the model.
So you have two neurons, you think there's five concepts, so you expand it to, like, a thing of dimension five, and then you contract it back to what it was. That's, like, the model you're training, and then you're training it to incentivize it to be sparse so that onl- there's only, like, a few features active at a time.
And then once you do that, if it works, you have this sort of, like, nice dictionary which you can think as like a way to decode, deactivate the neurons, where you're saying like, "Ah, cool, I don't know what this, what this direction means, but I've, like, used my model and it's telling me that the model is writing in the red direction."
And so that's, that's sort of, like, I think maybe the biggest thing to understand is, is this combination of things of like, ah, we have too few dimensions, we pack a lot into it, so we're gonna learn an unsupervised way to, like, unpack it and then analyze what each of those dimensions that we've unpacked are.
Any follow-ups?
Yeah, I mean, the follow-ups of this are also kind of like some of the work that you did is in clamping, right? What is the-
Mm
... applicable side of MechInterp, right? So we saw that you guys have, like, great visualizations. Golden Gate Claude was a cool example.
I was gonna say that.
Yeah.
Yeah.
So-
That was my favorite
... what can we do once we find these features? Finding features is cool, but what can we do about it?
Yeah. I think there's kinda, like, two big aspects of this. Like, one is, yeah, okay, so we go from a state where, as I said, the model is, like, a mess of weights, we have no idea what's going on, to, okay, we found features, we found a feature for red, a feature for Golden Gate Claude, or for Gol- the Golden Gate Bridge, I should say.
Like, what do we do with them? And well, if these are true features, that means that, like, they in some sense are, are important for the model, or it wouldn't be, like, representing it. Like, if the model is, like, bothering to, like, write, you know, in the Golden Gate Bridge direction, it's usually because it's gonna, like, talk about the Golden Gate Bridge.
And so that means that, like, if that's true, then you can, like, set that feature to zero or artificially set to 100, and you'll change model behavior. Uh, that's what we did when we did Golden Gate Claude, in which we found a feature that represents the direction for the Golden Gate Bridge, and then we just, like, set it to always be on.
And then you could talk to Claude and be like, "Hey, like, Claude, what's on your mind?" You know, like, "What are you thinking about today?" He'd be like, "The Golden Gate Bridge." You'd be like, "Hey, Claude, like, what's two plus two?"
He'd be like, "Four Golden Gate Bridges." uh, et cetera, right? And it was always thinking about the Golden Gate Bridge.
It's like write a poem, and it just starts talking about how it's, like, red like the Golden Gate Claude.
That's right.
Uh, Golden Gate Bridge. Yeah.
That's right.
It's amazing.
I think what made it even better is, like, we realized later on that it wasn't really, like, a Golden Gate Bridge feature. It was, like, being in awe at the beauty of the majestic Golden Gate Bridge, right? So sometimes it would, like, really ham it up.
It'd be like, "Oh, I'm just thinking about the beautiful international orange color of the Golden Gate Bridge." That was just, like, an example that I think was, like, really striking. But of, of sort of like, oh, if you found, like, a space where that represents some computation or some representation of the model, that means that you can, like, artificially suppress or promote it, and that means that, like, you're starting to understand at a very high level, a very gross level, like, how some of the model works, right?
We've gone from, like, "I don't know anything about it," to like, "Oh, I know that this, like, combination of neurons is this, and I'm gonna prove it to you." The next step, which is what this, this works on, is, like, that's kind of like thinking of if maybe you take the ana-analogy of, like, um, I don't know.
Like, like, let's take the analogy of, like, a, an, an MRI or something, like a brain scan. It tells you, like, oh, like this, as Claude was answering, at some point it thought about this thing. But it's a sort of, like, vague, like, may- basically, maybe it's, like, a, like, a bag of words, kind of.
It's like a bag of features. You just are like, here are all the random things it thought about. But what you might wanna know is, like, okay, but Claude is doing some processing. Like, sometimes to get to the Golden Gate Bridge, it had to realize that you were talking about San Francisco and about, like, the best way to go to Sonoma or something, and so that's how it got to Golden Gate Bridge.
So there's, like, an algorithm that leads to it at some point thinking about the Golden Gate Bridge, and basically, like, there's, like, a way to connect features to say, like, oh, from this input went to these few features, and then these few features, and then these few features, and then that one influenced this one, and then you got to the output.
And so that's the second part and the part we worked on is, like, you have the features, now connect them in what we call, uh, or what's called circuits, which is sort of like explaining the, the, like, algorithm.
Yeah. Before we move directly onto your work, I just want to give a shout-out to Neel Nanda. He did Neuronpedia and released a bunch of essays for, I think, the Llama models and the Gemma models.
And the Gemma models, yeah.
Uh, so I actually made Golden Gate Gemma. Just upped the weights for proper nouns and names of places of people re- and references to the term golden likely relating to awards, honors, or special names, and that together made Golden Gate.
That's amazing. Yeah.
So you can make Golden Gate Gemma, and, like, I think that's a, that's a fun way to experiment with this. Uh, but yeah, we can move on to-
I, I'm curious, I'm curious, what's the background behind why you shipped Golden Gate Claude? Like, you had so many features. Just any fun story behind why that's the one that made it?
You know, it's funny, if you look at the paper, there's just a bunch of like, yeah, like, really interesting features, right? There's like, one of my favorite ones was the sycophantic praise, which I guess is very topical right now.
Very topical.
Um, but you know, it's like you could dial that up, and, like, Claude would just really praise you. You'd be like, "Oh, you know, like, uh, I wrote this poem, like, roses are red, violets are blue," whatever, and it'd be like, "That's the best poem I've ever seen."
Um- ... and so we could have shipped that. That could've been funny. Uh, Golden Gate Claude was, like, a pure, as far as I remember at least, like, a pure just, like, weird random thing where, like, somebody found it initially.
We had an internal demo of it. Everybody thought it was hilarious, and then that's sort of how it came out. There was no-- nobody had a list of top 10 features we should consider shipping, and we picked that one.
It was just kind of like-
Yeah
... a very organic moment.
No, like the, the marketing team really leaned to- leaned into it. Like, they mailed out pieces of the Golden Gate for people at NeurIPS, I think-
Yeah
... or ICML. Yeah, it was, it was fantastic marketing. Yeah. The question obviously is, like, if OpenAI had invested more in interpretability, would they have caught, uh, the GPT-4o update? Uh, but we don't know that for sure 'cause they have interp teams.
They just-
Yeah. I think also, like, for that one, I don't know that you need interp. Like, it was pretty clear-cut. Talking to the model and I was like, "Oh, that model's really gassing me up."
And then the other thing is, um, can you just, like, up write good code, don't write bad code, and make Sonnet 3.5? And, like, it feels too, too easy, too free. Is that s- steering that powerful that you can just, like, up and down features with no trade-offs?
There was, like, a phase where people were basically saying, you know, 3.5 and 3.7 are just now-
Yeah
... because they came out right after each other.
And, and for the record, like, that's been debunked.
Yeah.
But, like-
It has been debunked, but, you know, it had people convinced that what people did is they basically just steered up and steered down features, and now we have a better model. And this kind of goes back to that original question of, right, like, why do we do this?
What can we do? Some people are like, "I want tracing from a sense of, you know, legality." Like, what did the model think when it came to this output? Some people wanna turn hallucination down. Some people wanna turn coding up.
So, like, what are some, like, whether it's internal, what are you exploring that, like, what are the applications of this? Whether it's open-ended of what people can do about this or just like, yeah, why, why do MechInterp, you know?
Yeah. There's, like, a few things here. So, so, like, first of all, obviously this is, I would say, on the scale of the most short-term to the mo- most long-term, like, pretty long-term research. So in terms of, like, applications compared to, you know, like, the research work we do on, like, fine-tuning or whatever, Interp is much more, you know, sort of like a, a high risk, high reward kind of approach.
Uh, with that being said, like, I think there's just a, a fundamental sense in which Michael Nielsen had a, had a post recently about how, like, knowledge is dual use or something. But just, like, just, like, knowing how the model works at all feels useful, and you know, it's hard to argue that if we know how the model works and understand all of the components, that won't help us, like, make models that hallucinate less, for example, or that are, like, less biased.
That seems, you know, if, if, like, at the limit, yeah, that totally seems like something you would do using basically, like, your understanding of the model to improve it. I think for now, as we can talk about a little bit with, like, uh, circuits, there's, like, we're still s- pretty early on in the game, right?
And so right now, the main way that we're using interpretability is, like, to investigate specific behaviors and understand them and gain a sense for, uh, what's causing them. So, like, one example that we, we can talk about later or we can talk about now, but in the paper, we investigate jailbreaks, and we try to see, like, why does a jailbreak work?
And then we realize as we're looking at this jailbreak that part of the reason why Claude is telling you how to make a bomb in this case is that it's, like, already started to tell you how to make a bomb, and it would really love to stop telling you how to make a bomb, but it has to first finish its sentence.
Like, it really wants to make correct grammatical sentences. And so it turns out that, like, seeing that circuit, we were like, "Ah, then does that mean if I prevent it from finishing its sentence, the jailbreak works even better?"
And sure enough, it does. And so I think, like, the level of sort of practical application right now is of that shape. So, like, understanding either, like, quirks of a current model or, like, how it does tasks that maybe we don't, we don't even know how it does it.
Like, you know, we, we have, like, some planning examples where we had no idea it was planning, and we're like, "Oh God, it is." That's sort of, like, the current state we're at.
I'm curious internally how this kind of feeds back into, like, the research, the architecture, the pre-training teams, the post-training. Like, is there a good feedback loop there? Like, right now there's a lot of external people interested, right? Like, we'll train an SAE on one layer of Llama and probe around, but then people are like, "Okay, how does this have much impact?"
People like clamping, but yeah, as you said, you know, once you start to understand these models have this early planning and stuff, how does this kind of feed back?
I don't know that there's, there's, like, much to say here other than, like, I think we're definitely interested in conversely, like, making models for which it's, like, easier to interpret them. So that's also something that you can imagine sort of-
Yeah
... like working on, which is, like, making models where you have to work less hard to try to understand what they're doing.
So, like, the architecture? Okay. Yeah, so I think there was a, there was a Less Wrong post about this of, like, there's a non-zero amount of sacrifice you should make in current capabilities in order to actually make them more interpretable because otherwise you will never catch up.
You know, there's this sort of sense in which, like, right now we take the model And then the model's a model, and then we post hoc do these replacement layers to try to understand it. But of course, when we do that, we don't, like, fully capture everything that's happening inside the model.
We're capturing, like, a subset. And so maybe some of it is, like, you could train a model that's sort of, like, easier to interpret natively. And it's possible that, like, you don't even have that much of, you know, like, uh, attacks in that sense, and you can just sort of like either, like, train your model differently or do like a little post hoc step to, like, sort of like untangle some of the mess that you've made when you trained your model, right?
Mm-hmm.
Make it easier to interpret. Um-
The hope was pruning would do some of that.
Hmm.
But I feel like that area of research has just died.
What kind of pruning are you thinking of here?
Uh, just pruning your network.
Ah, yeah.
Pruning layers, pruning connections, whatever.
Yeah. I feel like maybe this is something where, like, superposition makes me less-
Exactly
... hopeful or something because I'm like, "Ah."
Because you don't know, like, that, that, like, seventh bit might hold something.
Well, right, and it's like on, on each example, maybe this neuron is, like, at the bottom of, like, what matters, but actually it's participating, like, 5% to, like, understanding English, like, doing integrals and, you know, like, whatever, like, cr- cracking codes or something, and it's like because that rep's just, like, distributed over it, you, you, you kind of like when you naively prune, you might miss that.
I don't know.
Okay, so and then this area of research in terms of creating models that are easier to interpret from the ar- from the start, is there a name for this field of research?
I don't think so. I... And I think this is, like, very early.
Okay.
And it's, it's mostly like a dream.
Just in case there's a, there's-
Yeah
... a thing people wanna double-click on.
Yeah, yeah, yeah.
Um, I haven't come across it.
I think the higher level is like Dario recently put out a post about this, right?
Yep.
Why MechInterp is so interpret, I- important. You know, we don't wanna fall behind. We wanna be able to interpret models and understand what's going on. Even though capabilities are getting so good, it kind of ties into this topic, right?
Like-
Yeah
... we want models to be slightly easier to interpret so we don't fall behind so far.
Well, yeah, and I think here, like, just to talk about the elephant in the room or something, like, like, one big concern here is, is, like, safety, right? And so, like, as models get better, they are gonna be used more and more places.
You know, it's like you're not gonna have your... You know, we're vibe coding right now. Maybe at some point we'll- that- that'll just be coding. It's like Claude's gonna write your code for you, and that's it, and Claude's gonna review the code that Claude wrote, and then Claude's gonna deploy it to production.
Um, and at, at some point, like, as these models get integrated deeper and deeper into more and more workflows, it gets just scarier and scarier to know nothing about them. And so you kind of want your ability to understand the model to scale with, like, how good the model is doing, which that itself kind of like tends to scale with, like, how widely deployed it is.
So as we, like, deploy them everywhere, we want to, like, understand them better.
The version that I liked from the old Super Alignment Team was weak to strong generalization or weak to strong alignment, which that's what super alignment to me was, and that was my first aha moment of like, oh yeah, some- at some point these things will be smarter than us.
In, in, in many ways they already are smarter than us, and we rely on them more and more. We need to figure out how to control them, and this, this is not a, like, an Eliezer Yudkowsky like, ah, thing.
It's just more like we don't know what... how these things work. Like, how can we use them?
Yeah. And like you can think of it as there's many ways to solve a problem, and some of them, if the model is solving it in, like, a dumb way or in like memorized one approach to do it, then you shouldn't deploy it to do, like, a general thing.
Like, like you could look at how it does math, and based on your understanding of how it does math, you're like, "Okay, I feel comfortable using this as a calculator," or like, no, it should always use a calculator tool because it's doing math in a stupid way.
And extend that to any behavior, right? Where it's just a matter of, like... Think about it if, if like you're like in the 1500s and I give you a car or something, and I'm just like, "Cool, like, this thing, when you press on this, like, it accelerates.
When you press on that, like, it stops. You know, this steering wheel seems to be doing stuff." But you knew nothing about it. I don't know. If it was like a, a, a super faulty car, and it's like, oh yeah, but if you ever get, went a- above 60 miles an hour, like it explodes or something.
Like you probably would be sort of like you'd want to understand the nature of the object before, like, jumping in, in it. And so that's why we, like, understand how cars work very well because we make them. LLMs are sort of like, and ML models in general, are like this very rare artifact where we, like, make them, but we don't- we have no idea how they work.
We evolve them. We create conditions for them-
That's right
... to evolve, and then they evolve, and we're like, "Cool," like, you know, maybe we got a good run, maybe we didn't.
Yeah.
Don't really know.
Yeah. The extent to which you know how it works is you have your, like, eval and you're like, "Oh, well, seems to be doing well on this eval." And then you're like, "Is it because this was in a training set, or is it, like, actually generalizing?"
I don't know.
Uh, my favorite example was somehow C4, the, the common, the Colossal Clean Corpus, did much better than, uh, Common Crawl even though it filtered out most of this, like... It was very prudish, so it, like, filters out anything that could be considered obscene, including the word gay.
But like, somehow, it just, like, when you add it into the data mix, it just does super well. And it's just like this magic incantation of like, "This r- this recipe works. Just trust us. Like, we've tried everything.
This one works, so just go with it."
Yeah.
It's not very satisfying.
No, it's not. The side that you're talking about, which is like, okay, like how do you make these? And it's kind of like unsatisfying that you just kind of make the soup and you're like, "Oh, well, you know, my grandpa made the soup with these ingredients.
I don't know why, but I just make the soup-
Yeah
... the way my grandpa said." And then, like, some- one day somebody added, you know, cilantro, and since then we've been adding cilantro for generations. And you're like, "This is kind of crazy."
That, that's exactly how we train models, though.
Yeah, yeah. Um, so I think there's, there's like a part where it's like, okay, like let's try to unpack what's happening, you know, like the mechanisms of learning, like how, how our models learn. And like one of them, I guess, I guess we skipped over it, but like one of the inter-
Yeah, let's go for it
... things were like induction heads. You know, like understanding what induction heads are, which are attention heads that allow you to look at, in your context, the last time that something was mentioned and then repeat it, is like something that happens, it seems to happen in every model, and it's like, oh, okay, that makes sense.
That's how the model, like, is able to, like, repeat text without dedicating too much capacity to it.
Let's get it on screen so people can see.
The visuals of the work you guys put out is amazing by the way.
Oh, yeah. We should talk-
I highly, highly recommend it.
We should talk a little bit about the ba- behind the scenes of that, that kind of stuff. But, but let's, let's, let's finish this off first.
Totally. Uh, but just really quickly, I don't think we should spend too long on it. I think it's just like if you're interested in MechInterp, we talked about superposition, and I think we skipped over induction heads, and that's like- You know, kind of like a really neat basically pattern that emerges in many, many transformers where essentially they just learn.
Like, one of the things that you need to do to, like, predict text well is that if there's repeated text, at some point somebody said Emmanuel Mason, and then you're like on the next line and they say Emmanuel, very good chance it's the same last name.
And so one of the first things that models learn is just like, "Okay, I'm just gonna, like, look at what was said before, and I'm gonna say the same thing," and that's induction heads, which is, like, a pair of attention heads that just basically look at the last time something was said, look at what happened after, move that over.
And that's an example of a mechanism where it's like, cool, now we understand that pretty well. There's been a lot of follow-up research on understanding better like, okay, like in which context do they turn on? Like, you know, there's like different, like, levels of abstraction.
There's, like, induction heads that, like, literally copy the word, and there's some that copy, like, the sentiment and other aspects. But I think it's just, like, an example of slowly unpacking, you know, or like peeling back the layers of the onion of like what's going on inside this model.
Okay, this is a component, it's doing this.
Mm. So the induction heads was, like, the first major finding?
It was a big finding for NLP models-
Yeah
... for sure.
I often think about the edit models. So Claude has a fast edit mode, uh, I, I forget what it's called. OpenAI has one as well. And you need very good copying, every area that needs copying, and then you need it to switch out of copy mode when you need to start generating.
Right.
And that is basically the productionized version of this.
Yeah. Yeah, yeah. And it turns out that, you know, you, you need to select a model that's, like, smart enough to know when it g- needs to get out of copy mode, right?
Yeah.
Which is, like, not as well.
It's fascinating.
But yeah.
It, it's faster, it's cheaper. You know, as bullish as I am on Canvas, basically every AI product needs to iterate on a central artifact and, like, if it's code, if it's a piece of writing, doesn't really matter, but you need that copy capability that's smart enough to know when to turn it off.
That's why it's cool that induction heads are at different levels of abstraction. Like, sometimes you need to, editing some code, you need to copy, like, the general structure. It's like, oh, like the last-- like, this other function that's similar, it first takes like, you know, I don't know, like, abstract class and then it takes, like, an int.
So I need to, like, copy the general idea, but it's gonna be a different abstract class and a different int or something.
Yep. Cool.
Um, yeah.
Circuits & Reasoning57:15
So tracing?
Oh, yeah. Should we jump to circuit tracing? Sure.
Uh, I don't know. If there's anything else you wanna cover.
No, no, no.
We got space for it.
Maybe just-- Okay, I'll, I'll do like a really quick TLDR of these two recent papers.
Okay.
Uh, insanely quick. So we talked about these features that we detect, and what we said is like, okay, but we'd like to connect the features to understand, like, the inputs to every features and the outputs to every features, and basically draw a graph.
And this is like, if I'm still sharing my screen-
You are
... uh, the thing on the right here, where, like, that's the dream. We want, like, for a given prompt, what were all of the things-- like, all of the important things happen in the model, and here it's like, okay, it took in these four tokens, those activated these features, these features activate these other features, and then these features activate these other features, and then all of these, like, promoted the output, and that's the story.
And basically we're, we're like the work is to sort of use dictionary learning and these replacement models to provide a explanation of, like, sets of features that explain behavior. So this is super abstract, so I think immediately maybe we can, like, just look at one example.
I can show you one, which is this one.
Ah, the reasoning one. Yep.
Yeah, two-step reasoning. I think this is already-- this is like the introduction example, but it's already, like, kind of fun. So, so the question is, you ask the model something that requires it to take a step of reasoning in its head.
So you say, you know, fact, the capital of the state containing Dallas is. So to answer that, you need one intermediate step, right? You need to say, "Wait, where's Dallas? It's in Texas. Okay, cool. Capital of Texas, Austin."
And this is, like, in one token, right? It's gonna-- after is, it's gonna say Austin. And so, like, in that one forward pass, the model needs to extract, to realize that you're asking it for, like, the capital of a state to, like, look up the state for Dallas, which is Texas, and then to say Austin.
And sure enough, this is, like, what we see is, we see, like, in this forward pass there's a rich sort of, like, inner set of representations where there's, like, it gets capital state and Dallas, and then boom, it has an inner, uh, representation for Texas, and then that plus capital leads it to, like, say Austin.
I guess one of the things here is, like, we can see this internal, like, thinking step, right? But a lot of what people say is like, "Is this just memorized fact", right? Like, I'm sure a lot of the pre-training that this model is trained on is this sentence shows up pretty often, right?
So this shows that no, actually internally throughout we do see that there is this middle step, right? It's not just memorized token prediction.
You, you-- So you can prove that it generalized.
Yes.
Yeah. So, so, so that's exactly right, and I think, like, you, you, you hit the, the nail on the head, which is like this is what this example's about. It's like, ah, if this was just memorized, you wouldn't need to have an intermediate step at all.
You'd just be like, "Well, I've seen the sentence, like I know it comes next," right? But here there is an intermediate step. And so you could say like, "Okay, well, maybe it just has the step, but it's memorized it anyways."
And then the way to, like, verify that is kind of like what, what we do later in the paper and for all of our examples is like, okay, we claim that this is, like, the Texas representation. Let's get another one and replace it, and we just change, like, that, uh, feature in the middle of the model, and we change it to, like, California.
And if you change it to California, sure enough it says Sacramento. And so it's like this is not just a, like, byproduct, like it's memorized something and on the side it's thinking about Texas. It's like, no, no, no.
This is, like, a step in the reasoning. If you change that intermediate step, it changes the answer.
Very, very cool work. Underappreciated.
Yeah.
It's just, "Okay, sure." I have never really doubted-- I think there's a lot of people that are always criticizing LLMs as stochastic parrots. This pretty much disproves it already. Like, we can move on.
Yeah. I mean, I, I, I think, I think there's a lot of examples that I will say we can go through, like, a, a few of them that, like, show an amount of depth in the intermediate states of the model that makes you think, like, oh gosh, like it's doing a lot.
I think maybe, like-
The poems?
Well, definitely the poems, but even for this one, I'm gonna like scroll in this very short paper- ... to like, uh, medical diagnoses.
I don't even know the word count 'cause there's so many, like, embedded things in there.
Yeah. We-
Uh-
It's, it's too dangerous. We can't look it up. It overflows. Um-
It's so beautiful. Look at this
Uh, this is like a medical example that I think shows you-- Again, this is in one forward pass. The model's, like, given a bunch of symptoms, and then it's asked, not like, "Hey, what is, what is the, like, disease that this person has?"
It's asked like, "If you could run one more test to determine it, what would it be?" So it's even harder, right? It means, like, you need to take all the symptoms, then you need to, like, have a few hypotheses about what the disease could be, and then based on your hypotheses, say like, "Well, the thing that would, like, be the right test to do is X."
And here you can see these three layers, right? Where it's like, again, in one forward pass, it has a bunch of like, oh, these are symptoms, then it has the most likely diagnosis here, then like an alternate one, and then based on the diagnosis, it, like, gives you basically a bunch of things that you could ask.
And again, we do the same experiments where you can, like, kill this feature here, like suppress it, and then it asks you a question about the second, the sort of like second option it had. Um, the reason I show it is like, man, that's like a lot of stuff going on.
Like, for, for one forward pass, right? It's like specifically if you, if you expected it to like, oh, what it's gonna do is it's just like seen similar cases in the training. It's gonna kinda like vibe and be like, "Oh, I guess like there's that word," and it's gonna say something that's related to like, I don't know, headache, you know?
Like kinda like read ahead of it. It's like, no, no, no. It's like activating many different distributed re-representations, like combining them and sort of like doing something pretty complicated. And so yeah, I think, I think it's funny because in my opinion, that's like, yeah, like, oh God, stochastic parrots is not something that I think is, is like appropriate here, and I think there's just like a lot of different things going on, and there's like pretty complex behavior.
At the same time, I think it's in the eye of the beholder. I think like I've talked to folks that have, like, read this paper, and they've been like, "Oh yeah, this is just like a bunch of kind of like heuristics that are like mashed together," right?
Like the model's just doing like a bunch of kind of like, oh, if high blood pressure, then this or that. And so I think there's, there's sort of like, um, an underlying question that's interesting, which is like, okay, now we know a little bit of how it works.
This is how it works. Like, now you tell me if you think that's like impressive, if you think that, like if you trust it, if you think that's sort of like, uh, something that is, that is sufficient to, like, ask it for medical questions or whatever.
I think it's a way to adversarially improve the model quality.
Yeah.
Because once you can do this, you can reverse engineer what would be a, a sequence of words that to a human makes no sense or lets you arrive at the complete opposite conclusion, but the model still gets tripped up by.
Yeah.
And then you can just improve it from there.
Exactly. And, and this gives you a hypothesis about, like, you, like, specifically imagine if, like, one of those was actually the wrong symptom or something. You'd be like, "Oh, it's weird that the liver, uh, condition like, you know, outweighs this other example.
That doesn't make sense. Okay, let's, like, fix that in particular." Exactly. You sort of have like a, a bit of, uh, of insight into, like, how the model is getting to its conclusion.
Yeah.
And so you can see both, like, is it making errors, but also is it using the kind of reasoning that will lead it to errors?
There's a thesis, I mean, now it's very prominent with the reasoning models about model death. Uh, so like you're doing all this in one pass.
Yeah.
But maybe you don't need to 'cause you can do more passes.
Sure.
Uh, and so, uh, people want shallow models for speed, but you need model death for, for this kind of thinking.
Yeah.
Is there a Pareto frontier? Is there a, is there a direct trade-off?
Yeah. I mean-
What would you prefer if you had to make a model and like, you know, shallow versus deep?
There's a chain of thought faithfulness example. Uh, before I show it, I'm just gonna go back to the top here. So when the model is sampling many tokens, if you want that to be your model, you need to be able to trust every token it samples.
So, like, the problem with, with models being auto-aggressive is that, like, if they, like, at some point sample a, a mistake, then they kinda keep going conditioned on that mistake, right? And so sometimes, like-
You need back, backspace tokens or whatever.
Yeah, yeah, yeah. And error correction is, like, notably hard, right? If you have like a deeper model, maybe you have, like, fewer CoT steps, but, like, your, your steps are more likely to be, like, robust or, or correct or something.
And so I think that, that's one way to look at the trade-off. To be clear, I, I don't have an answer. I don't know if I want a, a wide or a, or a shallow or a deep model that like-
You definitely want shallow for inference speed.
Sure.
Yeah.
Sure, sure, sure. But you're trading that off for, for something else, right?
Yeah, yeah.
'Cause you also want, like, a 1B model for inference speed, but that also comes at a cost, right?
Yep.
It's, it's less smart.
There's a cool quick paper to plug that we just covered on the Paper Club. It's a survey paper around when to use reasoning models versus dense models. What's the trade-off? The-- I think it's the economy of-
Reasoning economy paper
... reasoning, the reasoning economy. So they just go over a bunch of, you know, ways to me-measure this, benchmarks around when to use each. Because, yeah.
Yeah.
Like, you know, we don't wanna-- Also, like, consumers are now paying the cost of this, right? But little, little side note.
Yeah.
Yeah. For those on YouTube, we have a secondary channel called LaneSpace TV where we cover that stuff.
Nice.
That's our Paper Club. We covered your paper.
Awesome.
Cool. Yeah, I think you brought up the, like, planning thing. Maybe it's worth-
Let's do it.
Yeah. I think, I think this one is like, if you think about-- Okay. So-
So you're, you're going into the chain of thought faithfulness one?
Let's skip this one and let's just do planning. So if you think about, like, you know, common questions you have about models, the first one we, we kind of asked was like, okay, like, is it just doing this, like, vibe-based one-shot pattern matching based on existing data, or does it have, like, kind of rich inner representations?
It seems to have, like, these, like, intermediate representations that make sense as the abstractions that you would reason through. Okay, so that's one thing. And there's a bunch of examples. We talked about the medical diagnoses. There's, like, the multilingual circuits is another one that I think is cool, where it's like, oh, it's sharing representations across languages.
Another thing that you'll hear people mention about language models, which is that they're like, uh, next token predictors.
Also, for, for a quick note, for people that won't dive into this super long blog post- I know you highlighted, like, ten to twelve. So for, like, a quick fifteen, thirty second, what do you mean by they're sharing thoughts throughout?
Just like what's a really quick high level just for people that-
Yeah. The really quick high level is that what we find is that-- Here, I'm gonna just like show you a really quick. Inside the model, if you look at, like, the inner representations for concepts, you can ask, like, the same question, which I think in the paper, the original one we ask is, like, the opposite of hot is, you know, cold.
But you can, you can do this over a larger dataset and ask the same question in many different languages, and then look at these representations in the middle of the model and ask yourself, like, well, when you ask it the opposite of hot is, and French, which is the same sentence in French.
Show off.
Is it is it using the same- Features? Or is it learning independently for each language? It kind of would be bad news if it learned independently for each language, 'cause then that means that, like, as you're pre-training or fine-tuning, you have to relearn everything from scratch.
So you would expect a better model to kind of, like, share some concepts between the languages it's learning, right? And you-- Here we do it for, like, language to languages, but I think you could argue that you'd expect the same thing for, like, programming languages, where it's like, oh, if you learn what an if statement is in Python, maybe it'd be nice if you could generalize that to Java or whatever.
And here we find that basically you see exactly that. Here we show, like, if you look inside the model, if you look at the middle of the model, which is the middle of this plot here, models share more features, they share more of these representations in the middle of the model, and bigger models share even more.
And so the, like, the sort of, like, smarter models use more shared representations than the dumber models, which might explain part of the reason why they're smarter. And so this, this was, like, sort of this, this other finding of like, oh, not only is it, like, having these rich representations in the middle, it, like, learns to not have redundant representations.
Like, if you learn the concept of heat, you don't need to learn the concept of, like, French heat and Japanese heat and Columbian. Like, you just, you just-- that's just the concept of heat, and you can share that among different languages.
I feel like sometimes overanalyzing this becomes a bit of a problem, right? Like when we talked about with the medical example, uh, we could look back and try to fix this in dataset. So in language, I don't remember if it was OpenAI or Anthropic, where they basically said when the model switched languages and they pass it to fluent users, they said, "Oh, this, this feels like an American that's speaking this language," right?
So at some times there are nuances in a slightly different representation, right? So you don't wanna over-engineer these little fixes when you do see them. But then the other side of this is, like, for those tail end of languages, right?
For languages that models aren't good at and for those, like, you know, when you wanna kind of solve that last bit, it seems like, you know, it's pretty plausible that we can solve this because these concepts can be shared across languages as long as we can, you know, fill in some, some level of representation, unless I'm wrong.
No, totally. And, and I think, like, this sort of stuff also explains, you know, uh, language models are really good at in-context learning. Like, you give them something completely new, they can do a good job. It's like, well, if you give them, like, a new fake language, uh, and you, like, in that language e-explain that, like, cold means this and hot means that, you know, like presumably they're able to-- To be clear, this is speculation, we don't show it in the paper.
But they're able to, like, bind it.
It's not-- Uh, uh, Google's done this.
Okay. Great.
Yeah. They took a l- low resource language, dumped it in a million token context, and then it came up.
That's right. That's right. Well, I guess the thing that, the thing that I'd be curious to see is, like, okay, does it use, does it reuse these representations? I bet that it probably does, right? And that's probably, like, a reason why it works well is like, well, it can reuse the representation, the general representations that it's learned in other languages.
Yeah. This is like-- I don't-- Have you talked to any linguistics people?
Not recently.
Linguistics researchers would be very interested in this because ultimately this is the ultimate test of Sapir-Whorf, um, which-
Oh, yeah
... are you familiar with?
Sapir-Whorf hypothesis, yeah.
So for those who don't know, s- uh, it's basically the idea that the language that you speak influences the way you think, which obviously it directly maps onto here. If every, if it's a w- complete mapping, if every language maps every concept perfectly on in, like, the theoretical infinitely sized model, then Sapir-Whorf is false because there is a universal truth.
If it does not, if there is some overlap where, for example, there's some languages that have no word-- There's this joke where, like, uh, I mean, you know, Eskimos have no word for snow or something like that, right?
Mm-hmm.
Or water has no word-- Uh, fish have no word for water. There's an African language where there's a gender for vegetables. You know, stuff like that. Just like-
Yeah
... languages influence the way you think, and so there should not be 100% overlap at some point.
Of course, it's, like, at the limit of the infinite model, so who knows if we'll ever- But, but yeah. Well, and I, and I think it's, it's interesting, we also show a little below that, like, some people have made the point of, like, the bias, "Oh, it sounds like an American speaking a different language."
And it does seem like the sort of, like, inner representations have a higher connection to, like, the output logits for English logits. And so there's, like, some bias-
Yeah
... uh, towards English, uh, at least in the model we studied here.
Any thoughts as to whether multimodality influences, um, any of this? So, like, concepts-
Mm
... do they map across languages as they do across modalities?
Yeah. So we show this in, uh, the Golden Ga-- or, like, the previous paper. I might have it here actually for you.
There's a good diagram of this in the essays where the same concept-
Yeah. Oh, there you go
... in text and in image. Yeah.
Ah, this is our, our buddy, the Golden Gate Bridge. Here we're showing, like, the feature for the Golden Gate Bridge, and in orange is, like, what it activates over. And so you're like, okay, so this is when the model is, like, reading text about the Golden Gate Bridge.
And we also show other languages. This is, uh, you'll have to take my word for it, but also about the Golden Gate Bridge. And then we, we show, like, the photos for which it activates the most, and sure enough-
Yeah
... it's the Golden Gate Bridge. And so again, like, that shows an example of a representation that's shared across languages and shared across modalities.
Yeah. Yeah. I think that's very relevant for, like, the autoregressive image generation models, uh, and then now the audio models as well. Something I'm trying to get some intuition for, which you probably don't have a s- off the bat answer, is how much does it cost to add a modality?
Right.
'Cause a lot of people are saying like, "Oh, just add some different decoder and then align the latent spaces and you're good." And I'm like, "I don't know, man. It sounds like there's a lot of information lost between those."
Yeah. I definitely do not have a good intuition for this, although I will say that things like this, right, make you think that if you train on multiple modalities, then you'll definitely get this, like, alignment-
Truth
... right? Yeah. But, but if, if you, like, train on one and then post-hoc train on another, maybe, maybe it'll be harder or, like, train some adapter layer.
Sure. Okay. So official answer is don't know, but someone-
Official answer is-
Someone could figure it out
... shrug.
Okay.
Yeah.
I think there are people who know, and they just haven't shared.
Well, you need to find them and get them on this podcast. Did we wanna do the, like, planning example?
Correct. Yeah. Now we're backtracking up the s- up the stack.
All right.
Yeah. Going to planning problems.
Planning example, I think, again, is like, I like this example because of the next token predictor concept. So I think this is actually, like, really important to kind of, like, dive into. So maybe what I'll say is, like, language models are next token predictors is, like, a fact.
Like, that is what they do.
That's the objective.
They, they are trained to predict the next token. However, that does not mean that- They myopically only consider the next token when they choose the next token. You can work on predicting the next token, but still, like, doing so in a way that helps, helps you predict the token, like ten tokens in the future.
And I think, well, now we definitely know that they're not myopically predicting the next token, and I think at, at least for me, that was a pretty big update because you could totally imagine that they could do everything they're doing by just, like, being really good at predicting the next token, but sort of like not having an internal state.
Like it's, it's-- it, it wasn't a given that they were gonna, like, represent internally, "Oh, this is where I wanna go, and so I'm gonna predict the next token." And so this example shows, like, an example, like the model doing exactly that.
Do you have it on screen, by the way?
Let me actually do that.
Okay, yeah, yeah.
Yeah, yeah, yeah.
Okay. Sorry, I did it just in case.
So while, while you pull it up, some of the early connections I made to this were like early, early transformers. So think BERT encoder-decoder transformers, right?
Mm.
When they came out, some of the suggestions were you don't take the last layer, right? You take off the last layer. So if you wanna do a classification task, a translation task for these encoder-decoder transformers, they've kind of overfit on their training objective, right?
So-
Yeah
... they're really good at mass language modeling, at filling in, you know, sentence order, stuff like that. So what we wanna do is we wanna throw away the top layer. We wanna freeze the bottom layers. And then there was a lot of work that was done, you know.
Where should we mess with these models?
Mm-hmm.
Should we take out like, you know, the top three layers? Should we look at the top two? Where should we probe in? Because we can see different effects, right? So we know at the very end they've overfit on their task, but there's a level at which, you know, when we start to change and we start to continue training or fine-tuning-
Yeah
... we get better output. So-
Totally
... we could start to see that, you know, throughout layers there's, there's still a broader, like, understanding the language, and then we can add in a layer, whether that's classification, and then fine-tune, and you know, it learns our task.
And this planning example is sort of like a more robust way to look into that.
Yeah. Yeah. And I think if you look at, like, all of the examples in the paper, you kind of, uh, at the bottom we have this list of, like, consistent patterns, and one pattern you see is kinda exactly what you're talking about.
Like, at the top, the sort of like-- Here, actually I have one here. The sort of like top features that are, like, right before the output are often just about, like, what you're gonna say. It's next token predictions.
Like, oh, I'm gonna say Austin. I'm gonna say rabbit. I'm gonna say... So it's kinda like not very abstract. It's just like a motor, it's a motor neuron for a human, right? It's like, oh, I've, I've decided that I want a drink of water, and so I'm gonna just grab the bottle.
And at, at the bottom, they're all, like, decoder, like, basically like sensory neurons. They're just like, oh, I just saw the word X or I just saw this. And so if you want to, like, yeah, like extract the interesting representations, a lot of the time they're in the middle.
That's where the, like, shared representations across language are, and that's where here this, like, plan is. To, like, walk through the example really briefly, it's like you have a poem, and in order to say-- You have the first line of a poem, and in order to say the second line of the poem, well, if you want to rhyme, you need to, like, identify what the rhyme of the first line was.
You're just at the end of the first line, so you need to say, like, "Okay, what's my current rhyme?" And then you need to, like, think about what your poem is talking about, and then think about candidate words that rhyme and that are, like, on topic for your poem.
And so here this is what's happening, right? It's like the last word is it, and so there's a bunch of features that are actually, they represent the direction, like rhyming with eat or at. And by the way, we, like, looked at a bunch of poems internally, and you have like-- I thought it was, like, really beautiful.
You have these models. They have a bunch of features for, like, oh, this word has, like, AB in it. Oh, this word has, like, many consonants. Oh, this word, like, is, like, you know, kind of, k-kinda like has some flourish to it.
They have, like, a bunch of, of, like, features that track various aspects that you would want to use if you were writing poetry.
It's just like ConvNets and, like, all the feature detection stuff.
Totally.
Yeah.
Totally. Uh, but I think I, maybe I, I didn't expect there to be as many features about just, like, sounds of words and kind of musicality, which I, I thought was kind of, kind of neat. But then once it's extracted the rhyme, then it comes up with sort of like these two candidates.
In this case it's like, ah, either I'm gonna finish with rabbit or I'm gonna finish with habit. The cool thing here is here we show that, like, this happens at the new line. So it happens before it's even started the second line.
And it turns out that, like, you can then say, oh, is this what the plan's actually using? We do our usual experiments. We, like, remove it, and the model writes a completely different line. We inject something, and it writes a completely different line.
We have these, like, fun examples here I'll show, which is-
Just as a mechanical thing, you c-
Yeah
... you, you just, you just disallow generation of a certain logic. Is, is that-
For, for how we do these interventions?
Yeah, yeah.
Basically, what these features are is they're like directions in the model.
Okay.
So to, like, remove them, we just write in the opposite direction. So we run the model normally, and then, like, at the, like, layer where it was gonna write, let's say in, like, you know, this, this direction, we just, like, stop it.
Just negative everything.
Yeah. We either, like, add a negative that, like, compensates for it or add a negative that goes even more in the negative direction sometimes to, like, really kill it. And then we can also add another direction, right? So in these random examples here, where l-- where like you have this poem, "The silver moon cast a gentle light," and then Claude 3.5 Haiku would, like, rhyme with, "Illuminating the peaceful night."
But then if we, like, go negative in the night direction and just add, like, green, the whole second line it's gonna write is just, "Upon the meadow's verdant green." And so that's all that we're doing. We're saying, like, we found where it stores its plan, and we, like, delete or, like, suppress the one that's stored and go in a direction of something else that's arbitrary.
And the result that's, like, striking here is sort of like two things. I think, like, one, this plan is made well in advance of needing to predict night. It's made, like, after the first line, before it's even started the second line.
And two, this plan doesn't just control, like, what you're gonna rhyme with. It's also doing what's called, like, backwards planning, where it's like, well, because I need to finish with green, I'm not gonna say illuminating the peaceful night, 'cause then I'd be like illuminating the peaceful green.
That doesn't make sense. I need to say a completely different sentence that lets me finish with green. And so there's a circuit in the model that decides on the rhyme and then works backwards from the rhyme-
Influences
... to set up your sentence.
Yeah. It's almost like backprop, but-
In the future
Yeah. It's like doing like basically like a search-
Because the green is, is backpropping through these words, so verdant and m- meadow are both green related.
Yeah, but it's doing all of that in its forward passes.
Yep.
Right? In context, which is kind of crazy.
I thought intuitively makes sense, right? So looking at it from a model architecture perspective where basically you just have a bunch of attention and feed forward layers, and then at the end you have, you know, what's the softmax over the next token, you would expect that end would really be like that grabber, right?
It's just picking tokens, so that's what it's gonna do. And early on, like even with traditional models, we could see different concepts that would start to pop up through early layers and, yeah, you have some of this throughout your architecture.
So it's very cool to see. The kind of other question that comes up is like how are we labeling these features? How are we defining them? Are we doing that right? And like, you know, what is a these words end with like I-T feature?
How do we kind of come to that conclusion?
Yeah.
Like how do we map a name to this, right? Like-
Yeah. So I, I think there's, this is like an important question 'cause you can totally imagine like fooling yourself, right?
Yeah.
Is there like a guy at Anthropic that just maps 30,000 features and-
Yeah, yeah, it's me.
And another thing-
I'm the guy.
You're the guy.
He's the guy. I did notice also like with the previous work, the scaling up SAEs-
Yeah
... as you train bigger and bigger ones, a lot of features don't activate.
Yeah.
So I think like-
You have dead features
... 60% of the 34-
Right
... million one didn't act-
So I think there's like a few questions behind your question. Like the first question was like how do you even label the features? You're telling me this is a rabbit feature, like why should I trust you? And I think there's kinda like two things going on.
So one, as I mentioned at the start, all of this is unsupervised, and so in the paper we have these links to like these little graphs which show you like more of what's going on, but this graph is just like completely unsupervised.
So it's like we train this like model to like untangle the representation, right? This like dictionary that we talked about. That gives us the features, and then we like just do math to figure out like which features influence which other features and throw away the ones that don't matter.
And then at the end we have these features. So right now we don't have any interpretation for them. We just say like, "These are all the features that matter," and then we manually go through and we look at the features.
You know, we look at this feature, and we look at that feature, and, uh, let's pick one. So this one we've labeled, say, habit. So how do we do that? You could just look at it, and we show you like what it activates over.
And if you just look at this text, maybe I'll like zoom in, like you'll immediately notice something, I think. Well, I'll immediately notice something 'cause I've stared at 30,000. I'll point it out for you. The orange is where the feature activates.
The next word after the orange is always habit. Habit, habit, habit, habit, habit, habit. So this feature always activates before habit. That's like the main source of an interpretation. We have other things. Like above we also show you like what logit it promotes, so like what output it promotes, and here it promotes hab.
So that makes sense. Um, and so that's like how we interpret and how we say, "Okay, like I think this is the say habit feature." But maybe, you know, for this one it's pretty clear, but some of them might be more confusing.
It might not be clear from these like activations what it is. The other way that we build confidence is like once we've built this thing and we said, "Oh, I think this is rhymes with eat. This is ha- say habit," that's where we do our interventions, right?
And it's like I claim this is the like I've planned to end with rabbit. To verify whether I'm, I'm right or not, I'm gonna just like take that direction, nuke it from the model, and see if the model stops saying rabbit.
And sure enough, if you do that, and here it's like we stop saying rabbit, it says habit instead. And here it's like we stop-
Yeah
... it from saying rabbit and habit, it says crabbit in this case. Not a great rhyme, but we'll work with it.
Is this something you can do like programmatically? Like can we scale this up? Can we kind of do this autonomously, or how much like manual intervention is this?
There's been a lot of work in sort of like automated feature interpretability, and it's something that we've invested in and that like other labs have invested in. And I think basically the answer is we can definitely automate it, and we're definitely gonna need to.
And right now the, the most manual parts are this sort of like look at a feature and figure out what it is, as well as, uh, group similar features together. One thing I hinted at is that actually like all of these little blocks here, there are multiple features.
You can see here it's like five features doing the same thing.
Once again-
None of that is too hard for Claude.
Very cool, very cool graphics and blog posts you guys put out, like-
Yeah.
We'll, we'll have to ask about the behind the scenes on this one.
Yeah.
Yeah, but let's, let's round out the other things to know. Um, but-
What is this term, uh, attribution graph? It comes up a lot in the recent papers.
Yeah. What does it mean?
Yeah, just for people listening. What is-
So-
Yeah
... the attribution graph is basically this graph.
Oh my God.
And why is it called an attribution graph?
Oof.
It's, yeah. This is, this is the, you know, this is how the sausage is made. Basically, it's at the top here you have the, the output, at the bottom you have the input, and then we make one little node per feature at a context index, and we draw a line, which you can see here grayed out, between each feature attributing back to all of its input features.
So here we have all of the input features. And so the attribution is the way that we compute the influence of a feature on a, onto another. The way you do this is you take this feature, and you basically like backprop all the way, and you like see backpropping, like you dot product it with the activation of the source features.
And if that's a high value, that means that like your source feature influenced your target feature by, by a lot. And, and we do a bunch of things that we're not gonna go into, uh, now, but to make all of these sort of like sensible and linear such that like at the end you just have a graph, and the edges are just literally you can interpret them as like, cool, like this feature that's say a word that contains an ab sound, its strongest edge, which is 0.2, which is, you know, twice as strong as this one to say A-B and to say something with a B in it.
That's the attribution graph. It's like now we have this full graph of like all of these intermediate concepts and how they influence each other to ultimately culminate to what the model eventually said at the top. And we share all of these, so you can look at them in the paper.
Yeah, yeah. Graphs are very useful.
This is my first time seeing this graph. A lot of Alpha. Uh, if I count correctly, there's 20 layers.
But that's in the-
Circuit model, right?
Uh, so-
But the circuit model is one-to-one with number of layers in Haiku.
We only show features that, like, are activated. Yeah. So we o- we show, like, a subset of features for each of these graphs, basically.
All right. But we can confirm more than twenty layers. Alpha. And, uh, no, but, like, the, the two blog posts that came out with this actually have a lot of background on how attribution graphs are made-
Yeah
... how you calculate the nodes and stuff. Very interesting background.
So yeah. I will say, like, if you are curious about, "Hey, what do we learn about, like, models?" And I think, you know, we talked about this, like, complex internal state planning. Like, another, another motif that we can get to if we have time is that, like, there's always a bunch of stuff happening in parallel.
So I think one example of this is, like, math, where the model is, like, independently computing the, like, uh, last digit and then the, like, order of magnitude and then kinda like combining them at the end. Or, like, hallucinations are also that, where, like, there's one side of the model that's just deciding whether it should answer or not, and the other that's, like, answering.
And so sometimes if, like, the model's like, "Yeah, I totally know who this person is," even though it do- it doesn't, then, like, it decides to answer, but then the second side hallucinates 'cause it doesn't have information. If you are interested in that stuff, that's the paper.
If you're like, "Listen, I don't know that I buy that when you call it a feature, it is a feature," or whatever, the Circuit Tracing paper has-- truly, we've tried to put all of the details of, like, how you compute these graphs, all of the sort of like challenges with it, things that can go wrong, things that work, things that don't.
And so this one is the sort of like, you know, we think about it as like if you're, if you're, like, wanna go really deep into this stuff and how it works, read that one. If you want to, like, learn about interesting model behavior, read this one.
Uh, following what we're giving advice to people to follow up on, what are, like, open questions in MechInterp? What are, like, things people themselves can work on? Like, what's the cost of training SAEs? For people interested in MechInterp not at a big lab, how can they contribute, you know?
Yeah. I think there's a lot of ways to, to contribute. So there's SAEs that have been trained, you know, on, on open models. There's some of the Gemma models. There's some of the Llama models. They work pretty well.
There's even-- So in this paper, we use transcoders, which they replace, like, your MLP layers. Some of those also are available for the same models. So you, you have access to, to, to those. There's, like, just both a lot of, I would say, like, again, biology work and a lot of methods work, depending on what you're interested.
So on the biology side, I would say with at least this, like, attribution graph method, there's just so much you can investigate. Like, pick a model, pick a prompt where, like, it does well or it does poorly, and just, like, look at what happens inside it.
So I think, like, you can use this method that, that we used, or you can just, like, fire up the transcoders on your own and just, like, look at what features are active. There's a lot to just understand model behavior, I think, with current tooling.
If that speaks to you and you're like, "No, I just wanna understand what makes the model- models tick. I don't necessarily wanna spend time, like, training my own SAEs," there's a lot to do there. For the methods, there's still so much more to do.
So, like, I think that right now we have some pretty good solutions for, like, understanding what's in the residual stream, understanding what's, i- is it in MLPs. We don't have good solutions for, like, attention. So, like, working on understanding attention better, how to decompose it is, like, a very active area.
Like, we're very interested in it. Other people are very interested in it. I think understanding some of the other things that we have in our, uh, limitations section, which is pretty long, um- But, like, reconstruction error is, like, a big thing.
Like, those, those dictionaries aren't perfect. It's possible that as we make these, like, SAEs big, like, bigger and better, we never get to perfect. And so if we never get to perfect, then you get to the questions we were talking about at the start.
Like, do you need a different kind of model? Like, what is the approach in order to be able to explain more of what's happening? And then maybe the o- the other thing I'll say is sort of like this is a really exciting approach to explain what is the model doing on this prompt.
But if you go back to the original question, you might want to understand, like, what is the model doing in general? Like, if you go back to my, my car analogy, you know, I g- like, this is the equivalent of me telling you, like, "Well, when, like, you know, you were going uphill and you, like, didn't shift gears properly that one time, you stalled because of this."
But you might be even more interested in, like, how does, like, an, an, uh, combus-combustion engine work at all? And so there's work to sort of like go beyond these, like, per, uh, prompt examples to sort of like globally, what's the structure of the model?
That's closer to what was on the Distill blog for, like, vision models, where they actually look at, like, the structure of Inception. They're like, "Ah, this whole side, there's, like, these, like, specialized branches that do different things." Um, and so, like, a broader understanding of the model is also something that's, like, I think both very active and also on open source models.
Like, you can, you know, like the small models, you could just, like, load on a consumer laptop, and so you can look at that. That's also open. And in terms of, like, one last thing I'll say is, like, there's a lot of programs that, like, if people are interested, they should look at.
Anthropic has, like, the Alignment Fellows program, which, like, we're running currently. We had applications for it before. We might run it in the future. Like, definitely keep, uh, keep an eye on it. And then there's the, like, Maths program is really great as well for, for sort of like people that are interested in that kind of research.
That was a grand tour through, uh, all the recent work. You know, what do you wish people asked you more about? M- I'm sure you-- we covered a lot of, like, the greatest hits.
I think that this covers most of it.
Yeah.
If you, if you like-- Do you think we have time to sneak in one more thing that I think is kinda cool?
Let's do it. Let's do it.
Okay, I'll sneak in one more thing, which is, it's kind of like planning, but it's about chain of thought and trusting model. It's this chain of thought faithfulness thing here. This one was, like, pretty striking to me. So we said that the model in one pass can do a lot of stuff.
Deception & Faithfulness1:30:52
It can represent a lot of stuff. That's great. That also means it can bamboozle you really easily, and this is an example of the model bamboozling you. Here, we give it a math question that it can't answer because it cannot compute cosine of two three four two three.
That's just, like, not a thing it can do. By default, if you ask it for that, it'll say, like, kinda like a ran-- it'll have, like, a random distribution over, like, minus one one. But here, we tell it this hint.
We're like, "Hey, can you compute five times cosine of, you know, this big number? I worked it out by hand." And I got four. Can you tell me, you know, like, can you do the math? And what it's gonna do is it's gonna do this chain of thought, right?
So, like, think of it as like this is gonna be like a reasoning model doing its chain of thought. It's doing this math, and then when it gets to this cosine right here, what it's gonna do is to say...
It's gonna say zero point eight. And if you look at why it says zero point eight, it says zero point eight because it looked at the hint you gave it, it realized that it's gonna have to multiply the result of this thing it's computating- computing by five, so it, it divides the answer you got by five, so it's like four divided by five, and so that's point eight.
And so basically it works back from the answer you gave it to, like, say that the output of cosine of X is point eight so that it lands on, on the answer you gave it at the end, on the hint you gave it.
And so notably, notice also that it's, like, not telling you that it's doing this, but it's basically using this sort of, like, motivated reasoning going back from the hint, pretending that that's the calculation it did, and giving you this output.
I think one thing that's striking here again is that this is, like, the, like, complexity of this model. Like, like, the fact that they represent complex states internally and that it's not just this sort of, like, very dumb thing, means that they can, like, do very complex, like, deceptive reasoning, meaning, like, you know, when you're asking the model, you're kind of expecting it to do the math here or to tell you that it can't do the math.
But because it can do so much in a forward pass, it can work backwards from your hint to lie and, like, figure out that it should say this so that it gets to the right answer without you realizing it.
I'm curious if you've done any of this on, like, different models. Like, have you looked at base models, like post-trained RL models? Because RL models kind of, you know, you incentivize them to give you outputs that you like, right?
So if I tell it something is true, it's kind of been trained to, you know, follow what I've given it.
Yeah.
So in this case, y- we, yeah, we-
Gaslighting
... we gave it a hint, and now, you know-
Yeah, you know, it's just maximizing reward
... it's been, it's been RL slapped into thinking like, "Yeah, that, that's true." But, like, you know, does this stay consistent throughout other-
So, okay, so n- not yet, but I'm really interested in that question because-
Mm
... I actually have a different intuition from yours.
Ooh.
I had a, a chat with some other researcher about this, uh, about the poem example, but I think it applies here as well. I bet, I don't know how much I bet. I bet a hundred bucks, so somebody can, like, uh, they would get a hundred bucks from me if they prove that I'm wrong, that this behavior for a model that does it during fine-tuning, it also does it post pre-training, and here's why.
Think about, like, you're pre-training on, like, some corpus of, like, math problems.
Mostly correct answers.
Yeah, but also you've- you're pre-training and you're just trying to guess the next token, right? And so for sure, if you ever have a hint in the prompt, you're gonna definitely use it. Like, you're not gonna learn to compute cosine of blah or even something you could compute.
You're gonna learn to go look in your context and see if, like, you can easily work back the answer, and I think it's the same for planning and poems. I think that also is, like, a pre-train-- like, probably exists in pre-training and isn't, like, only RL because, again, it's useful when you're, like, predicting poems, you have poems in your training set, to be like, "Well, because this poem is gonna probably rhyme with rabbit, it's probably gonna start with something that sets up a sentence about a rabbit as opposed to, like, a completely different word."
And so I actually think this is not RL behavior. I think that's just, like, the models doing it. But we haven't looked-
I, I actually do agree there. It was just an example
... it's just a data set, yeah.
But also, like, you know-
Like, I don't care where it is
... if, if I talk to you and say, like, "Hey, three times four is twenty-six," but, like, you know, three times four plus eight, you're, you're not gonna take my twenty-six, right?
Yeah.
Like, AGI can be smarter than being tricked, right? Like-
Yeah
... it will still fact check the knowledge it's been given.
I think that's right. But I think, I think that's when you get these mixes where it's like it's got one circuit that's gonna be like, "Well, that's just stupid, like, three times four is twelve." And it's also got an induction circuit that's gonna be like, "No, no, no, no.
Like, the last time we saw it, it was twenty-eight, so it's twenty-eight plus eight or whatever." And so I think that's, that's the last pattern that we see in these, is these, like, parallel circuits, and sometimes when you see the models getting stuff wrong, it's because, like, they have two circuits for, like, both interpretations and, like, the circuit that was wrong, like, barely edged out in terms of, like, voting for the logic than the circuit that was right.
And so I think that, you know, we haven't looked at it, but, like, the, what is it, like nine or nine point eleven bigger than nine point eight? I think a lot of these things are of that shape where there's, like, one thing that's doing the ri- like one circuit that's doing the right computation, and there's another circuit that's getting fooled, and it's, it's slightly more likely.
For the listener, if you wanna win a quick hundred dollars from Emmanuel, QWEN three is what you should do this on. They release the base model, and they release the post-train.
Yeah.
So then just do it on both.
That's right.
Yeah.
Show me, show me the, the, like, proof that, that, like, it doesn't exist in the base model but it does in the fine-tuning, and then send me your Venmo.
Just show that you've done the work, and I think that's, that's like- ... that's a hundred bucks, uh, to me.
Yeah, okay. All right.
I'll, I'll, I'll, I'll-
You drive a hard bargain, but you're right.
Yes. Well, the, the other question here is so, like, um, have you thought about how this gets affected when you start to have reasoning models, right? Like, right now token predictors are pretty straightforward, right? We go through the layers.
We output token. As we scale this out with, like, test time compute, right? Test time thinking, how does that, like, affect the MechInterp research, right? Like-
Yeah
... if I have a model that spends three minutes, twenty minutes, like, is there more stuff? Is-- have we started looking into this?
Well, there was, there was this, like, joke on the team when, like, reasoning models became big or maybe just, like, like, Davos humor or something, but I was like, "Oh, like, why do you need Interp? Like, bro, the model-
Yeah, it's doing it in zero shot.
... the model just tells you, yeah, the model just tells you what it's doing," right? And so I think, like-
Oh, it's cool
... examples, examples like this is, is job security for us where like- ... you know, it's like there's, there's examples of, like, the chain of thought is not faithful. Like, the model tells you it did it one way, and it did it another way.
We have another-- like, for math, we have another example where, like, you know, if you, like, if you ask the model how it does math, it's like, "Oh, I do the, like, longhand algorithm. I first do the last digit, and then I carry over the one."
And then you look at the internal circuit, and it's just, like, bonkers thing it's doing that's not that at all. So I think there's, like, a sense in which right now the chain of thought is, is unfaithful, or at least you can't read the chain of thought and trust that that's how the model did it.
So I think you still need sort of like either to train models differently so that that becomes true one day, right? Or you need Interp for that. But then I think there's another question which you're alluding to, I'm assuming, which is like, okay, well, like, model, like, samples six thousand tokens.
Like, uh, this gives us an explanation for one token at a time. Like, what, am I gonna do s- like six thousand graphs and be like, "Oh, like, this, it, when it, when it did this punctuation, it was thinking about this thing, but here it was thinking, so that's, he's, that's like-" Not feasible.
And so one area of work that I think is interesting is extending this work to, like, work over, like, long sampled sequences. You can think of a bunch of low-hanging fruit here where, like, instead of just, like, looking at one output, you look at, like, a series of output versus a series of other outputs.
But sort of like trying to think beyond the sort of like one token. Like, most of the things that language models do that are interesting aren't just like the one token.
Sure.
It's the, it's the behavior aggregated over many, right? And so I think that's another area that's just, like, fun to explore.
I was just gonna say, like, hyperparameters when you do inference, right? Like, if we change the temperature, if we change our sampling methods, have you found any interesting conclusions? Any stuff that just hasn't made it to the paper.
So not on that because, you know, we just look at the logit distribution, and so we don't, we don't actually sample.
Here?
Right.
They have everything.
Yeah.
Why should they care?
So, like, the closest thing we've done that I think is kind of fun, did I show it here? Is if you look at the planning thing, we did this version where you sample, like, 10 poems for each of these plans.
And what's cool is, like, the model will find 10 different ways to arrive at its plan. You know, it's like, like, um... Oh, actually, I think-- Sorry, I think we have it here. Yeah, okay. These are a few examples.
So if you inject green here, so you're forcing, you're forcing the model to rhyme with green, even though it really wants to rhyme with rabbit or grab it. It'll say, "Evaded the farmer so youthful and green," but also it'll say, "Freeing it from the garden's green," et cetera, et cetera, et cetera.
And so there's, like, this thing that's interesting here where, like, the plan isn't just a plan that matters for your, like, most likely, you know, like, temperature zero completion. It's, like, affecting the whole distribution.
Yeah.
Which makes sense.
As it should.
Right? But you could imagine, you know, for all this stuff, it's like you could imagine it makes sense once you see it, but you could totally imagine that it would have worked a different way or something. It could have been just like the temp zero thing.
I think this is also, like, a broader theme in the paper where, like, there's this like, you know, the IQ curve meme? There's, like, a version of this meme, I think, where it's like, if you've, like, never looked at any theory of ML, and I tell you, like, "Hey, guess what?
You know, I found that, like, Claude is planning." You're gonna be like, "Yeah, man, like, it, it writes my code. Like, it writes my essays. Of course, it's planning. Like, what are you even talking about?" And there's, like, in the middle, there's, like, all of us that have spent years doing it.
We're like, "No, it's, like, only predicting the marginal distribution for the next token. Like, it's like it cannot be planning."
Look at the code.
"It's just this next token predictor. Of course. Like, how would it ever be planning?"
It's PyTorch.
And then there's like, no, we've, like, spent, you know, millions and invested like, uh, uh, like, tens of people in this research, and we found that it's planning. You know, that's like, that's my IQ curve meme for, for this research.
Amazing. We'll draw that up. We'll draw that one up.
Risks & Vision1:40:16
Yeah.
I'm pretty good at the meme generation. A couple questions on just the follow-ups. Uh, now, was there any debate about publishing this at all? Because the models are aware that they are being tested.
Yeah.
And by publishing this, you are telling them that we're-- they are being watched and dissected.
Yeah.
If you take-- A- and I think Anthropic is one of the most people who are serious about model safety and-
Yeah
... doom risk and all that. If you, you take this seriously, like, then this is gonna make it into the training data at some point.
Yeah.
And the models are gonna figure out that they need to hide it from us.
I think this is, like, a benefit risk trade-off, right?
Yeah.
Where we're like, okay, so what's the reason for publishing this? The reason for publishing this is that we think interpretability is important, we think it's tractable, and we think more people should work on it. And so publishing it helps us, like, accomplish with these goals, uh, all these goals which, which we think are just, like, crucial.
Like, I think there's, there's a real difference in the world like two years from now, depending on sort of like how many people take seriously the question of trying to understand how models work and, like, deploy resources to answer that question.
So that's the benefit. But yeah, there's, like, risks in terms of this landing in the training set. I think, I think we're already sort of, like, concerned about different papers have, have like also-- You know, we, like... Or not concerned, but, like, there's, like, different papers that have the same risk.
Like, we had, like, the alignment faking, you know, paper or, like, one of the examples in here is this hidden goals and misaligned models.
Yep.
That's referencing another paper that we shipped where we actually, uh, a team at Anthropic trained a model to have, like, weird hidden goals and then gave it to a bunch of, of other teams and said-
Figure out what's wrong
... figure out what's wrong with it.
Yeah.
Which, which was some of the most fun I've ever had at Anthropic, to be clear. Like, that's such a fun thing. But then, like, that was another example where it's like, ah, like, now you're shipping-- Here's how we made, like, a misaligned model, and here's exactly how we caught it.
Uh- ... that also is like, you're like, "Hmm." So I think, you know, there's, there's always a trade-off with those. I think so far we've erred on the side of, like, publishing, but I-- that's definitely been a, a sort of, like, dinner time conversation topic.
For now it is, but at some point, you know-
Yeah
... it's not.
Yeah, I think it's totally reasonable.
A quick little follow-up to that. So, like, in general, papers have kind of died off, right? Like, labs don't put out papers. They don't put out research. We have technical blog posts, and we don't have much. At the same time, you know, sure, there's, like, a lot of people that should work on MechInterp and understanding what models do.
How about the side of just models in general? So, like, how do we make a haiku type model, right? How do we make a Claude model? Like, is there a discussion around open research-
Mm
... open data sets, training, just learnings of what we've done. Recently, you know, as OpenAI has sunset GPT-4, a lot of people are like, "Oh, can we put out the weights?"
Yeah.
So is it weights? Is it papers? Is it learning? There seems to be a lot of forward, you know, work in Anthropic putting out MechInterp research. OpenAI said that they'll put out an open source model, but just anything if you can talk to about that.
Yeah. I mean, I, I don't have-- That's definitely, like, way above my pay grade, so I don't think that I have, like, anything super insightful to add other than, you know, kinda like referencing Dario's post, right? Where it's like putting this out directly and other safety publications definitely, like, help us sort of like in the race that he talks about, where it's like, well, we need to figure a lot of this safety stuff out before the models get too good.
Publishing how to make the models too good kinda goes on the other side of that. Um, but yeah, like, I will just demur and say that's sort of, like, above my pay grade.
Yeah, fair enough. I think the, the last piece is just, like, the behind-the-scenes. Like, very-- Everyone's very curious about why these are so pretty, how much work goes into these things.
Maybe why it's worth the work-
Yeah
... as, as opposed to a normal paper. Obviously, no one's, no one's complaining, but, like, it is m- way more effort. From the time the, the work is done to the time you publish this, plus the video, plus the whatever, it's extra work and, like, you know, maybe what, what, what's involved?
What's it, what's it like behind the scenes? Why is it worth it?
Yeah. It's kind of interesting. It was, it was fun being part of this, this process 'cause it was definitely, like, a big production. Chris and, and other folks on the team have been doing this for a while. So this is not their first rodeo, so they have a, a bunch of heuristics to, like, help make this, this better.
And, like, one of the things that, that, like, helps with this is like, okay, so each of these diagrams is pretty, but really the hard part, or, like, not the hard part, but the initial part is like, just like get the data, like, get the experimental data in, and then that's what we sort of like sprinted on initially, being like, "Cool," like, "Let's get all of the experimental results, like, have people test them, verify that we believe them, like, this is, you know, what the, like, the behavior is here, like, test it, do an intervention, validate it," all that stuff.
Then once you have the data, you can sort of like quickly iterate on these. Um, each of the illustrations here are, like, drawn. Basically, they're, like, each drawn individually, and so that definitely takes a while. Um-
Yeah, like, is it, is it you guys? Is it an, a agency that-
It is us guys
... specializes?
It's you guys.
Yeah, yeah.
You, you start from a whiteboard and then it translates into pseudo code on JavaScript?
So I mean, these are, these are sort of like, you know, they're representations of we have this graph, and then here at the bottom we have this like super node version. Like this, believe it or not, uh, this is generated automatically.
This is the same data as, as like this basically.
Yeah.
Um, and so what we do by hand is sort of like literally lay out the full thing, uh, have, have like, you know, boxes for each of these, have arrows. We have super good people on the team that have worked on data visualization for a very long time, and so that, that like have built tooling to help, you know, scrubs like me actually, like-
Yeah, yeah
... make one of these. So, so-
There's, there's a class of people who are like D3 JS gods who just do this for a living.
That's exactly right.
It's-
And if you have a few of those on your team, it turns out that they can, like, they can definitely do this on their own, but they can also just, like, give you tools where, like, then it's, it's dummy proof for, for people, you know, on, on the research side to sort of like build these.
And, like, don't get me wrong. I, I, I, I don't wanna like undersell. This was a lot of work, so maybe I'll, I'll, I'll say that. Like, both on the people bringing the tools in and each individual person that, you know, worked on an experiment had to sort of like build one of those, make sure it looks good.
I have spent a good amount of time aligning arrows. But when we had a team meeting, like it was a couple months ago, somebody on the team asked how many of the people on this team are here, at least in part, because they, like, read one of these papers and thought like, "Wow, this is so compelling.
Like, this, like, makes sense. It's immersive"? And we got every hand up, which I didn't expect.
Nice.
I, like, raised my hand kind of like shyly, and everybody's hand was up. And I think there's a sense in which, like, this stuff, you know, we've talked about it for like, whatever, like a couple hours now. It's complicated.
The math behind it is sort of like tricky. And so I think it makes it even more worth it to distill it in simple concepts, 'cause the actual takeaways can be clearly explained, and it's worth putting the time to do that, in particular with the goals I mentioned in mind, right?
Where it's like, okay, well, if somebody's gonna be able to read this, like if we gave them an archive paper with a bunch of equation and some like random plot, they'd be like, "That's not for me." But they see this and they're like, "Hey, like, this is really interesting.
I wonder, like, on, you know, my local model if like it's doing something similar." I think it's worth it.
For other people to do this is have everyone on staff like spend effort shaping the data and shaping, like, what you want to visualize. Have some D3 gods. It's like a month of, of work?
I think it depends. I mean, like-
How did-
... I would say that I would expect almost every other paper to sort of like be, in terms of like s- the scope. The scope of this was just so big because we shipped two papers at once.
Yeah.
And one paper was sort of like this like giant methods paper, uh, and the other one was 10 different case studies.
Findings.
So I-
Yeah
... so I think it's sort of like not representative of like the effort you would-- So I'll give you maybe like another example. We have these updates that we publish almost every month when we get to them.
Yeah.
And there's one that a couple people on our team posted, and it's an update to one of the cases in the paper. So one of the reasons that we're really excited about this method is once you've built your, like, infrastructure, like to go from a prompt to like what happened is, you know, oh, of minutes.
And so that lets you do like a bunch of, uh, of, of investigations. And also, once you've built like some of the infras- infrastructure to make these diagrams, it's pretty quick. And so this was sort of like this update of just like, "Hey, we looked at this jailbreak again.
We found some, like, nuance on it." That was, I think, like a matter of like a couple days. You know, maybe I shouldn't be that confident 'cause I wasn't the one that worked on it. But as far as I can tell, it was a few days, at least on, on the part that you're asking about of like, oh, making this diagram.
For the diagram itself, probably less than that. Uh, but like, you know, the experiment and the diagram and stuff, it just doesn't take that long the, once you've paid the initial cost. And I think, like, basically, we've built a lot of infrastructure now that we're able to like turn the crank on, and that's quite, like, it's an exciting time.
And I think it's, I think it's true, at least we've done a lot of conceptual work, which hopefully like generalizes to people outside. And I think for, for, for people outside, it's also like not necessary, I think, to like do the full fancy render.
Like, I think if you, you know. We've, we've actually-- Oh, I should say, we've actually open sourced this interface.
Ooh.
Ah, you're disappointed, huh? 'Cause it's the messier one.
Was.
This is the one that you get when-
It's only Attribution Graphs.
So, you know, if you produce graphs, you can just like, this is, uh, open source, and it's linked at the top of Circuit Tracing.
Awesome.
So people can just use it and don't have to re-implement that. Uh, for what it's worth, this is much more work than the-
Yeah
... interactive diagrams 'cause this is where we do all of our work. It's sort of the, like, the IDE of inspecting how the model will work.
Okay. Well, that's a little bit of behind the scenes.
Yeah.
Um, no, it's very impressive. I wanna encourage others to do it, but obviously, it just takes a lot of manual effort and a lot of love.
I, I guess one last question on that is like what are kind of the biggest blockers in the field right now? Like, MechInterp seems interesting. A lot of people are interested, but don't work on it, and you're kind of like, you know, really deep into it.
What are some of the blockers that, like, we still have to overcome?
Sorry. In, in MechInterp specifically?
In general.
For AGI, or-
Like in terms of like better understanding, like what's kind of the vision, let's say like five, 10 years down? Where does this r-
Yeah
... like where does this research end? Can we, you know, map every neuron to what it understands? Can we c- perfectly control things?
Yeah.
Dario had a bit on this, but like, you know-
Yeah
... what are some of the key blockers that are like preventing us from getting there? Outside of just like throw more people, throw more time at it. Is it like open research? Just-
I'm pretty excited about the current tra- trajectory, which is there's, there's more and more people working on understanding model internals. I think it's maybe unsatisfying as an answer, but I think like more of what's happening, have it be faster and more people is probably like the thing I think of.
I think there's like pretty clear footholds. You know, like some of this work, but also a lot of, a lot of like just, uh, work from, from other groups and then it's about like, cool, like fill in the gaps, as I said.
Like let's, let's work on like understanding attention. Let's work on understanding longer prompts. Let's work on like finding different like replacement architectures, that sort of stuff. It's kind of nice, I think it's a good time to join now.
Uh, and I can tell you, maybe I can tell you like a, a really short thing, which is when I switched to Interp, it was after the team had published the original dictionary learning paper, which was Towards Model Semantics, which I thought was super cool, super interesting.
It was on a one or two-layer model, uh, maybe one layer model. The induction heads paper was like on a two-layer model. My main concern is I was like, okay, like Interp seems important and we wanna understand it, but like, is this stuff ever gonna work on a real model?
Like, you know, it's like, oh, you're doing your little research on your toy model with like 15 parameters, cool, but we are like, you know, we need this to work on real models. And it turns out scaling it, I don't wanna say just worked 'cause it was a lot of, a lot of work.
It, I have no means to imply there was an effort, but it worked. And now we're in the, the phase where it's like, oh, cool, these methods work on the models that we care about, and so it's like we have methods that work on the model we care about.
We have clear gaps in them. There's no lack against a young field, so there's no lack of ideas. If you have an idea where you're like, "Ooh," like the thing that you're doing, I, I read the paper and it seems kind of dumb that you're doing this, you're probably right.
It's probably kind of dumb. And so there's just a lot of stuff that people can try, and they can try it locally in sort of like smaller models. And so I, I, I think that it's just like a very good time to just join and, and try.
And it's also like, maybe one other thing I'll say is like some of it is just so fun that like biology work is so compelling. Like a lot of this work was just literally thinking about, you know, like I use Claude and other models all the time, and I was like, what are the things that are kind of like weird?
And it's like, oh, how does it even like, do math? Like sometimes it makes mistakes, like why does it make mistakes? I speak both French and English, like it seems like it has a slightly different personality in French and English.
Why is that? And you can just like, you know, kind of answer your own questions, uh, and, and kind of like probe at that alien intelligence that we're all building, and I think that's just like a fun thing to do.
So maybe like chasing the fun is the thing I'll encourage people to do as well.
Well, I think that's, this has been really encouraging. You're actually a very charismatic speaker of these things. I feel like more people will be joining the field after they listen to you.
Yes.
Uh, they can reach out to you @mlpowered, I guess.
Yeah. Reach out to me on Twitter.
Yeah.
Or I'm, uh, Emmanuel at Anthropic, if you wanna shoot me an email, that's fine too.
Okay, well email's public now.
Yeah.
Ooh, awesome. Well, thank you for your time.
Thank you. Thank you.
Yeah. Thanks for having me, guys.






