AI Frontiers, part 33: Mechanistic interpretability — looking inside the black box
Part 33from the AI Frontiers series · 65 parts in all
There are two ways to ask what a model is doing. The first is behavioral: give it inputs, watch the outputs, and build a theory that predicts both. Almost everything in this series so far has been about that approach, and it works well enough to ship products. The second is to open the network and identify the actual computation — which components read which information and transform it into what — and that is what mechanistic interpretability attempts.
It is the most scientifically interesting work in the field and the least obviously useful, which is why it has spent years being simultaneously oversold and underfunded. By late 2023 it had produced a handful of results that are genuinely load-bearing — a complete circuit for a specific task, a theory of why features share space, and a method for finding individual interpretable directions inside a network — and those results are worth understanding on their own terms rather than as a metaphor about brains.
The framework, and why it is the right starting point
Anthropic's mathematical framework for transformer circuits (Elhage et al.) did something unusual for the interpretability literature: it actually wrote down the architecture in a form that can be analyzed. Rewriting attention as a sum of paths through the network, and observing that the attention output at each position is a linear combination of values weighted by query-key similarity, gives you a calculus. Residual stream as a shared communication channel. Attention heads as operations that read from it and write back. MLP layers as key-value memories. QK circuits and OV circuits as the two functions a head performs.
Two consequences followed. First, the residual stream framing reframes the whole network as a set of components that read and write to a common workspace, which makes it meaningful to ask what a specific head is for in terms of information flow. Second, it explains a fact that matters for everything downstream: because the stream is a vector sum of everything written to it, components can communicate through a superposition of features, and pulling any one feature out requires disentangling what has been added together.
The second foundational result is the explanation of what superposition is doing and why. Toy Models of Superposition (Elhage et al.) trained tiny networks to represent more features than they have dimensions and found that the networks do it, by using nearly-orthogonal directions, accepting interference as the price. Critically, they found the behavior is sensible rather than pathological: the network allocates clean dimensions to features that are frequent or important and superposes the rest, in a way that resembles a compression code. The picture this gives of a large model is not a black box in which nothing can be read, but a compressed representation whose compression scheme would have to be inverted to read it.
The results that turned interpretability from a story into a science
Three specific findings, in ascending order of ambition, established that the program could produce real answers.
Induction heads. Olsson et al. identified a two-head mechanism that implements a simple but powerful pattern: a previous-token head writes the identity of the token before a given position, and an induction head learns to look for the same token elsewhere and predict what followed it. That is copy-what-came-after-this-before, and it accounts for a large share of in-context learning. Under the account in part 24 of why demonstrations work, this is the mechanism: the prompt is an index, and induction heads are the lookup.
A complete circuit for indirect object identification. Wang et al. went further and traced an end-to-end mechanism in GPT-2 small that resolves sentences like "when Mary and John went to the store, John gave a drink to ___". They identified the head that duplicates the repeated name, the heads that inhibit it, and the heads that promote the alternative — a composition of about twenty-six attention heads whose roles are individually interpretable and whose ablation produces exactly the predicted errors. This remains the strongest evidence that a real computation in a real model can be understood at the level of components, and the reason it is cited so often is that it came with causal tests rather than visualizations.
Grokking, explained. Nanda et al. took a small model that memorizes a modular arithmetic task and then, long after training loss flattens, suddenly generalizes, and showed what changes underneath: the network builds a discrete Fourier transform of the inputs and computes the answer by rotating in that space. The transition from memorization to a general algorithm is visible as the gradual emergence of a structured representation, and the generalizing solution can be identified, verified by interventions, and predicted. It converted a mysterious loss-curve kink into an engineering story about circuits forming.
The superposition problem, and the method that addressed it
If features are superposed, then individual neurons are not the right unit of analysis — a fact worth emphasizing because a great deal of early interpretability work, including OpenAI's attempt to have a model explain individual neurons in natural language (Bills et al.), was built on the assumption that it was. A single neuron responding to many unrelated concepts is not a malfunction; it is what a compressed representation looks like.
The insight that broke the impasse is that superposition is invertible under a sparsity prior. If the network represents many features in fewer dimensions, and only a small number are active at once, then the right tool is a sparse dictionary: learn an overcomplete set of directions such that any activation is a sparse combination of them. Sparse autoencoders had been suggested for exactly this (Sharkey et al.; Cunningham et al.), and Anthropic's monosemanticity work (Bricken et al.) turned it into a result — extracting thousands of features from a one-layer transformer, many of which activate on cleanly interpretable concepts, with a measurable cost in reconstruction quality that quantifies how much structure was lost.
The significance is that this is a method rather than an anecdote. It produces a decomposition with an evaluation criterion, it scales, and it changes the unit of analysis from the neuron to the feature — which is the unit that causal interventions can then be performed on. Earlier work on identifying the weights responsible for factual associations (Meng et al., ROME) and then editing many of them at once (Meng et al., MEMIT) had shown that targeted, localized modification was possible; dictionary learning supplied a principled way to find what to target.
What interpretability is not, and why the distinction matters
Several things get described as interpretability that are not the same activity, and conflating them has cost the field credibility.
Attention visualizations are not explanations. A heat map showing which tokens a head attended to is an input-output correlation with no causal content, and the literature has documented repeatedly that attention weights can be perturbed without changing behavior. It is a debugging aid and a source of hypotheses, not evidence.
Natural-language self-reports are not explanations either, for the reasons in the RLHF entry: a model's account of its reasoning is generated text that may or may not correspond to the computation. This is not a claim that models are deceptive. It is that the explanation and the mechanism are produced by different processes, and nothing forces them to agree.
Feature visualization in the vision-model tradition has the same status: a compelling image of what maximizes a neuron's activation is evidence about the neuron's input distribution and not about its role in the network.
The discipline that separates real interpretability from these is causal testing. A circuit claim is a claim that removing a component produces a specific, predicted degradation, and that the prediction was made before the intervention. The methods exist — activation patching, path patching, ablation, causal scrubbing (Chan et al.) — and the honest papers use them. A paper that ends with a heat map has not established anything, which is why the field's own literature is more modest than its public reputation.
What scaling the method would actually require
The demonstrations are on models with millions or billions of parameters, and the interesting ones are on models small enough to analyze exhaustively. The path to frontier models runs through four problems, and it is worth being explicit about them because the field's optimism often skipped this part.
Sparse autoencoders scale in cost. Training a dictionary large enough to capture the features of a big model means training another model of comparable size, which means the explanatory apparatus costs as much as what it explains. Feature extraction is not a one-time expense either, since each model revision requires redoing it.
Features compose into circuits, and circuits are harder. Identifying a clean feature is step one. Explaining how thousands of features interact across layers to produce an output is a qualitatively different problem, and the automated circuit discovery work (Conmy et al.) was an attempt to make the search tractable rather than manual.
Interpretability does not scale linearly with size. The GPT-2 result depended on reverse-engineering a mechanism by hand with human intuition about what the task required. It is not obvious that the same approach finds the equivalent structure in a model trained on trillions of tokens, where the number of interacting features grows far faster than the number of layers.
And the target moves. Frontier models change architecture, training recipe and post-training every few months. An interpretation method that requires a year of work per model is an archaeological practice rather than an engineering one, which is why the field's more pragmatic branch focuses on methods that generalize — dictionaries, probes, causal patching — rather than on any particular circuit.
Evaluating an interpretability claim
Because the field is easy to oversell, it helps to have criteria for reading a claim, and the interpretability community has largely converged on them.
A good claim is specific: it names components and predicts what happens when they are removed, rather than describing a correlation. It is causal: the prediction is tested by intervention, and the prediction was registered before the test. It is complete with respect to some behavior, in the sense that the described mechanism accounts for most of the effect rather than a sliver of it. And it is usable, meaning someone else can apply the method to a different model and get a result.
By those criteria, the four results this entry opened with hold up and most of what is described as interpretability in popular writing does not. That is not a criticism of the popular writing so much as an observation about why the field advances slowly: the standards are high, the objects of study are enormous, and the honest result is usually a single well-understood mechanism extracted at considerable cost. That is what a science looks like before it has instruments, and the instruments are what the next few years are for.
Why it matters, and the case for skepticism
The argument for doing this work is not that it will make models better, though it occasionally does — identifying a mechanism lets you remove a failure mode at its source rather than prompting around it. It is that every other method of controlling a model's behavior operates on inputs and outputs, and inputs and outputs are not a complete specification of a system that can learn things its developers did not intend.
The concrete places interpretability has already paid off are worth naming. Detecting whether a model is using a feature it should not, like a proxy for a protected attribute. Understanding which component produces a specific refusal or failure, rather than trying increasingly elaborate prompts. Establishing that a model's stated reasoning is not its actual computation, which is a safety claim in itself. And finding the representations that enable targeted editing, which is the mechanism behind everything from unlearning to steering.
The case for skepticism is that the gap between the demonstrations and the object of interest is enormous. Twenty-six heads in GPT-2 small is nine orders of magnitude away from a frontier model, and the honest answer to "will this scale?" in late 2023 was that nobody knew. Scaling interpretability is a research program, not a solved problem, and the failure mode would be a large body of beautifully labeled features for a model nobody deploys.
What keeps me interested is the same property that keeps the field interesting. Mechanistic interpretability is the only approach that tries to answer the question of what a model is doing by looking at the model rather than at its behavior, and the field keeps providing evidence that the two answers differ. A system whose internal computation diverges from its stated reasoning is a system you cannot supervise by reading its transcripts, and if that is the world we are building, the instruments for the internals are not a luxury.
Works Cited
Bills, Steven, et al. "Language Models Can Explain Neurons in Language Models." OpenAI, 2023, openai.com/research/language-models-can-explain-neurons-in-language-models. Accessed 2 Nov. 2023.
Bricken, Trenton, et al. "Towards Monosemanticity: Decomposing Language Models with Dictionary Learning." Transformer Circuits Thread, 2023, transformer-circuits.pub/2023/monosemantic-features. Accessed 2 Nov. 2023.
Chan, Lawrence, et al. "Causal Scrubbing: A Method for Rigorously Testing Interpretability Hypotheses." Alignment Forum, 2022, www.alignmentforum.org/posts/JvZhhzycHu2Yd57RN. Accessed 2 Nov. 2023.
Cunningham, Hoagy, et al. "Sparse Autoencoders Find Highly Interpretable Features in Language Models." arXiv, 2023, arxiv.org/abs/2309.08600. Accessed 2 Nov. 2023.
Dai, Damai, et al. "Knowledge Neurons in Pretrained Transformers." arXiv, 2021, arxiv.org/abs/2104.08696. Accessed 2 Nov. 2023.
Elhage, Nelson, et al. "A Mathematical Framework for Transformer Circuits." Transformer Circuits Thread, 2021, transformer-circuits.pub/2021/framework/index.html. Accessed 2 Nov. 2023.
---. "Toy Models of Superposition." Transformer Circuits Thread, 2022, transformer-circuits.pub/2022/toy_model/index.html. Accessed 2 Nov. 2023.
Li, Kenneth, et al. "Emergent World Representations: Exploring a Sequence Model Trained on a Synthetic Task." arXiv, 2022, arxiv.org/abs/2210.13382. Accessed 2 Nov. 2023.
Meng, Kevin, et al. "Locating and Editing Factual Associations in GPT." arXiv, 2022, arxiv.org/abs/2202.05262. Accessed 2 Nov. 2023.
---. "Mass-Editing Memory in a Transformer." arXiv, 2022, arxiv.org/abs/2210.07229. Accessed 2 Nov. 2023.
Nanda, Neel, et al. "Progress Measures for Grokking via Mechanistic Interpretability." arXiv, 2023, arxiv.org/abs/2301.05217. Accessed 2 Nov. 2023.
Olsson, Catherine, et al. "In-Context Learning and Induction Heads." arXiv, 2022, arxiv.org/abs/2209.11895. Accessed 2 Nov. 2023.
Sharkey, Lee, et al. "Taking Features Out of Superposition with Sparse Dictionary Learning." Alignment Forum, 2022, www.alignmentforum.org/posts/z6QQJbtpkEAX3Aojj. Accessed 2 Nov. 2023.
Wang, Kevin, et al. "Interpretability in the Wild: A Circuit for Indirect Object Identification in GPT-2 Small." arXiv, 2022, arxiv.org/abs/2211.00593. Accessed 2 Nov. 2023.