In brief

LLMs produce impressive results, but no one really knows why. Mechanistic interpretability is the research field that seeks to open this black box — not to observe outputs, but to understand the internal computations, step by step. Progress since 2021 is real: today we can isolate “circuits” responsible for specific behaviors, identify concepts encoded in activations, and even modify a model’s behavior by directly acting on its internal representations.

In short: a neurobiologist dissecting a brain does not ask “why is this thought correct?” — they ask “by which neural pathway was this thought produced?”. Mechanistic interpretability adopts the same posture: not judging an LLM’s outputs, but mapping the circuits that produce them. It is computational anatomy, not psychology.


The black box that fascinates and unsettles

An LLM transforms tokens into probabilities through hundreds of layers of matrix computation. This mechanism is perfectly deterministic and entirely visible — every multiplication can be inspected. Yet it is practically incomprehensible: billions of parameters, nonlinear interactions, and no engineer “programmed” French grammar or knowledge of capital cities. These properties emerge from training.

This is where interpretability comes in. The goal is not to understand why the model says something correct or incorrect — it is to understand the causal mechanism that produces the response. Like a biologist dissecting an unknown organism to map its organs.


Circuits: sub-networks with a function

The circuits approach rests on a simple idea: within a neural network, certain subsets of components (attention heads, MLP layers) work together to perform a specific task. These subsets form a “circuit” that can be isolated, analyzed, and tested.

The starting point is a foundational paper published by Anthropic in 2021, which formalizes transformers as a sum of separately analyzable functions. The authors distinguish the QK circuit — which determines where each attention head looks — from the OV circuit, which determines what it produces as output. This framework makes the network partially traceable.

In 2022, a concrete application confirmed the approach’s relevance: researchers identified induction heads, a two-layer circuit present in nearly all transformers, responsible for the ability to reproduce patterns already seen in context. This mechanism, simple to describe, explains a large part of the behavior known as “in-context learning” — the ability of a model to adapt to examples provided in the prompt.


The problem of polysemanticity

Circuit analysis quickly hits an obstacle: individual neurons are polysemantic. A given neuron activates for semantically very different contexts — legal text, DNA sequences, Hebrew pronouns. This is not a design flaw: it is a mathematical consequence of compression.

A model with 10 billion parameters must represent far more than 10 billion distinct concepts. The solution that networks spontaneously discover: superimpose multiple concepts in the same directions of the activation space. This is the superposition hypothesis, formalized in 2022 (Elhage et al., arXiv:2209.10652). Rare concepts can coexist in the same neurons with little interference, since they rarely activate simultaneously.

This compression makes neuron-by-neuron interpretation nearly impossible. A different approach is needed.

In short: a neuron that activates simultaneously on SQL code, biblical references, and Hebrew pronouns is not confused — it is efficient. The network “stores” multiple concepts in the same drawer, betting that they never activate at the same time. It is a form of intelligent compression: 10 billion parameters for far more than 10 billion concepts.


Sparse autoencoders: separating superimposed concepts

The response proposed in 2023 by Anthropic’s team is elegant: if concepts are superimposed in a low-dimensional space, they can be separated by projecting them into a much higher-dimensional space, forcing a sparse representation.

This is the principle of the sparse autoencoder (SAE) applied to dictionary learning. An auxiliary network learns to decompose the LLM’s internal activations into a sum of vectors, with the constraint that few vectors are active simultaneously. Each vector then corresponds to a “feature” — a direction in the activation space associated with a specific concept.

Applied to layer 6 of GPT-2 Small with a 16x expansion, the result is striking: approximately 15,000 distinct features, of which 70% correspond to a concept identifiable by humans. Features for Arabic script, DNA sequences, HTTP requests, Hebrew text. Comparable results were independently obtained by other teams (Cunningham et al., arXiv:2309.08600).

In 2024, the approach was applied to Claude 3 Sonnet, a production model. Tens of millions of features were identified, multilingual and multimodal. Among them, features associated with deception, sycophancy, and dangerous content. And a key demonstration: by artificially activating certain features (feature steering), one can modify the model’s behavior in a predictable way.

In short: a sparse autoencoder is a dictionary that turns an internal activation (a dense 4,096-dimensional vector) into a combination of a few concept-words chosen from a 65,000-word dictionary. The “few active words” constraint forces separation: each word acquires a distinct, human-verifiable meaning. On GPT-2 Small, 70% of the 15,000 learned features correspond to an identifiable concept.


Tracing computation steps: attribution graphs

In early 2025, Anthropic took a new step with attribution graphs: graphs that partially reconstruct the computational steps a model follows between a prompt and its response. The method is applied to Claude 3.5 Haiku.

The results are unexpected. In some translation cases, the model appears to operate in a shared conceptual space across languages before producing output in the target language — a form of intermediate “language of thought.” In some arithmetic reasoning cases, no trace of an actual computation is detectable: the model appears to guess the correct answer, then retrospectively reconstruct plausible intermediate steps.

These observations are not definitive conclusions — the companion paper “On the Biology of a Large Language Model” explicitly adopts an exploratory stance, treating the LLM as an organism whose physiology one seeks to map without guaranteeing completeness.


What the field does not yet know

Enthusiasm is tempered by serious limitations.

Scalability remains an open problem. Current techniques are difficult to apply to models with hundreds of billions of parameters without massive automation. Human annotation of features, manual circuit tracing — these processes do not scale (Bereska & Gavves, arXiv:2404.14082).

Are interpretations truly causal? A result published in 2025 shows that SAEs find “interpretable” features even in randomly initialized transformers, without any training (arXiv:2501.17727). This raises questions about the significance of discovered features: do they reflect genuine learning in the model, or simply structural properties of the architecture?

The link with safety remains to be demonstrated. The promise is strong: if we understand the circuits responsible for a problematic behavior, we can detect or correct them. But no published result yet demonstrates that a mechanical intervention at the scale of frontier models resolves a concrete safety problem. A 2025 paper (arXiv:2506.18852) argues that the field lacks rigorous conceptual frameworks for connecting computational mechanisms to normative objectives.

The unit of analysis is contested. Neurons, attention heads, directions in activation space, SAE features — each decomposition reveals certain aspects and obscures others. The review “Open Problems in Mechanistic Interpretability” (2025), signed by more than 30 researchers from Anthropic, Google DeepMind, MIT, and other institutions, explicitly identifies the absence of a unified theoretical framework as a major obstacle.

In short: mechanistic interpretability is to AI today what microscopy was to biology in the 19th century: structures are being discovered, some are named, but there is not yet a unified framework linking microstructure (neurons) to behavior (safety, alignment). The promises are strong, the tools improve, but the science is still under construction.


Key takeaways

  • Mechanistic interpretability seeks to understand the internal computations of LLMs, not just their outputs — using methods inspired by biology and reverse engineering.
  • “Circuits” are sub-networks responsible for specific behaviors; the “induction heads” identified in 2022 partly explain in-context learning capability.
  • The superposition of concepts in the same neurons complicates analysis; sparse autoencoders (SAE) enable their separation by projecting activations into a higher-dimensional space.
  • Feature steering — artificially activating a feature to modify behavior — is a promising technique but whose robustness at scale remains to be established.
  • Limitations are real: unresolved scalability, risk of interpretive projections without causal validity, and the link with safety still largely theoretical.