Are We Recovering Mechanisms? Objective-Level Recovery Gaps in Mechanistic Interpretability
Chuqin Geng, Li Zhang, Haolin Ye, Mark Zhang, Luke Zhang, Xujie Si
Abstract
Mechanistic interpretability aims to recover the internal computations responsible for model behavior. Progress in automated circuit discovery is often framed as a search problem: better attribution or optimization should identify better mechanisms. This assumes that the evaluation objective can recognize a better circuit once it is found. We show that intervention-defined faithfulness can instead prefer an equally sized circuit that reproduces the model's behavior less well, creating an objective-level recovery gap. Across four human-reference tasks and InterpBench, we compare validation faithfulness with behavior on held-out prompts under fixed ordinary resampling. The behavioral criterion is agreement with the intact model, including its mistakes, except on Greater-Than, where we use semantic accuracy. Controlled reference edits reveal misranking without any discovery algorithm, and outputs of EAP, EAP-IG, ACDC, and Edge-SP exhibit the same failure. Under resampling, KL misranks 9.4%-41.2% of candidate pairs across these methods on the human-reference tasks. We investigate context distortion as an explanation: replacing excluded signals changes the inputs on which retained components operate. Restoring selected signals from the recipient's intact-model execution repairs 96 of 100 persistent KL misrankings from the discovery pool on both validation and held-out prompts. The circuits and their original behavioral scores remain unchanged. These findings show why better discovery alone is insufficient when its objective rewards the wrong candidate.
Create a lesson
Related papers
TACO: Ternary Absolute-max Column-wise One-sparse Optimizer for LLM Fine-Tuning
Jichao Jiang, Cristian McGee, El Houcine Bergou et al.
FERPO: Forward Entropy-Regularized Policy Optimization
Sebastian Sanokowski, Alireza Sarmadi, Majid Khadiv
Cost-augmented Schrödinger bridges on graphs are exactly solvable: a Feynman-Kac tilt replaces learned control
Akshay Balsubramani
The Missing Primitive: Diagnosing and Repairing Mathematical Reasoning in Large Language Models
Shuo Xing, Zilin Dai, Chengyuan Qian et al.
Trust the Direction, Search the Step: Zero-and-First-Order Methods for LLM Fine-Tuning
Cristian McGee, El Houcine Bergou, Aritra Dutta
Generative modeling of intrinsically disordered protein regions by reinforcing sparse autoencoder features
Jason X. Liu, Sebastian Ibarraran, Frank Hu et al.