Circuit Condensation: Post-Training that Concentrates a Behavior's Causal Circuit
Sai Adith Senthil Kumar
Abstract
One approach to mechanistic interpretability explains behavior through circuits: the components and connections that carry it. Frozen discovery often returns hundreds of edges, making them hard to inspect, compare, or verify exhaustively. We introduce Circuit Condensation, which post-trains models to concentrate behaviors into smaller causal graphs. Each round prunes low-attribution edges and trains a low-rank adapter to match the original through what remains, retaining the cut only if task performance and general capability survive. Across four behaviors and eight models, condensed circuits are smaller than the strongest frozen baseline in 30 of 32 settings, by 8.1× on average and up to 316×. Repeating the search without weight updates produces larger circuits in 29 of 32 settings, showing that weight updates, rather than search alone, drive the reduction. Testing every subset of 19 circuits finds 11 that cannot be reduced and reveals removable edges in the rest. Pair ablations expose dependencies between edges, showing that their effects cannot be understood independently. On indirect object identification, condensation isolates 24 heads, 17 of them with documented roles, against 61 heads and 36 undocumented ones for the matched frozen circuit: a sufficient sub-circuit of the published mechanism rather than a reconstruction of it. The resulting circuit tracks the original model's next-token distribution and predicts its errors.
Create a lesson
Related papers
Understanding Evolution Strategies for LLM Reasoning: Broader Reasoning Coverage than GRPO
Yunpeng Ba, Zhi Zheng, Yue Xie et al.
Beyond Parallel Blindness: Information Floors and Model Gaps in Block Drafting
Xinwei Qiang, Xiang Fang, Chang Chen et al.
QuantumBoostNet: A Hybrid Classical-Quantum Architecture for Enhanced Accuracy in Cardiac Ultrasound View Identification
Mihai Udrescu-Milosav, Stefan-Alexandru Jura, Mihai Udrescu et al.
MM-Spectrum: Multimodal Multi-spectral Molecular Structural Elucidation with a Stable MoE Framework
Hai-tao Yu, Nan Min, Zheng Fang et al.
Making Latent Evolution Explicit: Operator-Structured Transitions for World Action Models
Xiaoxiao Lu, Yunlong Dong, Jiahao Shi et al.
Importance Scoring of Transformer Attention Heads in Learning Tabular Data
Ahmad Jad Allah, Kazi F. Akhter, Md. Kamrozzaman Bhuiyan et al.