arXiv cs.LG
7/16/2026

Targeted Recovery of Weight-Space Mechanisms From Neural Networks
Short summary
This paper proposes targeted parameter decomposition (tPD), which identifies only the computational components in a neural network that process specific inputs of interest, using a high-rank catch-all component for non-target data. This reduces the compute cost of mechanistic interpretability compared to full parameter decomposition. On transformer language models trained on The Pile, the authors extract a CSS-only submodel using 7% of the FLOPs of a full decomposition and demonstrate surgical ablation and rewiring of memorized sequences with negligible side effects.
- •Targeted parameter decomposition isolates circuits for specific inputs at 7% of full decomposition FLOPs
- •Validated on transformer LMs trained on The Pile, recovering mechanistically faithful circuits
- •Demonstrates surgical ablation and rewiring of memorized sequences in a 12-block transformer
Generated with AI, which can make mistakes.
Is this a good recommendation for you?