Faithfulness to Refusal: A Causal Audit of Neuron Selectors
For practitioners of model editing and interpretability, this work reveals that rank-stability is an unreliable proxy for causal validity and that refusal mechanisms are redundant, challenging uniqueness claims.
Attribution scores for neuron rows in LLMs are rarely tested for causal importance. This paper introduces two causal audits showing that attribution methods outperform baselines at identifying dispensable rows, and that attributed rows can install refusal behavior while preserving fluency, but different methods find disjoint sufficient sets, indicating redundancy.
Attribution scores increasingly identify which neuron rows of a language model matter for applications such as pruning, interpretability, and editing for safety, yet whether they identify causally important rows is rarely tested directly. We address this with two paired audits built on one-shot neuron-row zeroing. We first audit selectors at the language-modeling level: attribution methods substantially outperform activation and magnitude-based baselines at identifying dispensable rows across five LLMs. We then adapt the same intervention into a behavior test by driving it with a contrastive harmful-versus-benign signal; the attributed rows are sufficient to install refusal on hate and crime while keeping benign over-refusal low and preserving language model fluency, and specific in that layer-matched random controls at the same depths fail. Highly rank-stable selectors can be among the least causally valid. Refusal moreover lives in a redundant subspace, where different attribution methods install it through largely disjoint row sets, so the recovered edit is one realization of a sufficient set rather than a unique mechanism. Together, these findings show that rank-stability proxies miss the kinds of selector failures a direct causal audit can surface.%