LGJul 8

Distributed Sparse Interventions in Language Models

arXiv:2607.0712819.3h-index: 11
Predicted impact top 4% in LG · last 90 daysOriginality Incremental advance
AI Analysis

This work provides a more precise and interpretable method for controlling and understanding task-specific computations in large language models, addressing a key bottleneck in model interpretability.

The authors introduce Distributed Sparse Interventions (DSI), a method for steering language models by intervening on as few as 0.01% of neurons, which activates task behavior more effectively than prior global direction approaches by capturing neuron-specific nonlinear effects.

Language models perform a wide range of tasks at varying levels of abstraction with the capacity to flexibly infer tasks from context, execute multiple tasks simultaneously, and select among competing tasks. To study the role of model components in task behaviour, their causal influence can be investigated through interventions. Prior work on model steering has largely focused on interventions along global directions in activation space, modeling task representations as approximately linear and additive. By studying interventions at the neuron level, we find substantial, neuron-specific nonlinear effects on model outputs that are not captured by current steering approaches. We introduce Distributed Sparse Interventions (DSI), an intervention approach that considers nonlinearities and interactions between neurons across layers to identify sparse sets of neurons that elicit task-relevant computations. Across a range of tasks, we demonstrate that DSI can activate task behaviour in instruction-tuned language models by localising and intervening on as few as 0.01% of neurons, highlighting the effectiveness of sparse, distributed interventions in the neuron basis. Additionally, adopting a set-based perspective enables computations over the identified neuron sets, offering insights into the roles of individual neurons by analysing their effects across tasks. Through sparse interventions, DSI enables fine-grained control over model behaviour, localisation of task-relevant neuron sets, and furthers our understanding of task composition.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes