LGAIMay 30

Query Lens: Interpreting Sparse Key-Value Features with Indirect Effects

arXiv:2606.0761711.9h-index: 11
Predicted impact top 29% in LG · last 90 daysOriginality Incremental advance
AI Analysis

Provides a more comprehensive and faithful method for interpreting sparse features in neural networks, addressing a key bottleneck in mechanistic interpretability.

Query Lens extends Logit Lens to interpret sparse features in autoencoders by jointly considering encoder-side key features and decoder-side value features, and accounting for indirect effects. It yields coherent token signatures for previously uninterpretable features.

While sparse autoencoders provide features more interpretable than individual neurons, reliably characterizing them remains challenging. We propose Query Lens, which extends Logit Lens to enable more comprehensive and faithful interpretations of sparse features. By jointly considering encoder-side key features and decoder-side value features, we identify both the inputs that activate a feature and the outputs it promotes. We also account for indirect, module-mediated effects that arise when the feature is processed by downstream modules, going beyond the direct effect captured by Logit Lens. In experiments, we find that Query Lens yields coherent token signatures for features that remain uninterpretable under Logit Lens. Finally, we propose the Subspace Channel Hypothesis, suggesting that downstream modules read features through layer-specific subspaces.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes