SDAIJun 12

Listening with Attention: Entropy-Guided Explainability for Transformer-Based Audio Models

arXiv:2606.14647v18.9
Predicted impact top 47% in SD · last 90 daysOriginality Incremental advance
AI Analysis

For researchers and practitioners using transformer-based ASR models, LEAF-X provides a more faithful and interpretable explanation method, addressing the need for transparency in high-stakes applications.

The paper introduces LEAF-X, an explainability framework for transformer-based ASR models that combines entropy-guided attention weighting and multi-layer attention rollout to produce faithful, sparse token-to-frame attributions, achieving 32% improved faithfulness and 35-39% stronger locality/sparsity over existing methods.

Transformer-based automatic speech recognition (ASR) models such as Whisper are highly accurate, but their predictions remain difficult to interpret. Existing explainable AI (XAI) methods often lack faithfulness and precise temporal grounding. We propose Listening with Entropy-guided Attention for Faithful explainability (LEAF-X), a model-intrinsic XAI framework for transformer-based ASR. LEAF-X combines entropy-guided attention weighting, multi-layer attention rollout, and optional causal ablations to identify low-entropy, high-impact heads and layers, producing sparse token-to-frame attributions. Unlike perturbation-based explainers or raw attention maps, LEAF-X exploits the internal structure of encoder-decoder and speech-augmented decoder-only models to generate explanations that better reflect model computation. Results show 32% improved faithfulness, 35-39% stronger locality/sparsity, and the most stable attributions, supporting more transparent and auditable ASR.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes