SDJul 7

InsideSSL: Understanding Self-Supervised Speech Representations using a Model-Centric Perspective

arXiv:2607.0639210.4
Predicted impact top 26% in SD · last 90 daysOriginality Incremental advance
AI Analysis

For researchers studying self-supervised speech representations, this work provides a systematic understanding of how different SSL objectives shape internal representations, though it is an incremental analysis framework rather than a new method.

The paper proposes InsideSSL, a model-centric framework to analyze internal layer-wise dynamics of SSL speech models (Wav2Vec2, HuBERT, WavLM). It reveals distinct regimes of compression and manifold unfolding across layers, and introduces a Generative Compatibility Matrix to evaluate functional transferability, showing stable phonetic cores and deep-layer semantic pruning.

Self-supervised learning (SSL) models, such as Wav2Vec2, HuBERT, and WavLM, have become foundational across a wide range of speech and audio tasks. Despite their success, understanding their internal layer-wise dynamics remains an ongoing challenge. To address this, we propose a two-part model-centric framework called InsideSSL. First, we establish a task-agnostic analysis from three intrinsic per-layer perspectives: compression (entropy), geometry (curvature), and robustness to perturbations. We show that varying training objectives induce distinct regimes of acoustic compression and manifold unfolding. Second, we introduce the cross-layer Generative Compatibility Matrix (GCM) to evaluate functional transferability, exposing stable phonetic cores, identity volatility, and deep-layer semantic pruning. In addition to these evaluations, linear probing connects the model-centric perspective to downstream tasks, demonstrating how layer topology dictates phoneme, pitch, and speaker encoding.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes