Phone Segmentation and Recognition through Phonological Activation Mapping

CMU
arXiv:2607.0902022.3h-index: 30
Predicted impact top 4% in AS · last 90 daysOriginality Incremental advance
AI Analysis

For speech processing researchers, this work shows that phonetic structure is latent in self-supervised models and can be efficiently extracted for both segmentation and recognition tasks with minimal supervision.

The paper introduces SPAM, a method that uses phonological feature activations from self-supervised speech models to perform phone segmentation and recognition with less than a minute of transcribed data, achieving strong performance across diverse datasets.

Phone segmentation and recognition are inherently related tasks, yet modern approaches typically model them separately. We argue that phonetic structure is already latent in the representations of self-supervised speech models (S3Ms), and one only needs to steer them to solve both tasks. We leverage S3M-based Phonological Activation Mapping (SPAM), which maps each S3M representation frame to a vector of phonological feature activations, such as voicing and nasality. On top of SPAM, we introduce two simple but effective lightweight, gradient-descent-free prediction heads: a recognition head and a segmentation head. Our method requires less than a minute of phonetic transcriptions, and generalizes to unseen phones during training. Across a diverse range of datasets, our approach attains strong segmentation and recognition performance.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes