SDJun 17

Constraining to Generalize: Subspace Tuning for Few-shot Generalization of Audio-Language Models

arXiv:2606.185607.1
Predicted impact top 64% in SD · last 90 daysOriginality Incremental advance
AI Analysis

For practitioners adapting audio-language models to new tasks with limited data, SubT addresses the critical trade-off between seen and unseen class performance.

Few-shot adaptation of audio-language models suffers from a base-to-new trade-off due to zero-shot drift in text embeddings. Subspace Tuning (SubT) uses geometry-constrained adaptation to improve unseen-class generalization, achieving strong results across 11 audio benchmarks without text-encoder backpropagation.

Few-shot adaptation of pretrained Audio--Language Models (ALMs) often improves seen-class performance at the cost of unseen-class generalization, leading to the base-to-new trade-off. We attribute this failure to zero-shot drift in the text embedding space: few-shot tuning can distort inter-class structure and move adapted embeddings far from their pretrained anchors. We therefore propose Subspace Tuning (SubT), a geometry-constrained adaptation framework with two complementary controls on drift. Structured Subspace Parameterization limits structural deformation, and Residual Anchoring stabilizes adaptation around the zero-shot prior. At inference time, Subspace-aware Gating further suppresses negative transfer for weakly aligned unseen classes. Across 11 audio benchmarks, SubT delivers strong few-shot generalization while remaining efficient, operating directly on precomputed text embeddings without text-encoder backpropagation.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes