SDASJun 25

wav2tok 2.0: Scalable Audio Tokenization Maintaining Explicit Pairwise Token Alignment for Efficient Audio Retrieval

arXiv:2606.268245.1
Predicted impact top 76% in SD · last 90 daysOriginality Incremental advance
AI Analysis

For researchers in spoken term detection, this work provides a more scalable and effective tokenizer for variable-length audio retrieval.

wav2tok 2.0 introduces a scalable audio tokenizer for query-by-example spoken term detection that enforces pairwise token alignment via staged training, outperforming BEST-STD and general-purpose tokenizers while maintaining efficiency.

Learning discrete speech representations that preserve similarity across variable-length utterances is central to query-by-example spoken term detection (QbE-STD). While wav2tok introduced CTC-based sequence alignment to enforce token consistency, its tightly coupled clustering and alignment training recipe limits scalability. We propose wav2tok 2.0, a scalable alignment-aware speech tokenizer built on the BEST-STD backbone. wav2tok 2.0 employs staged training, first learning discriminative, speaker-invariant representations via contrastive learning and vector quantization, and then enforcing pairwise token consistency using a CTC alignment loss and a novel DTW-aligned framewise prediction objective with adaptive weighting. Experiments show that wav2tok 2.0 consistently outperforms BEST-STD and general-purpose tokenizers on QbE-STD while remaining efficient and scalable.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes