A Generalized Formalism of Auto-Regressive Decoding for Speech Processing
For researchers in speech processing, this work provides a systematic way to understand and evaluate auto-regressive decoding strategies, though it is primarily a theoretical contribution without empirical results.
The paper addresses the lack of a unified formalism for auto-regressive decoding strategies in speech processing. It proposes a generalized theoretical framework to categorize and compare these strategies, enabling simplified benchmark design and focused ablation studies.
In speech processing, most state-of-the-art sequence prediction models rely on auto-regressive (AR) strategies to generate output sequences based on the raw predictions of the model. Despite their crucial role in the inference process, a comprehensive overview of AR strategies as a unified field is lacking, due largely to implicit and multiple definitions of next-token decoding. This context complicates the choice, comparison, and evaluation of strategies, while creating inconsistencies in the characterization of approaches as auto-regressive or not. We begin by setting explicit inclusion criteria for the field of AR search in speech processing, and derive a generalized theoretical framework to categorize and report on search strategies for neural models. We show the capabilities of this formalism in simplifying the design of benchmarks centered around the decoding process, allowing for ablation studies that are focused on search strategies.