Preference-ASR: A Preference-Aware Test Set for Benchmarking ASR in the Era of Speech LLMs
For ASR researchers and practitioners, this work provides a new benchmark to evaluate user-preference alignment, addressing a gap in current evaluation that ignores stylistic preferences.
Current ASR benchmarks fail to measure whether models follow user preferences for output style (e.g., numbers, disfluencies, casing). The authors introduce Preference-ASR, a test set with preference instructions across four categories, and show that model rankings shift across preference types, revealing quality differences hidden by traditional evaluation.
Popular ASR test sets adopt inconsistent conventions for numbers, disfluencies, entities, and casing, while standard normalizers erase the format distinctions users care about. Current benchmarks therefore cannot measure whether a model follows user preferences for output style. We introduce PreferenceASR, a test set evaluating ASR systems on their ability to follow natural-language preference instructions across four categories: normalization, entities, disfluencies, and case. Built from seven open-source corpora via a two-stage LLM-assisted pipeline with human verification, it is evaluated with a preference-aware normalizer that selectively skips steps matching the active instruction. Benchmarking four models shows rankings shift across preference types, exposing quality differences traditional evaluation obscures. We publicly release the dataset.