SDAIJun 19

AOR-Bench: Do Large Audio Language Models Over-Refuse Pseudo-Harmful Queries?

arXiv:2606.2114717.6
Predicted impact top 10% in SD · last 90 daysOriginality Incremental advance
AI Analysis

For researchers and developers of LALMs, this work highlights a critical safety alignment issue (over-refusal) in the audio domain and provides a benchmark to evaluate and mitigate it.

The paper introduces AOR-Bench, the first benchmark for over-refusal in Large Audio Language Models, containing 3,000 pseudo-harmful audio samples. Evaluating 12 LALMs reveals widespread over-refusal, and two lightweight mitigation strategies are explored.

Large Audio Language Models (LALMs) have demonstrated strong performance across a wide range of audio tasks. As they are increasingly deployed in real-world applications, ensuring their safety alignment has become more important. Although refusal mechanisms serve as a key safeguard by preventing LALMs from responding to harmful requests, they can also lead to {\em over-refusal}, where models incorrectly reject benign queries. This issue is especially challenging in the audio domain because speech that appears harmful in isolation may become benign when interpreted together with the surrounding acoustic context, such as background sounds. To study this problem, we introduce \textbf{AOR-Bench} (\textbf{A}udio \textbf{O}ver-\textbf{R}efusal \textbf{Bench}mark), the first benchmark for over-refusal specifically designed for LALMs. AOR-Bench contains 3,000 pseudo-harmful audio samples across six scenario categories. Evaluating 12 representative LALMs from six major model families, we find that over-refusal is widespread (Figure~\ref{fig:overall_performance}) and uncover several important patterns in their safety judgments. As a preliminary effort to mitigate this issue, we further explore two lightweight strategies (e.g., Chain-of-Thought and activation steering) to reduce over-refusal.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes