ASCLJun 17

IndicContextEval: A Benchmark for Evaluating Context Utilisation in Audio Large Language Models Across 8 Indic Languages

arXiv:2606.1915710.7
Predicted impact top 29% in AS · last 90 daysOriginality Incremental advance
AI Analysis

For researchers and developers of AudioLLMs, this benchmark exposes the gap between claimed context usage and actual model behavior, highlighting the need for explicit evaluation of contextual grounding.

IndicContextEval is a 56-hour multilingual benchmark across 8 Indian languages and 23 domains, designed to test whether AudioLLMs genuinely use contextual prompts (e.g., entity lists, domain descriptions) for transcription. Evaluation of five models shows substantial differences in context utilisation, indicating that current models often rely on parametric knowledge rather than provided context.

AudioLLMs enable speech recognition conditioned on textual prompts such as domain descriptions or entity lists. However, it remains unclear whether these models genuinely utilise such context or rely on parametric knowledge learned during pretraining. Existing benchmarks cannot answer this question because they evaluate transcription under fixed prompting conditions and rarely include explicit contextual inputs. We introduce IndicContextEval, a 56-hour multilingual benchmark of natural speech from 555 speakers across 8 Indian languages and 23 professional domains. We design a 7-level prompting framework that progressively introduces contextual signals, including metadata, natural-language descriptions, entity lists in English and native script, and adversarial prompts with incorrect entities. Evaluating five models reveals substantial differences in context utilisation behaviour, highlighting the need for explicit evaluation of contextual grounding in AudioLLMs.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes