CLAIJun 25

SamaVaani: Auditing and Debiasing Multilingual Clinical ASR for Indian Languages

arXiv:2606.2690115.2Has Code
Predicted impact top 66% in CL · last 90 daysOriginality Incremental advance
AI Analysis

For healthcare AI practitioners deploying ASR in multilingual Indian clinical settings, this work identifies and mitigates systematic biases that could lead to inequitable care.

The paper audits eight ASR models on multilingual psychiatric interview data from India, revealing performance disparities across languages, speaker roles, and genders. It proposes SamaVaani, a debiasing technique that improves both ASR accuracy and fairness, achieving equitable performance across demographic groups.

Automatic Speech Recognition (ASR) is increasingly used to document clinical encounters, yet its reliability in multilingual and demographically diverse Indian healthcare context remains largely unknown. In this study, we first conduct the systematic audit of ASR performance on real-world psychiatric interview data spanning Kannada, Hindi and Indian English, comparing eight state-of-the-art models including IndicWhisper, WhisperLargeV3, Sarvam, GoogleS2T, Gemma3n, OmniLingual, Vaani, and Gemini. Our results reveal substantial variability across models and languages, with some systems performing competitively in Indian English but failing in regional speech. We further fine-tune two of the best performing opensource models, i.e., Gemma3n and OmniLingual, using various methods. With this, we uncover systematic performance gaps tied to speaker role and gender, raising concerns about equitable deployment in clinical settings, which are further mitigated by fairness-aware fine-tuning. To this end, we propose SamaVaani, a unified debiasing technique that simultaneously improves ASR performance and improves fairness across demographic groups.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes