SDLGASJun 29

RA-QA: A Benchmarking System for Respiratory Audio Question Answering Under Real-World Heterogeneity

arXiv:2602.184529.7h-index: 15
Predicted impact top 27% in SD · last 90 daysOriginality Incremental advance
AI Analysis

It provides a standardized benchmark for evaluating multimodal AI in respiratory health, addressing the lack of realistic evaluation in this domain.

The paper introduces the RA-QA benchmark for respiratory audio question answering, comprising 9 million QA pairs from public datasets, and shows that current models fail under real-world heterogeneity.

As conversational multimodal AI tools are increasingly adopted to process patient data for health assessment, robust benchmarks are needed to measure progress and expose failure modes under realistic conditions. Despite the importance of respiratory audio for mobile health screening, respiratory audio question answering remains underexplored, with existing studies evaluated narrowly and lacking real-world heterogeneity across modalities, devices, and question types. We hence introduce the \textbf{Respiratory-Audio Question-Answering (RA-QA) benchmark}, including a standardized data generation pipeline, a comprehensive multimodal QA collection, and a unified evaluation protocol. RA-QA harmonizes public RA datasets into a collection of 9 million format-diverse QA pairs covering diagnostic and contextual attributes. We benchmark general audio-language models as well as domain-specific architectures, establishing reproducible reference points and showing how current approaches fail under heterogeneity.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes