RMS-AQA: A Two-Stage Spatial Audio Question Answering Benchmark for Real-World Domestic Environments
Peihao Chen, Qing Wang, Lichun Fan, Yufeng Hao, Zhifeng Kong, Mengyao Zhu, Hengyi Hong, Hang Chen, Hang Su, Yujie Jian, Chao-Han Huck Yang, Shichao Hu, Jun Du, Jian Luan, Ke Li
Abstract
Embodied assistants in domestic environments must infer what happened, where and when it occurred, and how to respond. To address this, we introduce RMS-AQA, a spatial audio question answering (SAQA) benchmark for real-world domestic environments. The benchmark features a two-stage question-answering (QA) format to comprehensively assess the ability of audio-language models (ALMs) to first ground audible sound events and subsequently perform complex spatio-temporal reasoning based on that grounding. To maximize acoustic realism, our dataset combines authentic real-world first-order Ambisonics (FOA) recordings with high-fidelity synthetic data generated using measured room impulse responses (RIRs). Furthermore, we provide a lightweight spatial plug-in that injects FOA-format data into frozen audio-language backbones. Experimental results reveal that the primary challenges stem from concurrent sources, far distance, and sim-to-real domain gap between RIR-synthesized and authentic recordings.
Create a lesson
Related papers
LAST: Looped Audio Spectrogram Transformer
Haider Al-Tahan, Sean O'Brien, Anastasia Razdaibiedina et al.
From Isolated Feature to Orbits: Discovering Music Concepts via Multi-SAE Alignment
Liwei Lin, Gus Xia
Beyond Decodability: Do Acoustic Factors Drive Predictions in Speech-Based Alzheimer's Assessment?
Serli Kopar, Alkis Koudounas, Roshan P. Rane et al.
Q-SPT: Learnable Query-Based Compression for Low-Frame-Rate Speech Tokenization
Jeeyoung Yun, Seohwan Yun, Sungwoong Kim
Multi-Party Backchannel Prediction: a Diagnosis, a Benchmark, and a Ceiling
Mohammed Hafsati, Ahmed Loughzali
AudioJev: Direct Audio Decisions with Order-Calibrated Probabilities
Sihan Lv, Zhen Li, Zhiqi Cao et al.