Auditory Illusion Benchmark for Large Audio Language Models
Hayoon Kim, Eunice Hong, Kyogu Lee
Abstract
Perceptual illusions have long served as crucial probes into human cognition, revealing biases and limitations of perception. In the auditory domain, such illusions provide a unique lens for testing whether Large Audio Language Models (LALMs) replicate human perceptual tendencies. Despite their importance, most benchmarks focus on visual illusions or general audio tasks, leaving auditory illusions underexplored. To this end, we present AIB, the first auditory illusion benchmark for LALMs, covering ten representative illusions across music, sound, and speech, each annotated for the presence of knowledge-based priors. Our methodology pairs model evaluation with controlled human listening studies, enabling direct comparison of responses. Results show systematic differences: while most LALMs remain signal-faithful on low-level acoustic illusions, several exhibit more human-like responses when linguistic or musical priors are involved, although no model matches the human perceptual profile. These findings highlight the current limitations of LALMs as cognitive models. By establishing auditory illusions as a rigorous testbed, our work offers a new perspective for probing neural black-box models and advancing understanding of auditory cognition. AIB is publicly available at https://github.com/gillosae/aib.
Create a lesson
Related papers
Understanding Automatic Mixing: A Subtask-Oriented Analysis of Two-Stage Mixing System
Jinjie Shi, Wei Hua, Kunzhu Xie et al.
Scalable Direction-Following TTS via Voice Impression-Guided Pseudo Triplet Construction
Kenichi Fujita, Yusuke Ijima
Removing Speech, Keeping Activities: A Privacy Firewall for Acoustic Sensing in Assisted Living
Pavlos Nicolaou, Christos Efstratiou
SonicCaps: Large-Scale Diverse and Fine-Grained Captioning for Improved Audio-Retrieval
Zineb Lahrichi, Marc Ferras, Gaël Richard et al.
Efficient Passive Acoustic Monitoring of Killer Whales Using a Two-Stage Detection and Ecotype Classification Cascade
Daniela Ruiz, Manuel Castellote, Zhongqi Miao et al.
Soft Posterior Speaker Injection for Multi-Talker Speech Recognition
Jian Zhu, Cheng Luo