Direct or Mediated? Task-Dependent Audio Information Routing in Large Audio Language Models
Yizhou Zhang, Wangjin Zhou, Xin Gu, Yichi Wang, Wei Tan, Yi Zhao, Zhi Gong, Keisuke Imoto, Tatsuya Kawahara
Abstract
Large Audio Language Models (LALMs) have demonstrated strong performance across a wide range of audio understanding tasks. However, they are typically evaluated on single, coherent audio segments, leaving their behavior under less familiar input configurations underexplored. We study this issue through a controlled setting in which two audio segments are concatenated into a single input. Across multiple LALMs, we observe a striking task-dependent robustness gap: automatic speech recognition (ASR) remains comparatively stable, whereas audio question answering (AQA) degrades substantially. To investigate the mechanisms underlying this disparity, we analyze how audio information is routed through LALM decoders using layer-wise attention knockout. The results reveal distinct task-dependent pathways. ASR relies primarily on direct retrieval from audio tokens by answer tokens, whereas AQA depends more strongly on a mediated route in which audio information is first integrated into prompt tokens and subsequently accessed during generation. We further probe prompt-token representations under audio concatenation and find that task-relevant audio attributes remain readily decodable, particularly in middle and later decoder layers, even when AQA performance deteriorates sharply. This dissociation indicates that the failure cannot be explained by complete loss of audio information from the decoder states and is instead consistent with a downstream bottleneck in retrieving or utilizing prompt-mediated information during answer generation. Together, our findings reveal task-dependent audio information routing in LALMs and highlight information utilization as a potential limitation on their generalization.
Create a lesson
Related papers
SpeechGym: An Audio-Native Gym for Training Voice Agents via Reinforcement Learning
Jiajun Fan, Jingyuan Li, Prashanth Gurunath Shivakumar et al.
AudioSpan: Spanning the Duration and Depth of Audio Comprehension
Wen Huang, Yunfei Chu, Meng Gao et al.
Decay-Region Group Delay as a Forensic Cue for AI-Generated Impulsive Sounds
JaeHyeong Chang, Chengzhe Sun, Siwei Lyu
StreamAV-Bench: A Comprehensive Benchmark for Streaming Audio-Video Generation
Kaiqi Liu, Haoxuan Zeng, Jingqi Liu et al.
Attention-Guided Reliability Scaling for Contrastive Decoding in Robust Audio-Visual Speech Recognition
YoungChae Kim, Da-Hee Yang, Joon-Hyuk Chang
Dissonance Spectrum explicitly models perceptual frequency interactions for better music understanding
Tianle Wang, Xinyi Tong, Liangke Zhao et al.