Dissonance Spectrum explicitly models perceptual frequency interactions for better music understanding
Tianle Wang, Xinyi Tong, Liangke Zhao, Jishang Chen, Sirui Zhang, Haoxin Zhang, Xin Jin, Duo Xu, Xiaobing Li, Song-Chun Zhu
Abstract
Conventional music representations describe acoustic energy over time and frequency but do not explicitly expose relations among simultaneous frequency components. We introduce the Dissonance Spectrum (DS), a nonnegative time--frequency representation that applies a tolerance-based rational pitch-relation kernel with logarithmic harmonic distance to a constant-Q spectrum and attributes aggregate pairwise interactions back to individual frequency bins. Controlled music-theory tests show strong ordinal agreement for intervals, harmonic-function connections, and church modes, and weaker but significant agreement across diverse chord voicings. DS is then encoded by a lightweight parallel branch whose zero-initialized residual projection preserves the baseline function at initialization. Across six paired training seeds in open-ended music question answering and categorical and dimensional music emotion recognition, DS obtains the highest mean on every reported endpoint relative to the unchanged baseline, a parameter-matched Gaussian-input branch, and an architecture-matched magnitude-CQT branch. These results support DS as an interpretable, complementary representation, while listener-specific perception and broader task coverage remain open problems.
Create a lesson
Related papers
Direct or Mediated? Task-Dependent Audio Information Routing in Large Audio Language Models
Yizhou Zhang, Wangjin Zhou, Xin Gu et al.
SpeechGym: An Audio-Native Gym for Training Voice Agents via Reinforcement Learning
Jiajun Fan, Jingyuan Li, Prashanth Gurunath Shivakumar et al.
AudioSpan: Spanning the Duration and Depth of Audio Comprehension
Wen Huang, Yunfei Chu, Meng Gao et al.
Decay-Region Group Delay as a Forensic Cue for AI-Generated Impulsive Sounds
JaeHyeong Chang, Chengzhe Sun, Siwei Lyu
StreamAV-Bench: A Comprehensive Benchmark for Streaming Audio-Video Generation
Kaiqi Liu, Haoxuan Zeng, Jingqi Liu et al.
Attention-Guided Reliability Scaling for Contrastive Decoding in Robust Audio-Visual Speech Recognition
YoungChae Kim, Da-Hee Yang, Joon-Hyuk Chang