MuSP-Bench: Advanced Multimodal Benchmarking of Music Understanding across Score and Performance
Milan Liessens Dujardin, Song-Ze Yu, Kevin Miao
Abstract
Musicians commonly communicate music through scores and performances. Scores encode musical intent, while performances realize it in sound. To investigate whether models can meaningfully engage with both modalities, we introduce MuSP-Bench, a human-authored benchmark of 490 questions targeting understanding across Musical Scores and Performances. The benchmark distinguishes itself by spanning score-based, performance-based, interpretive, and long-horizon reasoning across classical piano and orchestral works. We evaluate frontier multimodal large language models under multiple input conditions. Our results show that these models struggle substantially to understand scores, while facing even greater challenges when reasoning about performance audio. The benchmark is available at https://musp.vaclis.net/.
Create a lesson
Related papers
The Missing Temporal Link: Temporal Context Routing for Script-Driven Audio-Video Generation
Yichen Liu, Quanwei Zhang, Haozhe Wang et al.
Transfer Safety Awareness for Cross-Modal Safety Drift in Multimodal Large Language Models
Tianqi Xiao, Shiyao Cui, Minghao Zhang et al.
Can LLMs Design Video Coding Tools? A Case Study on Planar Mode
Yingwen Zhang, Meng Wang, Liqiang He et al.
Agentic Artifact Creation: Systems, Evaluation, Principles, and Opportunities
Tianfu Wang, Zhezheng Hao, Xilin Xia et al.
A Mixed-Behavior Vote Model for Multimedia Subjective Quality Votes, Means, and Variances
Jaden Pieper, Stephen D. Voran
How AI Experiences Art: Emergent Aesthetic Structure in a Self-Supervised Multimodal Embedding Space
Corey D. C. Heath