TUTTI: Toward generalizable audio-to-score transcription via fully synthesized data
Jianhuai Hu, Yashan Wang, Shangda Wu, Zhancheng Guo, Shijie Liang, Wuna Meng, Chuanqi Yang, Xiaobing Li, Feng Yu, Maosong Sun
Abstract
Generalizable Audio-to-Score (A2S) transcription is fundamentally constrained by the severe scarcity of high-quality, real-world paired data. Relying solely on existing human-annotated datasets often restricts the generalization of A2S models, limiting their efficacy primarily to single-instrumentation domains. To break this dependency on scarce real-world data, we introduce TUTTI (Transformer for Unified audio-To-score Transcription trained on Synthetic multi-Instrumentation Data), a pre-training paradigm driven by a purely synthetic, large-scale dataset. Rather than using human-composed scores, we leverage a symbolic music generation model to generate a massive, highly scalable multi-instrumentation corpus and create audio-score pairs with expressive acoustic characteristics. Capitalizing on the generated data, we employ a standard Transformer encoder-decoder architecture. We empirically demonstrate that pre-training a unified attention-based model on generated, multi-instrumentation data yields a consistently stronger foundational representation than single-instrumentation training. When fine-tuned with downstream real-world datasets, TUTTI outperforms previous approaches, establishing new overall state-of-the-art results across various A2S baselines. Notably, TUTTI shows remarkable cross-instrument transferability, effectively adapting to unseen instruments with highly competitive performance. The source code and the TuttiCorpus dataset will be made publicly available at https://github.com/a-musiclover/TUTTI.
Create a lesson
Related papers
Understanding Automatic Mixing: A Subtask-Oriented Analysis of Two-Stage Mixing System
Jinjie Shi, Wei Hua, Kunzhu Xie et al.
Scalable Direction-Following TTS via Voice Impression-Guided Pseudo Triplet Construction
Kenichi Fujita, Yusuke Ijima
Removing Speech, Keeping Activities: A Privacy Firewall for Acoustic Sensing in Assisted Living
Pavlos Nicolaou, Christos Efstratiou
SonicCaps: Large-Scale Diverse and Fine-Grained Captioning for Improved Audio-Retrieval
Zineb Lahrichi, Marc Ferras, Gaël Richard et al.
Auditory Illusion Benchmark for Large Audio Language Models
Hayoon Kim, Eunice Hong, Kyogu Lee
Efficient Passive Acoustic Monitoring of Killer Whales Using a Two-Stage Detection and Ecotype Classification Cascade
Daniela Ruiz, Manuel Castellote, Zhongqi Miao et al.