Overview and Meta-Analysis of DCASE 2026 Challenge Task 6: Audio Moment Retrieval from Long Audio
Hokuto Munakata, Tatsuya Komatsu, Keisuke Imoto, Taichi Nishimura, Huang Xie, Tuomas Virtanen
Abstract
This paper presents an overview of the Detection and Classification of Acoustic Scenes and Events (DCASE) 2026 Challenge Task 6, Audio Moment Retrieval (AMR) from Long Audio. Given a several-minute-long audio recording and a free-form text query, AMR aims to retrieve temporal moments in the recording that match the query, where each moment is represented by a pair of start and end timestamps. This task requires effective cross-modal alignment and long-range temporal modeling. We describe the task definition, the evaluation metrics, the development and evaluation datasets, and a baseline system that combines a pre-trained MS-CLAP feature extractor with a Detection Transformer (DETR)-based moment-detection network. On the development data, the baseline trained on a manually annotated dataset and a synthetic dataset achieved Recall1@0.7 of 13.56%, indicating that AMR in long audio remains a challenging problem. The challenge attracted 21 teams, which submitted 59 systems in total. The three best systems achieved Recall1@0.7 of 48.59%, roughly 3.5 times the baseline score. The results show that strengthening the audio-text feature extractor and the moment-detection network led to substantial performance improvements. Furthermore, the top three teams boosted performance by applying confidence score calibration or ensembling across different temporal resolutions of features.
Create a lesson
Related papers
MP-Bench: Evaluating Voice Agents as a Multiparty Conversation Participant
Yi-Jen Shih, Shih-Yun Shan Kuan, Guan-Ting Lin et al.
Objective Intelligibility Prediction Using Distance Metrics on Speech Foundation Model Representations
Lyonel Behringer, Andreas Brendel
AlignDPO: Preference-Gated Alignment for Reducing Hallucination in Decoder-Only TTS
Xiao Zhou, Oisín Turbitt, Kit Bower-Morris et al.
A Device to Control and Manipulate Occlusion Effects for Own Voice Perception Studies
Rouben Rehman, Simon Kersten, Aron Schliep et al.
X-Pred MeanFlow for Streaming Token-to-Mel Speech Decoding
Hanke Xie, Xiaming Ren, Qirui Zhan et al.
Location-based Training with Complementary Folded Linear Orderings for Multichannel Speech Separation
Kaixuan Yang, Stijn Kindt, Nilesh Madhu