TAG-Bench: Benchmarking Temporal Audio Grounding in Large Audio Language Models
Yuhang Dai, Xin Shu, Zengxi Li, Lei Xie, Xiangang Li, Jianwei Yu
Abstract
Large audio language models (LALMs) can describe what is heard, but their ability to localize when queried content occurs remains less systematically evaluated. We present TAG-Bench, a benchmark for temporal audio grounding in which a model returns every time interval that matches a natural-language query. TAG-Bench contains 1,750 human-verified query-recording pairs covering 149.5 hours, with eight source-dependent subsets spanning query categories and audio durations from 7 s to 20 min; 22.1% of the queries have multiple ground-truth intervals. Across 21 evaluated systems, the best-performing model achieves 31.2 mIoU and is the only system above 20 mIoU on the two long subsets, yet even this top performer reaches only 21.5% recall at IoU >= 0.7. Moreover, 9 of 21 systems fall below 5 mIoU, and every model under-reports the number of occurrences on one-to-many queries, with none exceeding 13.2% count accuracy. Because responses are free-form, we report parsing-failure rate and MAE coverage: parsing failures remain in mIoU, Recall, gIoU, and count metrics as empty predictions but do not enter MAE. The results separate precise localization, occurrence enumeration, and output-format reliability within a benchmark whose cross-subset comparisons are descriptive rather than controlled estimates of query abstraction or duration. We will release the TAG-Bench data and evaluation code to support future research.
Create a lesson
Related papers
VibeVoice-ASR-Streaming Technical Report
Yujie Tu, Zhiliang Peng, Jianwei Yu et al.
VAANI Noise Event Dataset: A curated spontaneous speech dataset annotated with timestamps for noise events
Pavan Kumar J, Agneedh Basu, Pranav Bhat et al.
Sensing Bone-Conducted Speech with Earbuds
Christoph Weyer, Peter Jax
Ontology-based Target Sound Extraction
Carlos Hernandez-Olivan, Marc Delcroix, Tsubasa Ochiai et al.
U-PAST: A Phase-Aware Audio Spectrogram Transformer-U-Net for Single-Channel Speech Enhancement
Cao Duong Ly, Jörn Anemüller
Likelihood-Constrained Acoustic Reranking for Training-Free Hallucination Mitigation in LLM-Based ASR
Jiasheng Kuang, Linru Zheng, Hongjin Song et al.