SonicCaps: Large-Scale Diverse and Fine-Grained Captioning for Improved Audio-Retrieval
Zineb Lahrichi, Marc Ferras, Gaël Richard, Geoffroy Peeters
Abstract
Recent advances in audio-language modeling have been driven by large-scale audio captioning datasets. However, existing datasets remain limited by low semantic diversity, generic descriptions lacking acoustic details, and one-to-one audio-caption mappings that poorly reflect the inherent ambiguity of auditory perception. We introduce SonicCaps, a large-scale audio captioning dataset comprising ~15M captions paired with ~700k audio clips, generated using a multi-modal large language model (Qwen3-Omni) conditioned on both audio and text. To explicitly promote diversity, we generate around 24 captions per audio via structured prompt engineering and few- shot generation, spanning main descriptions, rephrased variants (verbosity, style) and semantic tags. Human evaluation shows that SonicCaps is rated significantly higher than existing captioning datasets, with fine-grained analyses indicating that our captions are perceived as more descriptive and precise, which strongly correlates with quality judgments. Finally, training CLAP models on SonicCaps with a multi-caption sampling strategy consistently improves audio retrieval and zero-shot classification, with stronger generalization across public and commercial benchmarks. We release both SonicCaps and two specialized CLAP models on hugging face: https://huggingface.co/datasets/Zineb/SonicCaps.
Create a lesson
Related papers
Understanding Automatic Mixing: A Subtask-Oriented Analysis of Two-Stage Mixing System
Jinjie Shi, Wei Hua, Kunzhu Xie et al.
Scalable Direction-Following TTS via Voice Impression-Guided Pseudo Triplet Construction
Kenichi Fujita, Yusuke Ijima
Removing Speech, Keeping Activities: A Privacy Firewall for Acoustic Sensing in Assisted Living
Pavlos Nicolaou, Christos Efstratiou
Auditory Illusion Benchmark for Large Audio Language Models
Hayoon Kim, Eunice Hong, Kyogu Lee
Efficient Passive Acoustic Monitoring of Killer Whales Using a Two-Stage Detection and Ecotype Classification Cascade
Daniela Ruiz, Manuel Castellote, Zhongqi Miao et al.
Soft Posterior Speaker Injection for Multi-Talker Speech Recognition
Jian Zhu, Cheng Luo