AVSD-Scenes: A Dataset for Audio-Visual Description of Urban Scenes
Dhanunjaya Varma Devalraju, Arshdeep Singh, Mark D. Plumbley
Abstract
Natural language descriptions can provide rich semantic representations of audio-visual urban scenes, yet datasets that jointly describe both auditory and visual information remain limited. In this paper, we introduce AVSD-Scenes, a paired audio-visual scene description dataset for urban environments. The dataset contains 12,291 audio-visual scene descriptions generated from the TAU Urban Audio-Visual Scenes dataset. To construct the dataset, we first generate audio- and visual-based descriptions using Qwen2-Audio-7B and Qwen2.5-VL-7B, respectively. These modality-specific descriptions are then combined using large language models, namely Qwen3-14B, Mistral-Small-3.2-24B-Instruct-2506, and Gemma-3-27B-it, to produce multimodal descriptions that capture complementary information from both modalities. We benchmark AVSD-Scenes using semantic alignment, cross-modal retrieval, scene classification, LLM-as-a-judge evaluation, and human subjective assessment. Results show that multimodal descriptions improve semantic alignment and cross-modal retrieval performance compared with modality-specific descriptions while preserving strong scene-discriminative information. The generated descriptions achieve up to 94.5% accuracy in urban scene classification, while combining audio, visual, and description embeddings further improves accuracy to 95.4%. Furthermore, the descriptions remain highly scene-discriminative even when scene labels are removed from the prompting instructions, indicating that they capture semantic information derived from the audio-visual content rather than merely reflecting label information.
Create a lesson
Related papers
ScholarCatalyst: A Benchmark for Retrieving Papers That Inspire New Research
Sohyeon Kim, Yoonho Lee, Bo Liu et al.
VISTA: A Visual Harness for Reasoning in an Interactive World
Qiushi Han, Keya Hu, Linlu Qiu et al.
A Comparative Explainability Framework for DeBERTa-v3 in Zero-Shot Medical Abstract Classification
Javier Diaz Esteban-Herreros, David Muñoz-Valero, Raquel Martínez-España et al.
Homomorphic Advantage Operator: Stabilizing Reinforcement Learning Under Fully Homomorphic Encryption Constraints
Abid Mohamed Nadhir, Ahmad Al Hanbali, Beggas Mounir
PyPottery: an AI-powered end-to-end suite for pottery processing and publication
Lorenzo Cardarelli
Causal Memory Policy: Making Memory Utility Identifiable by Intervening on Retrieval
Arman Behnam, Binghui Wang