SsgCaps: A controlled dataset for the evaluation of sound scene generation algorithms
Modan Tailleur, Junwon Lee, Laurie M Heller, Mathieu Lagrange, Keunwoo Choi, Brian McFee, Keisuke Imoto, Yuki Okamoto
Abstract
Sound Scene Generation is about the automatic synthesis of artificial sound scenes. We introduce SsgCaps, a publicly available dataset of human-engineered sound scenes wherein each scene matches a precisely structured prompt that guides the sampling process. The corresponding prompts are sampled from a predefined action-based typology that allows extensive sampling while retaining plausibility. SsgCaps is a sound scene dataset derived from the unpublished reference dataset for Task 7 of the 2024 DCASE Challenge edition, which contained private-and public-domain audio samples. In contrast, SsgCaps contains only public-domain audio samples, allowing us to open this dataset to the community. To make this dataset useful to the community, we first elaborate on the rationale for the prompt and dataset structure. We then perform a comparative quantitative analysis of the 2 versions of the dataset. To do so, we compare both versions to the audio synthesized by the SSG algorithms submitted to the challenge using Fréchet Audio Distance (FAD) and Kernel Audio Distance (KAD) as well as perceptual ratings. This analysis shows only small differences, which enables us to recommend the open version for further benchmarking of SSG algorithms.
Create a lesson
Related papers
ARS-Avatar: Animatable and Relightable Surfel Avatars with Learnable Ambient Occlusion
Jiateng Liu, Hao Gao, Junxin Sun et al.
Constrained Program Generation for 3D Reaction Animation with a 0.8B Model
Hongyuan Wang, Daming Luo, Nico Pietroni et al.
MoSAT: Human Motion Generation from Spatial Audio and Textual Description
Shuyang Xu, Zhiyang Dou, Yiduo Hao et al.
Edge-centric Brain Transformer: An Edge-centric Functional Connectivity Learning Framework for fMRI-based Brain Disorder Diagnosis
Dengyi Zhao, Zhiheng Zhou, Mengyao Zhou et al.
Physically Based Rendering in the Latent Space
Vuk Radovanovic, Vishesh Gupta, Adrien Gruson et al.
S4R: Scaling for Rigid-Body Interpenetration Resolution
Zhiyang Dou, Ang Zhao, Chen Peng et al.