Ontology-based Target Sound Extraction
Carlos Hernandez-Olivan, Marc Delcroix, Tsubasa Ochiai, Naohiro Tawara, Shoko Araki
Abstract
Target sound extraction (TSE) aims to isolate a sound source of interest from a mixture, given a semantic query. Existing TSE systems are conditioned on fixed class representations tied to individual sound categories, limiting their ability to handle the hierarchical relationships that naturally organize environmental sounds. In this paper, we introduce ontology-based TSE, a new task formulation in which a single model extracts sounds queried at any level of a sound ontology, from fine-grained leaf classes such as cat and dog to high-level categories such as animal. We propose a learnable class embedding table defined over all nodes of an AudioSet-derived ontology, regularized with a Cophenetic Correlation Coefficient (CPCC) loss that aligns embedding distances with shortest-path distances in the ontology tree. Our experiments across different approaches show the benefit of considering the ontology structure when training TSE systems.
Create a lesson
Related papers
VibeVoice-ASR-Streaming Technical Report
Yujie Tu, Zhiliang Peng, Jianwei Yu et al.
VAANI Noise Event Dataset: A curated spontaneous speech dataset annotated with timestamps for noise events
Pavan Kumar J, Agneedh Basu, Pranav Bhat et al.
Sensing Bone-Conducted Speech with Earbuds
Christoph Weyer, Peter Jax
TAG-Bench: Benchmarking Temporal Audio Grounding in Large Audio Language Models
Yuhang Dai, Xin Shu, Zengxi Li et al.
U-PAST: A Phase-Aware Audio Spectrogram Transformer-U-Net for Single-Channel Speech Enhancement
Cao Duong Ly, Jörn Anemüller
Likelihood-Constrained Acoustic Reranking for Training-Free Hallucination Mitigation in LLM-Based ASR
Jiasheng Kuang, Linru Zheng, Hongjin Song et al.