Training DeepFilterNet with Accurate Room Acoustic Simulations Improves Single-Channel Speech Enhancement
Alessia Milo, Georg Götz, Steinar Guðjónsson, Daniel Gert Nielsen, Jesper Pedersen, Finnur Pind
Abstract
We investigate how the realism of synthetic room impulse response (RIR) datasets affects the training of DeepFilterNet3 for single-channel speech enhancement. We compare a DNS4 image-source-method (ISM) RIR dataset with a higher-acoustic-fidelity dataset generated using hybrid wave-based and geometrical acoustics simulation. Rather than isolating individual simulation factors, we compare complete RIR generation pipelines while keeping the enhancement model unchanged. Models are evaluated on unseen measured RIRs using objective speech enhancement metrics and downstream automatic speech recognition (ASR). Training with the higher-fidelity dataset consistently yields modest improvements in objective metrics and substantially lower ASR word error rates than the ISM dataset. Although the experiments do not attribute these gains to individual modelling components, they show that increasing the overall realism of synthetic acoustic training data improves the generalization of DeepFilterNet3 to unseen measured environments.
Create a lesson
Related papers
GrainSpeech: Less Context, More Detail for Compact Speech Synthesis
Zitao Liang, Chang Gao
Absolute Quality Ratings of Speech Enhancement Systems by Listeners of Different Ages and Degrees of Hearing Loss
Matteo Torcoli, Chih-Wei Wu, Andrea Esposito et al.
Mask-Based Speech Enhancement for Spatial Audio: A Comparison of Ambisonics, Beamforming, and Microphone Channels
Sheli Hendel, Boaz Rafaely, Dorothea Kolossa
Reviving Etter method for autoregressive inpainting: Generalization, evaluation, implementation
Ondřej Mokrý, Matěj Hrdlička, Pavel Rajmic
Correlation-Guided Encoder Selection for Multi-Encoder Large Audio-Language Models
Pei-Jun Liao, Hung-Shin Lee, Wenze Ren et al.
Task-oriented neural FOA encoding for SELD from irregular microphone arrays
Jiachen Liu, Yin Cao, Ming Wu et al.