Towards Real-world Environment-aware Zero-shot Text-to-speech Synthesis via Disentangled Audio Infilling
Ye-Xin Lu, Xin Wang, Yang Ai, Hui-Peng Du, Zhen-Hua Ling, Junichi Yamagishi
Abstract
Recent zero-shot text-to-speech (TTS) systems achieve remarkable naturalness and speaker similarity but typically require high-quality speaker prompts and either strip away or entangle the acoustic environment with speaker characteristics, limiting their real-world applicability. We present an extended DAIEN-TTS, an environment-aware zero-shot TTS framework that disentangles and jointly models speech, background noise, and reverberation, enabling independent control over timbre and acoustic environment through separate speaker and environment prompts. Built upon the flow-matching-based F5-TTS, it uses a speech-environment separation module to decompose environmental speech into speech, noise, and reverberation components, which are injected into the Diffusion Transformer for environment-aware generation. Training uses simulated data constructed by mixing clean speech with noise and room impulse responses, together with a cross-speaker conditioning strategy that suppresses speaker information leakage from the environment branch. When real-world data are available, the system can be further fine-tuned to bridge the simulated-to-real domain gap.At inference, a triple classifier-free guidance mechanism enables fine-grained control over speech, noise, and reverberation, and a signal-to-noise-ratio adaptation strategy aligns the synthesized speech with the environment prompt. Experiments on simulated and real-world test sets show that DAIEN-TTS generates environmental personalized speech with high naturalness, strong speaker similarity, and faithful noise and reverberation reproduction, while offering controllability beyond prior environment-aware TTS systems.
Create a lesson
Related papers
GrainSpeech: Less Context, More Detail for Compact Speech Synthesis
Zitao Liang, Chang Gao
Absolute Quality Ratings of Speech Enhancement Systems by Listeners of Different Ages and Degrees of Hearing Loss
Matteo Torcoli, Chih-Wei Wu, Andrea Esposito et al.
Mask-Based Speech Enhancement for Spatial Audio: A Comparison of Ambisonics, Beamforming, and Microphone Channels
Sheli Hendel, Boaz Rafaely, Dorothea Kolossa
Reviving Etter method for autoregressive inpainting: Generalization, evaluation, implementation
Ondřej Mokrý, Matěj Hrdlička, Pavel Rajmic
Correlation-Guided Encoder Selection for Multi-Encoder Large Audio-Language Models
Pei-Jun Liao, Hung-Shin Lee, Wenze Ren et al.
Task-oriented neural FOA encoding for SELD from irregular microphone arrays
Jiachen Liu, Yin Cao, Ming Wu et al.