Multi-sample Synthetic Supervision for Accent Conversion
Yangyang Qu, Michele Panariello, Massimiliano Todisco, Nicholas Evans
Abstract
Accent conversion (AC) requires changing accent while preserving speaker identity and linguistic content, yet parallel recordings are scarce. Speech synthesis provides an alternative source of supervision, but generated targets vary in accent realization and source preservation. We propose a multi-sample synthetic supervision framework that constructs conversion targets by jointly assessing these properties across candidate waveforms for each source--accent condition. Selected candidates provide associated discrete speech codes, representing linguistic content, prosody, and speaking style, as targets for accent-conditioned autoregressive adaptation. Compared with using one generated target per training example, our method improves target-accent classification accuracy by 2.49 percentage points, while speaker similarity remains nearly unchanged, and word error rate increases by 0.29 percentage points. Across six target accents, our method achieves 7.81% WER and the highest mean listening ratings among evaluated systems.
Create a lesson
Related papers
Shared-State Local Translations for Training-Free Voice Conversion
Yangyang Qu, Michele Panariello, Massimiliano Todisco et al.
Teaching LLMs to Hear Who Spoke What: Metadata-Supervised Pretraining for Encoder-Free Speech-LLMs
Mohan Shi, Ruchao Fan, Sunit Sivasankaran et al.
Code-Switching Spoken Language Identification as Multi-Label Set Prediction
Shunsuke Mitsumori, Matthew Wiesner, Shigeo Morishima et al.
PADP: Perceptual Audio Data Perturbation for Probing Perception Awareness in Audio Quality Models
Guanxin Jiang, Andreas Brendel, Pablo M. Delgado et al.
A Federated Deepfake Speech Detection Method Based on Layer-Wise Center-Guided Weighting Aggregation
Yingjian Yu, Haiyan Guo, Tianshun Wang et al.
FedCFM: Federated Continual Domain Generalization for Fake Speech Detection via Conditional Flow Matching
Yingjian Yu, Haiyan Guo, Tianshun Wang et al.