Shared-State Local Translations for Training-Free Voice Conversion
Yangyang Qu, Michele Panariello, Massimiliano Todisco, Nicholas Evans
Abstract
In one-shot training-free voice conversion (VC), the source and reference utterances may contain different linguistic content, so reliable frame-level correspondence between them cannot be assumed. We propose StateVC, which jointly defines a common set of local regions from pooled frame-level WavLM representations of the source and reference utterances; we refer to these regions as states. These shared states are obtained by fitting a pair-specific Gaussian mixture model to the pooled representations, without explicit source--reference frame matching. Within each state, StateVC estimates a source-to-reference mean shift in the original WavLM space. Source-frame posterior probabilities then combine the state-specific shifts so that different frames can receive different local updates. For the LibriSpeech one-shot protocol, StateVC achieves the lowest word error rate (WER) and character error rate (CER) among the evaluated systems, at 8.01% and 3.22%, respectively, with a speaker similarity (SIM) of 0.9512. It also achieves the highest mean perceived speaker similarity among the evaluated systems and the highest mean naturalness among the evaluated training-free systems.
Create a lesson
Related papers
Multi-sample Synthetic Supervision for Accent Conversion
Yangyang Qu, Michele Panariello, Massimiliano Todisco et al.
Teaching LLMs to Hear Who Spoke What: Metadata-Supervised Pretraining for Encoder-Free Speech-LLMs
Mohan Shi, Ruchao Fan, Sunit Sivasankaran et al.
Code-Switching Spoken Language Identification as Multi-Label Set Prediction
Shunsuke Mitsumori, Matthew Wiesner, Shigeo Morishima et al.
PADP: Perceptual Audio Data Perturbation for Probing Perception Awareness in Audio Quality Models
Guanxin Jiang, Andreas Brendel, Pablo M. Delgado et al.
A Federated Deepfake Speech Detection Method Based on Layer-Wise Center-Guided Weighting Aggregation
Yingjian Yu, Haiyan Guo, Tianshun Wang et al.
FedCFM: Federated Continual Domain Generalization for Fake Speech Detection via Conditional Flow Matching
Yingjian Yu, Haiyan Guo, Tianshun Wang et al.