A Federated Deepfake Speech Detection Method Based on Layer-Wise Center-Guided Weighting Aggregation
Yingjian Yu, Haiyan Guo, Tianshun Wang, Zirui Ge, Chi Liu
Abstract
The advancement of deep learning-based speech synthesis has significantly increased the diversity of deepfake speech, posing threats to voice authentication. While centralized training is effective for deepfake speech detection (DSD), it requires considerable computational resources and raises privacy concerns. To address these issues, we propose a Federated DSD (FedDSD) method that enables collaborative model training across decentralized speech datasets without sharing raw audio. Specifically, each client trains a local model using the FedProx algorithm to mitigate the effects of data heterogeneity and uploads model parameters to a central server. To improve global model aggregation, we further propose a layer-wise center-guided weighting aggregation (L-CGWA) strategy that adjusts each client's contribution per layer based on its distance to a reference center, capturing inter-client and inter-layer discrepancies and enhancing the robustness of model aggregation. Experimental results demonstrate that models trained under the proposed FedDSD method achieve equal error rates (EERs) comparable to those obtained via centralized co-training, while significantly out-performing models trained on individual corpora. Furthermore, the proposed FedDSD method demonstrates robust generalization capabilities across diverse cross-domain datasets.
Create a lesson
Related papers
Multi-sample Synthetic Supervision for Accent Conversion
Yangyang Qu, Michele Panariello, Massimiliano Todisco et al.
Shared-State Local Translations for Training-Free Voice Conversion
Yangyang Qu, Michele Panariello, Massimiliano Todisco et al.
Teaching LLMs to Hear Who Spoke What: Metadata-Supervised Pretraining for Encoder-Free Speech-LLMs
Mohan Shi, Ruchao Fan, Sunit Sivasankaran et al.
Code-Switching Spoken Language Identification as Multi-Label Set Prediction
Shunsuke Mitsumori, Matthew Wiesner, Shigeo Morishima et al.
PADP: Perceptual Audio Data Perturbation for Probing Perception Awareness in Audio Quality Models
Guanxin Jiang, Andreas Brendel, Pablo M. Delgado et al.
FedCFM: Federated Continual Domain Generalization for Fake Speech Detection via Conditional Flow Matching
Yingjian Yu, Haiyan Guo, Tianshun Wang et al.