Code-Switching Spoken Language Identification as Multi-Label Set Prediction
Shunsuke Mitsumori, Matthew Wiesner, Shigeo Morishima, Shinji Watanabe
Abstract
Code-switched (CS) speech leaks through the monolingual language identification (LID) filters used to curate massive speech corpora, calling for CS-aware LID (CS-LID). We formulate utterance-level CS-LID as multi-label language-set prediction and propose a set generator that directly outputs the languages in an utterance, comparing it against atomic-pair and score-based classification baselines. Oracle Top-k is the strongest baseline, but thresholding fails because no single threshold separates CS from monolingual speech. Our set generator predicts the correct language count on unseen pairs without assuming the number of languages, but underperforms oracle Top-k in exact set accuracy. Our analysis identifies the key obstacles to robust CS-LID: oracle cardinality, threshold instability, language bias in CS training data, and the synthetic-to-real gap.
Create a lesson
Related papers
Multi-sample Synthetic Supervision for Accent Conversion
Yangyang Qu, Michele Panariello, Massimiliano Todisco et al.
Shared-State Local Translations for Training-Free Voice Conversion
Yangyang Qu, Michele Panariello, Massimiliano Todisco et al.
Teaching LLMs to Hear Who Spoke What: Metadata-Supervised Pretraining for Encoder-Free Speech-LLMs
Mohan Shi, Ruchao Fan, Sunit Sivasankaran et al.
PADP: Perceptual Audio Data Perturbation for Probing Perception Awareness in Audio Quality Models
Guanxin Jiang, Andreas Brendel, Pablo M. Delgado et al.
A Federated Deepfake Speech Detection Method Based on Layer-Wise Center-Guided Weighting Aggregation
Yingjian Yu, Haiyan Guo, Tianshun Wang et al.
FedCFM: Federated Continual Domain Generalization for Fake Speech Detection via Conditional Flow Matching
Yingjian Yu, Haiyan Guo, Tianshun Wang et al.