CAL-MOS: Bridging Layers with Adapters for Robust MOS Prediction Across Speech Foundation Models
Alef Iury Siqueira Ferreira, Pedro Lustosa Rege Botelho, Fernanda Silva, Daniel Casanova, Rafael Faustino, Frederico Oliveira, Arlindo Galvão Filho, Anderson da Silva Soares
Abstract
Speech Quality Assessment (SQA) is essential for modern speech technologies, and recent non-intrusive SQA predictors increasingly rely on Speech Foundation Models (SFMs). However, because SFMs expose representations from many layers, it remains unclear which depths are most informative for MOS prediction and how multi-layer information should be combined reliably across backbones and datasets. We benchmark ten SFMs on four MOS datasets under three regimes: full fine-tuning, last-layer probing with a frozen encoder, and naive cross-layer weighted aggregation. We find that the best layer is strongly backbone- and dataset-dependent, and that naive weighted fusion can be unstable across settings. We further evaluate a layer-calibrated aggregation variant that applies per-layer adapters before pooling, which improves the robustness of multi-layer fusion and narrows the gap to full fine-tuning while keeping the backbone frozen.
Create a lesson
Related papers
Multi-Dimensional Prosody Judgment For Live Streaming Speech Synthesis
Zifan Guan, Longyu Lu, Junan Zhang et al.
CircleMatch: Prototype Matching with Circular Temporal Statistics for Tiny Keyword Spotting
Jiajun Sun, Zhe Gao
Robust Workflow Generation via Adversarial Learning for Audio Deepfake Detection
Xiang Li, Pin-Yu Chen, Wenqi Wei
CoRELoop: Parameter-Efficient Controlled Recurrent Refinement for Audio Deepfake Detection
Kunyu Feng, Yuxiang Wang, Li Wang et al.
A Cross-Lingual Acoustic Disease-Alignment Framework for Respiratory Health Assessment from Spontaneous Speech
Roksana Khanom, Raghib Asfak Tasnim, Bodrun Nahar Bithi et al.
A State-Space Model of Figured-Bass Realization: Local Constraints, Coupled Voices, and Polynomial-Time Solvability
Evan Unit Lim