Objective Intelligibility Prediction Using Distance Metrics on Speech Foundation Model Representations
Lyonel Behringer, Andreas Brendel
Abstract
High-dimensional representations of pretrained speech foundation models have proven beneficial for objective speech quality and intelligibility prediction. While existing work on neural intelligibility prediction usually leverages such representations for task-specific fine-tuning, in this work we evaluate the usefulness of such representations for intelligibility prediction without any further training. We conduct a layer-wise analysis of multiple speech foundation models, correlating various embedding distances with subjective intelligibility scores. The results show that embeddings extracted from Whisper speech recognition models are best suited, with the last encoder and decoder layers yielding the best correlations when using the Fréchet Audio Distance. Notably, the evaluated distances outperform classical intelligibility metrics and are more robust than Word and Character Error Rates. Further, correlations improve with increasing size of the Whisper model from which embeddings are extracted.
Create a lesson
Related papers
MP-Bench: Evaluating Voice Agents as a Multiparty Conversation Participant
Yi-Jen Shih, Shih-Yun Shan Kuan, Guan-Ting Lin et al.
AlignDPO: Preference-Gated Alignment for Reducing Hallucination in Decoder-Only TTS
Xiao Zhou, Oisín Turbitt, Kit Bower-Morris et al.
A Device to Control and Manipulate Occlusion Effects for Own Voice Perception Studies
Rouben Rehman, Simon Kersten, Aron Schliep et al.
X-Pred MeanFlow for Streaming Token-to-Mel Speech Decoding
Hanke Xie, Xiaming Ren, Qirui Zhan et al.
Location-based Training with Complementary Folded Linear Orderings for Multichannel Speech Separation
Kaixuan Yang, Stijn Kindt, Nilesh Madhu
Overview and Meta-Analysis of DCASE 2026 Challenge Task 6: Audio Moment Retrieval from Long Audio
Hokuto Munakata, Tatsuya Komatsu, Keisuke Imoto et al.