Skip to content

Objective Intelligibility Prediction Using Distance Metrics on Speech Foundation Model Representations

Lyonel Behringer, Andreas Brendel

eess.ASarXiv:2609.13046

Abstract

High-dimensional representations of pretrained speech foundation models have proven beneficial for objective speech quality and intelligibility prediction. While existing work on neural intelligibility prediction usually leverages such representations for task-specific fine-tuning, in this work we evaluate the usefulness of such representations for intelligibility prediction without any further training. We conduct a layer-wise analysis of multiple speech foundation models, correlating various embedding distances with subjective intelligibility scores. The results show that embeddings extracted from Whisper speech recognition models are best suited, with the last encoder and decoder layers yielding the best correlations when using the Fréchet Audio Distance. Notably, the evaluated distances outperform classical intelligibility metrics and are more robust than Word and Character Error Rates. Further, correlations improve with increasing size of the Whisper model from which embeddings are extracted.

Create a lesson