Accommodating Sample Size Effect on Similarity Measures in Speaker Clustering
Alexander Haubold, John R. Kender
Abstract
We investigate the symmetric Kullback-Leibler (KL2) distance in speaker clustering and its unreported effects for differently-sized feature matrices. Speaker data is represented as Mel Frequency Cepstral Coefficient (MFCC) vectors, and features are compared using the KL2 metric to form clusters of speech segments for each speaker. We make two observations with respect to clustering based on KL2: 1.) The accuracy of clustering is strongly dependent on the absolute lengths of the speech segments and their extracted feature vectors. 2.) The accuracy of the similarity measure strongly degrades with the length of the shorter of the two speech segments. These effects of length can be attributed to the measure of covariance used in KL2. We demonstrate an empirical correction of this sample-size effect that increases clustering accuracy. We draw parallels to two Vector Quantization-based (VQ) similarity measures, one which exhibits an equivalent effect of sample size, and the second being less influenced by it.
Create a lesson
Related papers
Direct or Mediated? Task-Dependent Audio Information Routing in Large Audio Language Models
Yizhou Zhang, Wangjin Zhou, Xin Gu et al.
SpeechGym: An Audio-Native Gym for Training Voice Agents via Reinforcement Learning
Jiajun Fan, Jingyuan Li, Prashanth Gurunath Shivakumar et al.
AudioSpan: Spanning the Duration and Depth of Audio Comprehension
Wen Huang, Yunfei Chu, Meng Gao et al.
Decay-Region Group Delay as a Forensic Cue for AI-Generated Impulsive Sounds
JaeHyeong Chang, Chengzhe Sun, Siwei Lyu
StreamAV-Bench: A Comprehensive Benchmark for Streaming Audio-Video Generation
Kaiqi Liu, Haoxuan Zeng, Jingqi Liu et al.
Attention-Guided Reliability Scaling for Contrastive Decoding in Robust Audio-Visual Speech Recognition
YoungChae Kim, Da-Hee Yang, Joon-Hyuk Chang