PADP: Perceptual Audio Data Perturbation for Probing Perception Awareness in Audio Quality Models
Guanxin Jiang, Andreas Brendel, Pablo M. Delgado, Jürgen Herre
Abstract
This paper presents a collection of audio transforms that introduce perceptually irrelevant distortions and demonstrates their use as perceptual stress tests for audio quality models. We refer to these methods as Perceptual Audio Data Perturbation (PADP). PADP exploits the insensitivity of the human auditory system to certain fine-grained signal variations to substantially alter the waveform while preserving perceived audio quality and content. The audibility of PADP and the selection of its parameters are evaluated through controlled listening tests, ensuring that the transformations achieve transparent or near-transparent quality for both non-critical and critical items. We further probe the robustness of state-of-the-art (SOTA) perception-motivated objective audio quality models and foundation models. The results reveal a misalignment between model responses and human auditory perception, highlighting the limited perceptual awareness of these models for certain proposed transforms.
Create a lesson
Related papers
Multi-sample Synthetic Supervision for Accent Conversion
Yangyang Qu, Michele Panariello, Massimiliano Todisco et al.
Shared-State Local Translations for Training-Free Voice Conversion
Yangyang Qu, Michele Panariello, Massimiliano Todisco et al.
Teaching LLMs to Hear Who Spoke What: Metadata-Supervised Pretraining for Encoder-Free Speech-LLMs
Mohan Shi, Ruchao Fan, Sunit Sivasankaran et al.
Code-Switching Spoken Language Identification as Multi-Label Set Prediction
Shunsuke Mitsumori, Matthew Wiesner, Shigeo Morishima et al.
A Federated Deepfake Speech Detection Method Based on Layer-Wise Center-Guided Weighting Aggregation
Yingjian Yu, Haiyan Guo, Tianshun Wang et al.
FedCFM: Federated Continual Domain Generalization for Fake Speech Detection via Conditional Flow Matching
Yingjian Yu, Haiyan Guo, Tianshun Wang et al.