Tracing the Origins: Legacy Codec Identification in Neural Audio Transcoding
Wonje Heo, Shinee Youn, Yooshin Kim, Chuck Chae, Donghoon Shin
Abstract
Residual Vector Quantization (RVQ)-based neural audio codecs (NACs) enable high-fidelity audio distribution at unprecedentedly low bitrates through discrete token-based representations. However, this shift disrupts traditional forensics, as non-linear neural transcoding obscures the underlying traces of legacy compression. This study defines the forensic gap and proposes a Transformer-based framework designed to leverage the hierarchical and temporal dependencies inherent in RVQ sequences. By modeling inter-layer causal relationships and dynamic forensic significance, our model effectively disentangles superimposed artifacts from legacy-to-neural transcoding. Experimental results achieve 97%+ accuracy for codec identification and robust joint identification performance across 32-128 kbps. These results demonstrate that traditional codec traces persist even after neural transcoding, supporting the feasibility and necessity of neural-codec-aware audio forensics.
Create a lesson
Related papers
Multi-Dimensional Prosody Judgment For Live Streaming Speech Synthesis
Zifan Guan, Longyu Lu, Junan Zhang et al.
CircleMatch: Prototype Matching with Circular Temporal Statistics for Tiny Keyword Spotting
Jiajun Sun, Zhe Gao
Robust Workflow Generation via Adversarial Learning for Audio Deepfake Detection
Xiang Li, Pin-Yu Chen, Wenqi Wei
CoRELoop: Parameter-Efficient Controlled Recurrent Refinement for Audio Deepfake Detection
Kunyu Feng, Yuxiang Wang, Li Wang et al.
A Cross-Lingual Acoustic Disease-Alignment Framework for Respiratory Health Assessment from Spontaneous Speech
Roksana Khanom, Raghib Asfak Tasnim, Bodrun Nahar Bithi et al.
A State-Space Model of Figured-Bass Realization: Local Constraints, Coupled Voices, and Polynomial-Time Solvability
Evan Unit Lim