Q-SPT: Learnable Query-Based Compression for Low-Frame-Rate Speech Tokenization
Jeeyoung Yun, Seohwan Yun, Sungwoong Kim
Abstract
Neural speech codecs increasingly serve as tokenizers for speech language models (SLMs). Lowering the frame rate reduces the computational and memory costs of SLMs, but makes it difficult to preserve both linguistic information and acoustic detail. Existing approaches rely on rule-based compression: average pooling can discard linguistic information, whereas similarity-based merging uses a fixed threshold on adjacent-frame similarity and applies the resulting boundaries to the acoustic stream. We propose Q-SPT, a low-frame-rate dual-stream speech tokenizer with separate, context-aware, learnable query-based compressors specialized for semantic and acoustic representations. In particular, queries at a fixed rate independently attend to the semantic and acoustic streams as separate key-value sources, enabling stream-specific, context-aware aggregation through two separately learned compressors. In addition, an autoregressive text loss explicitly supervises the semantic compressor to preserve linguistic information. Experimental results show that Q-SPT achieves the best reconstruction among the evaluated codecs at the same frame rate. In downstream SLMs, it yields the best speech recognition accuracy and text-to-speech perceptual quality with competitive intelligibility.
Create a lesson
Related papers
LAST: Looped Audio Spectrogram Transformer
Haider Al-Tahan, Sean O'Brien, Anastasia Razdaibiedina et al.
From Isolated Feature to Orbits: Discovering Music Concepts via Multi-SAE Alignment
Liwei Lin, Gus Xia
Beyond Decodability: Do Acoustic Factors Drive Predictions in Speech-Based Alzheimer's Assessment?
Serli Kopar, Alkis Koudounas, Roshan P. Rane et al.
Multi-Party Backchannel Prediction: a Diagnosis, a Benchmark, and a Ceiling
Mohammed Hafsati, Ahmed Loughzali
AudioJev: Direct Audio Decisions with Order-Calibrated Probabilities
Sihan Lv, Zhen Li, Zhiqi Cao et al.
Toward Elastic Speech Inference: Training-Free Wake-Word Detection from Pretrained ASR
Hwayeon Kim, Youngwon Choi, Hyeonyu Kim