X-Pred MeanFlow for Streaming Token-to-Mel Speech Decoding
Hanke Xie, Xiaming Ren, Qirui Zhan, Jingbin Hu, Wenhao Li, Haoyu Zhang, Ruonan You, Chengyou Wang, Yunxiang Chen, Houdun Liu, Su Feng, Lei Xie
Abstract
Recent advancements in discrete token-based speech generation have highlighted the importance of efficient token-to-waveform synthesis in streaming and dialogue scenarios. Flow-matching acoustic decoders achieve high-quality token-to-mel generation, but their iterative sampling requires multiple neural function evaluations, limiting low-latency speech synthesis. MeanFlow reduces the sampling budget by modeling the average velocity over a temporal interval, yet maintaining high acoustic quality under extremely few-step token-to-mel generation remains challenging. To address this challenge, we propose X-Pred MeanFlow, a few-step streaming token-to-mel decoder that reparameterizes MeanFlow with mel-space prediction. The decoder predicts a generalized mel field and analytically derives the corresponding average velocity for sampling, thereby preserving the MeanFlow formulation while providing a direct acoustic prediction target. We further introduce layer-selective block-wise attention to enable continuous chunk-wise generation with bounded context. Experiments show that X-Pred MeanFlow improves few-step token-to-mel synthesis over Direct-u MeanFlow and supports stable streaming generation. Speech samples are available.https://renxiaming.github.io/xpred-meanflow-stream-demo
Create a lesson
Related papers
MP-Bench: Evaluating Voice Agents as a Multiparty Conversation Participant
Yi-Jen Shih, Shih-Yun Shan Kuan, Guan-Ting Lin et al.
Objective Intelligibility Prediction Using Distance Metrics on Speech Foundation Model Representations
Lyonel Behringer, Andreas Brendel
AlignDPO: Preference-Gated Alignment for Reducing Hallucination in Decoder-Only TTS
Xiao Zhou, Oisín Turbitt, Kit Bower-Morris et al.
A Device to Control and Manipulate Occlusion Effects for Own Voice Perception Studies
Rouben Rehman, Simon Kersten, Aron Schliep et al.
Location-based Training with Complementary Folded Linear Orderings for Multichannel Speech Separation
Kaixuan Yang, Stijn Kindt, Nilesh Madhu
Overview and Meta-Analysis of DCASE 2026 Challenge Task 6: Audio Moment Retrieval from Long Audio
Hokuto Munakata, Tatsuya Komatsu, Keisuke Imoto et al.