AlignDPO: Preference-Gated Alignment for Reducing Hallucination in Decoder-Only TTS
Xiao Zhou, Oisín Turbitt, Kit Bower-Morris, Jonathan Carlton, Jamie Stacey, Kris Y. Hong
Abstract
Decoder-only text-to-speech (TTS) models scale efficiently but remain prone to content hallucinations that arise from weak text-speech alignment during autoregressive generation. We find that robustness is governed by a non-monotone relation to the sharpness of the alignment-bearing attention heads: a moderate degree is best, whereas over-sharpening is no better than the unaligned backbone and even less robust. Guided by this, we present AlignDPO, a post-training method that reaches this moderate regime by folding a lightweight connectionist-temporal-classification (CTC) alignment term into Direct Preference Optimization (DPO), applied only to the chosen samples, with no architectural or inference-time change. On the Seed-TTS-Eval English set, this significantly reduces the content-hallucination and word error rates relative to a strong DPO baseline and lowers the severe content-hallucination rate to ~0.6% (from 4.4%); a listening study further finds it preferred for naturalness over both the backbone and that baseline. Alignment is thus best learned and kept moderate rather than maximized or imposed at decoding. Audio samples are available at https://align-dpo-demo.vercel.app.
Create a lesson
Related papers
MP-Bench: Evaluating Voice Agents as a Multiparty Conversation Participant
Yi-Jen Shih, Shih-Yun Shan Kuan, Guan-Ting Lin et al.
Objective Intelligibility Prediction Using Distance Metrics on Speech Foundation Model Representations
Lyonel Behringer, Andreas Brendel
A Device to Control and Manipulate Occlusion Effects for Own Voice Perception Studies
Rouben Rehman, Simon Kersten, Aron Schliep et al.
X-Pred MeanFlow for Streaming Token-to-Mel Speech Decoding
Hanke Xie, Xiaming Ren, Qirui Zhan et al.
Location-based Training with Complementary Folded Linear Orderings for Multichannel Speech Separation
Kaixuan Yang, Stijn Kindt, Nilesh Madhu
Overview and Meta-Analysis of DCASE 2026 Challenge Task 6: Audio Moment Retrieval from Long Audio
Hokuto Munakata, Tatsuya Komatsu, Keisuke Imoto et al.