Multi-Dimensional Prosody Judgment For Live Streaming Speech Synthesis
Zifan Guan, Longyu Lu, Junan Zhang, Zhizheng Wu, Meiguang Jin, Junfeng Ma
Abstract
Evaluating live streaming speech synthesis (TTS) requires assessing fine-grained, highly expressive prosody such as emotion, intonation, and energy which traditional MOS predictors fail to capture. While proprietary Large Language Models (LLMs) like Gemini can evaluate these aspects, they are too costly for massive inference and reinforcement learning feedback. To address this, we first introduce Live-ProsodyJudge (LPJ), a cost-effective pairwise evaluator distilled from Gemini into Qwen3-Omni. However, we identify a critical flaw in standard multi-dimensional evaluation: verdict coupling. The judge tends to lazily align all individual dimension scores with its overall preference, collapsing a rich multi-dimensional rubric into a single preference bit. To resolve this, we further propose Decoupled-Live-ProsodyJudge (D-LPJ). D-LPJ eliminates the overall verdict target to prevent blind following, masks uncertain pair-dimensions during Supervised Fine-Tuning(SFT), and introduces a novel span-local GRPO strategy that applies normalized advantages strictly to their corresponding rationale spans. Evaluated on highly curated human-annotated test sets, 10 sample balanced-order LPJ achieves higher point accuracy than a single Gemini call, while D-LPJ successfully produces independent,decoupled dimension judgments. Furthermore, in a Best-of-8 TTS candidate selection tournament, the LPJ-selected utterance falls within the human top-3 in 85.29% of high-confidence cases, demonstrating its efficacy for fine-grained TTS preference optimization.
Create a lesson
Related papers
CircleMatch: Prototype Matching with Circular Temporal Statistics for Tiny Keyword Spotting
Jiajun Sun, Zhe Gao
Robust Workflow Generation via Adversarial Learning for Audio Deepfake Detection
Xiang Li, Pin-Yu Chen, Wenqi Wei
CoRELoop: Parameter-Efficient Controlled Recurrent Refinement for Audio Deepfake Detection
Kunyu Feng, Yuxiang Wang, Li Wang et al.
A Cross-Lingual Acoustic Disease-Alignment Framework for Respiratory Health Assessment from Spontaneous Speech
Roksana Khanom, Raghib Asfak Tasnim, Bodrun Nahar Bithi et al.
A State-Space Model of Figured-Bass Realization: Local Constraints, Coupled Voices, and Polynomial-Time Solvability
Evan Unit Lim
FRAUDSkill: Structured Frozen-Weight Skill Optimization for Audio Anti-Fraud Detection
Chengxian Hu, Zhiming Ma, Mingjun Pan et al.