MP-Bench: Evaluating Voice Agents as a Multiparty Conversation Participant
Yi-Jen Shih, Shih-Yun Shan Kuan, Guan-Ting Lin, Kai-Wei Chang, Siddhant Arora, Shu-wen Yang, Abdelrahman Mohamed, Shinji Watanabe, Hung-yi Lee, David Harwath
Abstract
Conversational voice agents have advanced significantly, offering increasingly natural human-machine interactions through both cascaded and end-to-end architectures. However, while recent benchmarks extensively evaluate dyadic interactions and passive audio comprehension, they largely overlook a prevalent real-world scenario: multi-party conversations. Evaluating agents in these settings is fundamentally more challenging than in dyadic interactions due to the exponentially greater conversational complexity. For voice agents to integrate seamlessly into human group dynamics, they must not only generate contextually appropriate responses but also demonstrate a nuanced understanding of open turn-taking. To address this gap, we introduce Multiparty Bench (MP-Bench), the first benchmark specifically designed to objectively evaluate conversational speech systems as active participants within multi-party contexts. MP-Bench assesses agent behavior along two primary dimensions: turn-taking awareness and response appropriateness. Additionally, we incorporate comprehension-based question-answering tasks as a complementary evaluation. By benchmarking 12 voice agents, we find that real-time voice agents stay at or below 22% on multiparty comprehension and remain near chance on multiparty turn-taking, exposing an open challenge for real-time voice agents under multiparty scenario.
Create a lesson
Related papers
Objective Intelligibility Prediction Using Distance Metrics on Speech Foundation Model Representations
Lyonel Behringer, Andreas Brendel
AlignDPO: Preference-Gated Alignment for Reducing Hallucination in Decoder-Only TTS
Xiao Zhou, Oisín Turbitt, Kit Bower-Morris et al.
A Device to Control and Manipulate Occlusion Effects for Own Voice Perception Studies
Rouben Rehman, Simon Kersten, Aron Schliep et al.
X-Pred MeanFlow for Streaming Token-to-Mel Speech Decoding
Hanke Xie, Xiaming Ren, Qirui Zhan et al.
Location-based Training with Complementary Folded Linear Orderings for Multichannel Speech Separation
Kaixuan Yang, Stijn Kindt, Nilesh Madhu
Overview and Meta-Analysis of DCASE 2026 Challenge Task 6: Audio Moment Retrieval from Long Audio
Hokuto Munakata, Tatsuya Komatsu, Keisuke Imoto et al.