VibeVoice-ASR-Streaming Technical Report
Yujie Tu, Zhiliang Peng, Jianwei Yu, Li Dong, Songchen Xu, Yaoyao Chang, Wenhui Wang, Zilong Wang, Zehua Wang, Yan Xia, Jiajun Zhang, Xie Chen, Furu Wei
Abstract
Traditional speaker-attributed ASR systems treated ASR and speaker diarization as two separate tasks. Recently, end-to-end models such as VibeVoice-ASR have unified the two tasks within a single model. However, existing unified models still mainly support offline recognition, making it difficult to meet the low-latency requirements of real-time voice assistants and agents. To tackle this issue, we present VibeVoice-ASR-Streaming, one of the first LLM-based end-to-end approaches to streaming speaker-attributed ASR. It interleaves fixed-size audio chunks, a small amount of lookahead audio and previous text. This allows the model to produce ''who said what'' as speech arrives, without a separate diarization stage. For transcription accuracy, our 7B model achieves the lowest average WER/CER across five evaluation sets. For speaker attribution, it achieves the best or tied-best on 12 of 13 evaluation settings. We release the 1.5B and 7B model weights together with inference code.
Create a lesson
Related papers
VAANI Noise Event Dataset: A curated spontaneous speech dataset annotated with timestamps for noise events
Pavan Kumar J, Agneedh Basu, Pranav Bhat et al.
Sensing Bone-Conducted Speech with Earbuds
Christoph Weyer, Peter Jax
TAG-Bench: Benchmarking Temporal Audio Grounding in Large Audio Language Models
Yuhang Dai, Xin Shu, Zengxi Li et al.
Ontology-based Target Sound Extraction
Carlos Hernandez-Olivan, Marc Delcroix, Tsubasa Ochiai et al.
U-PAST: A Phase-Aware Audio Spectrogram Transformer-U-Net for Single-Channel Speech Enhancement
Cao Duong Ly, Jörn Anemüller
Likelihood-Constrained Acoustic Reranking for Training-Free Hallucination Mitigation in LLM-Based ASR
Jiasheng Kuang, Linru Zheng, Hongjin Song et al.