ASSEMBLE: Atomic Skills for Evidence-Grounded Video Reasoning
Xiyang Wu, Zongxia Li, Shengxin Zhang, Zhichao Liu, Dinesh Manocha
Abstract
Complex video reasoning often depends on evidence scattered across distant moments, entities, and events, yet a correct answer alone does not reveal whether a model relied on the right parts of the video. We introduce ASSEMBLE, a framework that makes supporting evidence explicit throughout long-video reasoning. ASSEMBLE organizes local observations and cross-clip narratives into timestamped evidence catalogs traceable to the source video. A grounding-aware reader then composes question-specific atomic skills whose structured outputs contain explicit evidence references and support assessments. We use correctness-gated citation alignment as a direct grounding signal: after teacher-supervised fine-tuning, Group Relative Policy Optimization (GRPO) jointly optimizes answer correctness and citation alignment. This produces inspectable intermediate traces while keeping final predictions linked to explicit supporting evidence. Using a 9B reader supervised by a 235B teacher and shared precomputed evidence catalogs, ASSEMBLE achieves 59.2% macro-averaged answer accuracy across three long-video reasoning benchmarks, compared with 58.3% for Gemini-2.5-Pro, while improving macro-averaged overlap-based Grounded accuracy by 6.7%, with gains on all three benchmarks. Ablations further show that, with the same post-trained reader and inference budget, structured skill inference improves Grounded accuracy over free-form reasoning. Together, these results show that explicit evidence grounding can be integrated directly into long-video reasoning without sacrificing answer accuracy.
Create a lesson
Related papers
Supporting Perspective Acquisition and Opinion Formation on Societal Issues Through AI-Generated Japanese Rap Battle Debates
Ryota Mibayashi, Toru Urakawa, Dai Takanashi et al.
PrecipJEPA: JEPA-Regularized Future-State Prediction with Motion-Source Rendering for Precipitation Nowcasting
Yufeng Zhu, Dan Niu, Qiliang Wu et al.
Toward Generative Video Communication: A Dual-Stream Digital Transmission Framework
Bingyan Xie, Longyu Zhou, Tianhao Liang et al.
ReVR: Dual-Path Concept Reasoning for Multimodal Fake News Detection
Zhikai Tan, Yuzhou Yang, Qichao Ying et al.
TemplateCraft: Agentic Visual Template Generation
Hongjie Yu, Zhiyuan Fan, Yuzhe Zhang et al.
TempQ-Jail: Query-Constrained Candidate Ranking for Text-to-Video Jailbreak Attacks
Tianmeng Fang, Jiancheng Wang, Chen Wang et al.