From Reasoning Failures to Composable Video Spatial Intelligence
Pengzhan Sun, Junbin Xiao, Ramanathan Rajaraman, Shiu-hong Kao, Angela Yao
Abstract
Spatial reasoning benchmarks evaluate vision-language models across diverse tasks, but task-level scores do not reveal which underlying capabilities account for success or failure. Each task requires recovering spatial evidence, representing geometry, and reasoning over it. We disentangle these capabilities by comparing predicted and ground-truth spatial context under a shared schema and coordinate contract. This comparison reveals four recurring sources of error: inaccurate perception, missing information in the spatial context, selection of the wrong measurement, and errors in reference frames or in tracking position and orientation. Guided by this diagnosis, we develop CROSS, a training-free library of typed geometric operators and spatial skills that function over available evidence to support reliable video spatial reasoning. The resulting library supplies verified context to non-coding VLMs or callable skills to a SpatialClaw agent. We evaluate on five benchmarks. raises the average score from 55.9\% to 60.2\% on ReVSI and improves the SpatialClaw result from 62.8\% to 66.3\% on DSI-Bench. These gains demonstrate that explicit handling of spatial conventions can repair systematic reasoning failures without additional training.
Create a lesson
Related papers
Moore, Escher, Penrose: A Conformal Golden Braid
Sophia Feldman, Assaf Shocher
Sphere Encoder 2
Kaiyu Yue, Sean McLeish, Ruchit Rawal et al.
One Basis to Animate Them All: Gaussian Blendshape Distillation for Real-Time Avatars
Ramazan Fazylov, Stamatis Lefkimmiatis, Ivan Laptev
ROWBench: Do Video Models Render What the Program Specifies?
Zheng-Hui Huang, Guixu Lin, Yu-Ju Tsai et al.
Embedding Prediction Helps Image Generation
Sihan Xu, Ji Xie, Zilin Wang et al.
SILSA: Sliding-Window Slice Latents for Topology-Preserving High-Resolution 3D Generation
Tianjiao Yu, Xinzhuo Li, Yifan Shen et al.