ATI-VLA: Action-Centric Predictive Vision-Language-Action Models via Actionable Alignment Then Adaptive Injection
Yijie Zhu, Rui Shao, Jie He, Wei Li, Bo Zhao, Yelin Wang, Xiaochen Yuan, Tao Tan, Miao Zhang, Xiaojiang Peng, Zitong Yu
Abstract
Predictive Vision-Language-Action (VLA) models aim to improve robotic manipulation via future observation or world dynamics forecasting. However, existing approaches often fail to realize this potential and underperform direct action prediction models. We argue that these limitations stem from modality misalignment between observations and actions, together with joint optimization conflicts that drive learning away from an action-centric objective. To this end, we introduce ATI-VLA, an Action-Centric Predictive Vision-Language-Action framework via Actionable Alignment Then Adaptive Injection. Specifically, it follows a two-step design: 1) Actionable Representation Alignment via a Shared Codebook. It aligns predictive observation and action representations by mapping both modalities into a shared discrete latent space via a unified codebook, making predictive observation latents readily usable for action generation and mitigating modality misalignment. 2) Action-Centric Adaptive Injection of Predictive Latents. Building upon this, it then injects predictive observation latents into action decoding as explicit predictive priors via a lightweight adaptive side-path, enabling adaptive predictive guidance under a single action-centric objective. Extensive experiments on both simulation and real-world robotic tasks demonstrate that ATI-VLA achieves state-of-the-art performance with faster convergence.
Create a lesson
Related papers
Moore, Escher, Penrose: A Conformal Golden Braid
Sophia Feldman, Assaf Shocher
Sphere Encoder 2
Kaiyu Yue, Sean McLeish, Ruchit Rawal et al.
One Basis to Animate Them All: Gaussian Blendshape Distillation for Real-Time Avatars
Ramazan Fazylov, Stamatis Lefkimmiatis, Ivan Laptev
ROWBench: Do Video Models Render What the Program Specifies?
Zheng-Hui Huang, Guixu Lin, Yu-Ju Tsai et al.
Embedding Prediction Helps Image Generation
Sihan Xu, Ji Xie, Zilin Wang et al.
SILSA: Sliding-Window Slice Latents for Topology-Preserving High-Resolution 3D Generation
Tianjiao Yu, Xinzhuo Li, Yifan Shen et al.