SpaceVLA: Spatially Grounded VLA for Robotic Manipulation with User-Authored Grasp and Place Anchors
Daniia Zinniatullina, Iaroslav Kolomiets, Mikhail Konenkov, Miguel Altamirano Cabrera, Dzmitry Tsetserukou
Abstract
Vision-language-action (VLA) models follow language commands but often lack explicit spatial intent for manipulation. We present Visual Intent Anchors, an XR pipeline that lets users specify grasp and placement regions and renders them as image-space overlays for VLA control. We collect 200 Unity pick-and-place demonstrations and fine-tune OpenVLA-7B with LoRA on temporally subsampled annotated observations. The policy predicts tokenized 7-DoF incremental actions from marked RGB observations and language. We evaluate the policy in closed-loop Unity trials, achieving a grasp success rate of 91.25% and mean grasp and placement errors of 0.5 cm and 0.7 cm, respectively.
Create a lesson
Related papers
Calmables: Demonstrating Closed-Loop Infrared Earables for Thermal Biofeedback and Relaxation Support
Valeria Zitz, Michael Küttner, Jonas Hummel et al.
"Okay, I've Actually Softened My Take on This": How People in Decentralized Social Media Reason about the Appropriateness of Generative AI
Romina Mahinpei, Manoel Horta Ribeiro, Andrés Monroy-Hernández et al.
Integrating Flipped Learning and Generative AI for Practice-Based Design Education: Evidence from a Knit Yarn Design Course
Hong Qu, Zichao Ling, Yadie Yang
EasyFashion: A Human-AI Co-Creation System for Personalized Fashion Design and Sewing Pattern Generation
Hong Qu, Zhaoxiang Xu, Jinbo Luo et al.
Verify, Offload, Extend & Recommend: Selective Complementarity in AI Support for Physical Activity Planning with Longitudinal Patient Data
Pavithren V S Pakianathan, Rania Islambouli, Diogo Branco et al.
Building a Cultural Perspective on Doctor-Patient Conversations
Krithi Shailya, Siddharth D Jaiswal, Ashish Makani et al.