Agile-WAM: An Agile Tactile World Action Model for Contact-Rich Robot Control
Hanchu Zhou, Brendan Lynch, Raman Goyal, Dechen Gao, Begum Kasap, Boqi Zhao, Junshan Zhang
Abstract
World Action Models (WAMs) advance beyond conventional visuomotor policies by jointly predicting future world states and robot actions, enabling the policy to learn physical dynamics that support effective control. However, recent tactile WAMs often rely on large-scale pretrained generative backbones to capture contact-rich physical dynamics, which limit their inference efficiency and flexible deployment. In this paper, we present , an agile tactile World Action Model for contact-rich robot control. encodes visual and tactile observations into a shared latent that serves as the source of a direct vision-tactile-to-action flow-matching process, which can jointly generate latent representations of action chunks and future visual/tactile latents. A key observation is that vision and tactile signals evolve at inherently different timescales: adjacent visual frames are often highly similar, whereas tactile signals can change abruptly upon contact. We therefore introduce multi-horizon multimodal prediction in , which provides supervision for visual latent at a larger temporal offset while predicting the tactile latent in the next frame to capture fine-grained contact dynamics. Across nine simulated and five real-world contact-rich manipulation tasks, demonstrates strong and robust performance, outperforming the strongest baseline in success rate while maintaining low inference latency. In particular, in five real-world experiments, yields a relative gain of 29.4\% in overall success rates while achieving inference latency of 11.9 ms. These results demonstrate that multimodal WAM can be achieved with an agile architecture suitable for precise and high-frequency robot control. More details are available on our project page: https://hanchuzhou.github.io/TAROprojectpage/.
Create a lesson
Related papers
GeoAAC: Geometry-Based Adaptive Action Chunking from Denoising Trajectories in VLA Policies
Xin Chen, Sen Chen, Yujuan Ding et al.
OPTED: On-Policy Fine-Tuning for End-to-End Driving using a Render-Free Teacher
Damiano Da Col, Maximilian Igl, Peter Karkus et al.
MILER: Semantic Mid-Level Representation for Sim-to-Real Reinforcement Learning in Unstructured Autonomous Driving
Thomas Steinecker, Denis Trescher, Alexander Bienemann et al.
Underwater Visual Target Tracking with Target-Specific Depth Estimation and Adaptive Model-Fusion Predictive Control
Yuheng Zhou, Haiyang Cheng, Yanqi Feng et al.
MoWAM: Explicit Future Motion Prediction for Efficient World Action Models
Jiayu Wang, Bin Zhu, Yue Yu et al.
HOPHY: A Hierarchical Hypergraph Representation for Off-Road Path and Mission Planning
Pranay Meshram, Charuvahan Adhivarahan, Prithvi Poddar et al.