MoWAM: Explicit Future Motion Prediction for Efficient World Action Models
Jiayu Wang, Bin Zhu, Yue Yu, Jingjing Chen
Abstract
World Action Models (WAMs) improve robot policy learning by incorporating future dynamics, yet explicitly generating future videos at inference introduces substantial computational overhead. Removing future generation improves efficiency, but leaves future dynamics only implicitly encoded in observation features, which can limit robustness under distribution shifts. We propose MoWAM, an efficient WAM that replaces future video generation with explicit future motion prediction. Instead of reconstructing the complete future scene, MoWAM models structured robot motion as a compact abstraction of the future, capturing how the robot is expected to evolve under the current scene and interaction constraints. A Mixture-of-Transformer architecture learns future visual dynamics during training while jointly predicting motion and action, allowing video generation to be removed entirely at inference while retaining an explicit representation of the future. The compact motion representation further enables efficient inference-time scaling by sampling multiple candidates of motion and action pairs and selecting among them with a motion-aware task-progress verifier. Experiments on LIBERO, LIBERO-Plus, and real-world manipulation tasks demonstrate that MoWAM achieves strong in-distribution performance, improved out-of-distribution robustness, and higher average real-world success than representative WAM baselines. In addition, performance improves as more candidates are explored, demonstrating that explicit future motion provides an effective and efficient basis for inference-time scaling.
Create a lesson
Related papers
GeoAAC: Geometry-Based Adaptive Action Chunking from Denoising Trajectories in VLA Policies
Xin Chen, Sen Chen, Yujuan Ding et al.
Agile-WAM: An Agile Tactile World Action Model for Contact-Rich Robot Control
Hanchu Zhou, Brendan Lynch, Raman Goyal et al.
OPTED: On-Policy Fine-Tuning for End-to-End Driving using a Render-Free Teacher
Damiano Da Col, Maximilian Igl, Peter Karkus et al.
MILER: Semantic Mid-Level Representation for Sim-to-Real Reinforcement Learning in Unstructured Autonomous Driving
Thomas Steinecker, Denis Trescher, Alexander Bienemann et al.
Underwater Visual Target Tracking with Target-Specific Depth Estimation and Adaptive Model-Fusion Predictive Control
Yuheng Zhou, Haiyang Cheng, Yanqi Feng et al.
HOPHY: A Hierarchical Hypergraph Representation for Off-Road Path and Mission Planning
Pranay Meshram, Charuvahan Adhivarahan, Prithvi Poddar et al.