Dynin-Robotics: Omnimodal Unified Diffusion Vision-Language-Action Model
Hoeun Lee, Jaeik Kim, Jusang Oh, Jinhyeok Kim, Geon Choi, Hyeonggeun Kim, Jaeyoung Do
Abstract
Visual goal and dynamics prediction can provide language-conditioned robot policies with both a target outcome and a representation of action-dependent scene changes. We bring these predictions into action generation and selection through a shared trajectory model. Dynin-Robotics implements this formulation on Dynin-Omni, an omnimodal masked-diffusion backbone, representing language, visual observations, goals, and actions as discrete tokens. By varying conditioning and target spans, the same model learns action prediction, action-conditioned next-observation prediction, terminal goal-state prediction, and trajectory-to-instruction reconstruction. These interfaces support test-time scaling through goal prediction, action-candidate evaluation, and joint refinement of action and future-state predictions. We continually pretrain the model on approximately 1.33 million trajectories from 48 Open X-Embodiment datasets and adapt it separately to downstream domains. On two VLABench tasks, robot pretraining improves adaptation within a fixed Stage-2 step budget, and the full objective mixture improves shifted-instruction success over Policy-only post-training under the same coupled decoder. Combining goal guidance with joint action-next-state denoising further improves shifted-instruction success over action-only decoding; the benefit depends on how the predictions are composed. Dynin-Robotics achieves competitive performance on LIBERO and zero-shot LIBERO-Plus, together with a 78.4% average success rate across four manipulation conditions on a Franka Research 3 robot. An optimized block-parallel implementation accelerates model-side action decoding by up to 29.2x relative to the base implementation under the reported profiling setup. These results support shared trajectory modeling as a common interface for learning complementary robot objectives and composing their predictions during control.
Create a lesson
Related papers
ASTRIL-MPC: Autonomous Traversal Framework of Articulated Tracked Robots with Language-Guided Neural-Kinematic MPC
Zhenfeng Gan, Yanbo Chen, Lirong Che et al.
Global Path Planner with Multi-Model Switching
Pietro Gori, Francesco Iotti, Eduard Zelenay et al.
Comfort by Construction: Adaptive, Comfort-Bounded Action Spaces for Learned Driving Policies
Anna Rothenhäusler, Daniel Jost, Raghu Rajan et al.
Tuning ROS 2 for Energy-Efficient Navigation: Empirical Insights from Costmap 2D Configurations
Michel Albonico, Andreas Wortmann, Ivano Malavolta
Distributed Stochastic Optimal Control for Pattern-Oriented Swarms
Qingrui Zhang, Chenghao Yu, Feng Xue et al.
A Robot Among People:From Social Imitation to the Social Becoming of Human Groups
Victor Tuan Vu Pham, Judith Dörrenbächer, Thomas H. Weisswange et al.