PROMO: Preference-conditioned Multi-Objective Reinforcement Learning for Quadrupedal Robots
Amr Mousa, Rifny Rachman, Neil Karavis, Michele Caprio, Richard Allmendinger
Abstract
Quadrupedal locomotion requires balancing conflicting objectives such as command tracking, stability, and energy efficiency, yet conventional reinforcement learning (RL) hardcodes these priorities into a fixed scalar reward at training time. We present PROMO (Preference-Conditioned Multi-Objective Reinforcement Learning), a semantic multi-objective approach that makes this trade-off an explicit runtime input to a single locomotion policy. PROMO conditions the policy on deployment facing preferences while keeping embodiment-specific locomotion priors fixed, thereby separating operator intent from reward shaping terms required for viable gait generation. Compared with fixed-objective controllers, multi-objective baselines, and independently trained specialists, PROMO achieves objective specialization and robustness from a single deployable policy. Across 100 sampled preferences in simulation, 67 behaviors are non-dominated under exact Pareto dominance, with a mean preference-objective correlation of 0.843, demonstrating broad Pareto coverage and predictable preference response. The same policy transfers zero-shot to a Unitree Go2, where preference changes alone reduce specific energy by up to 30.4%, position error by 38.7%, and peak body-attitude deviation by 59.0% relative to the balanced preference. These results establish preference-conditioned multi-objective RL as a practical runtime interface for adaptive legged locomotion, extending its role beyond offline Pareto-set construction. Open-source code and videos are available at https://amrmousa.com/promo/.
Create a lesson
Related papers
Reconstruct, Practice, Go Real: Guided Self-Improvement for Embodied Agents
Yen-Jen Wang, Haozhe Jiang, Shuying Deng et al.
InterEvolve: Test-Time Evolution of Reward Programs for Humanoid Loco-Manipulation
Zhuo Lin, Sirui Xu, Liuyu Bian et al.
Watch, Infer, Coordinate: Inferring Robot Partner Constraints for Zero-Shot Coordination
Suyu Ye, Zheyuan Zhang, Vaishnav Tadiparthi et al.
DuoMind: Enabling Distributed Multi-Robot Coordination with Semantic Communication
Hanchu Zhou, Dechen Gao, Hang Wang et al.
SkeleWAM: Skeleton World-Action Modeling for Efficient Robotic Manipulation
Juyi Sheng, Hua Wang, Mengyuan Liu
GlassGuard: Verified Glass Plane Mapping for Robot Navigation
Hanwen Guo, Zhengzhi Lin, Yusen Xie et al.