A Unified and Constrained View of Regularization-Based Robust Reinforcement Learning
Amine Andam, Jamal Bentahar, Mustapha Hedabou
Abstract
Regularization-based methods have become a standard approach for training Deep Reinforcement Learning policies against adversarial input perturbations. In this paper, we unify these methods by deriving new upper bounds on the performance gap between the nominal and worst-case policies. Each upper bound is expressed as an existing regularization objective plus a KL-divergence penalty between the nominal and worst-case policies, which further explains why adding a KL penalty improves robustness in practice. Building on these bounds, we formulate robust training as a constrained optimization problem, showing that existing methods correspond to the special case of a fixed Lagrange multiplier. We instead update the multiplier jointly with the policy to automatically tune the regularization weight. Finally, we conduct extensive adversarial evaluations across several continuous control tasks to validate our theoretical analysis.
Create a lesson
Related papers
MAxBench: A Multinomial Concept Recovery Benchmark
Divya Appapogu, Freya Behrens, Yonatan Belinkov et al.
CanvasAnneal: Curriculum Reinforcement Learning for Diffusion Language Models
Blake Olson, Yuhang Song, Emmett McQuinn et al.
Benign Loss Landscapes Can Coexist with Worst-Case Hardness
Zach Furman, Stephan Wäldchen, Yangda Bei et al.
MCRL2: Multi-resource Cross-attention-based Representation Learning-augmented Reinforcement Learning for Cloud Microservice Scheduling
Tiangang Li, Shi Ying, Xiangbo Tian et al.
Robust Policy Optimization via Adversarial Importance Sampling
Amine Andam, Jamal Bentahar, Mustapha Hedabou
DynSHAP: Towards Explainable Dynamic Survival Analysis
Nastasya Anokhina, Jonas Jürß, Pietro Liò