Physics-informed Reinforcement Learning for Stochastic Reach-Avoid Analysis
Hikaru Hoshino, Yorie Nakahira
Abstract
Stochastic reach-avoid analysis of controlled dynamical systems is an important tool for safety-critical control under uncertainty, in which the reach-avoid probability is characterized by a Hamilton-Jacobi partial differential equation (PDE). However, solving this PDE using conventional numerical methods becomes computationally intractable as the system dimension increases. Physics-informed neural networks (PINNs) may converge to inaccurate local minima when trained primarily through PDE-residual minimization. Reinforcement learning (RL) offers a scalable alternative, but its learned value functions may be inaccurate or inconsistent with the governing PDE. This paper proposes a physics-informed RL (PIRL) framework that combines the complementary strengths of PINNs and RL for stochastic reach-avoid analysis. We develop a scheduled PIRL algorithm in which temporal-difference actor-critic learning first guides the critic toward a meaningful approximation of the reach-avoid value function. PDE-residual and boundary-condition losses are then introduced progressively to enforce consistency with the governing PDE and its boundary conditions. The proposed method mitigates the failure modes of conventional PINN techniques while achieving accuracy comparable to that of successfully trained PINNs. The effectiveness of the proposed framework is demonstrated through two case studies.
Create a lesson
Related papers
Leader-Follower Formation Control with Prescribed Convergence Rates under Bearing Persistence of Excitation
Tarek Bouazza, Zhiqi Tang, Soulaimane Berkane et al.
On asymptotic stability of the time-varying Kalman filter for unstabilizable linear systems: an optimization perspective
James B. Rawlings, Titus Quah, Matthias A. Müller
Designing Grid-Aware Dynamic Specifications for Large Data Center Loads
Ashutossh Gupta, Vassilis Kekatos
Time-Optimal Operation of a Load-Hoisting Gantry Crane
Eric Mountain, Tarunraj Singh
Learning to Solve Two-Stage Stochastic Unit Commitment Problems with Quality Guarantees
Andrea Fusco, Andrea Lodi, Lavanya Marla
Towards Interaction Regulation from Human Feedback via Free Energy Minimization
Maria Paula Diaz Monfort, Cinzia Tomaselli, Michael Richardson et al.