Stable Multi-Step Rollouts via Uncertainty-Guided Hybrid Dynamics
Andrei Maalberg, Axel Neumann, Jens Knobloch
Abstract
Multi-step rollouts are essential for model-based reinforcement learning (RL) and predictive control, yet learned dynamics models often become unstable when recursively applied, leading to divergence and unreliable policy updates. This paper proposes a model-agnostic hybrid dynamics framework that blends a provably contracting nominal model with a flexible excursion model through an uncertainty-guided switching law. The switching signal is derived from calibrated epistemic uncertainty and activates only when the system leaves the nominal region, ensuring that each model operates within its reliability regime. Under clearly stated smoothness and boundedness assumptions, we show that the resulting hybrid predictor yields globally bounded recursive multi-step rollouts: trajectories remain Lyapunov-stable in the nominal region and exhibit at most affine growth during excursions. To illustrate the theory in practice, we instantiate the hybrid dynamics framework within a model-based RL scheme that uses real one-step transitions for value learning and hybrid rollouts for policy improvement. Experiments on a nonlinear Duffing oscillator demonstrate stable long-horizon prediction and improved cost-effort trade-offs relative to a stabilizing baseline.
Create a lesson
Related papers
Leader-Follower Formation Control with Prescribed Convergence Rates under Bearing Persistence of Excitation
Tarek Bouazza, Zhiqi Tang, Soulaimane Berkane et al.
On asymptotic stability of the time-varying Kalman filter for unstabilizable linear systems: an optimization perspective
James B. Rawlings, Titus Quah, Matthias A. Müller
Designing Grid-Aware Dynamic Specifications for Large Data Center Loads
Ashutossh Gupta, Vassilis Kekatos
Time-Optimal Operation of a Load-Hoisting Gantry Crane
Eric Mountain, Tarunraj Singh
Learning to Solve Two-Stage Stochastic Unit Commitment Problems with Quality Guarantees
Andrea Fusco, Andrea Lodi, Lavanya Marla
Towards Interaction Regulation from Human Feedback via Free Energy Minimization
Maria Paula Diaz Monfort, Cinzia Tomaselli, Michael Richardson et al.