Reachability-Informed Reinforcement Learning for Multi-Impulse Interplanetary Transfers
Yashdeep Chaudhary, Roberto Armellin, Harry Holt
Abstract
Reinforcement learning offers the prospect of a reusable sequential decision-making mechanism for spacecraft trajectory design, motivating policy interfaces that connect learned decisions to the underlying maneuver geometry. This paper develops Reachability Analysis-Informed Reinforcement Learning (RARL) for deterministic multi-impulse interplanetary transfers, placing intermediate waypoint selection at the center of the learned decision process. Local first-order reachability maps bounded velocity perturbations into an ellipsoidal set of next-node positions, within which the policy selects its waypoint. Lambert reconstruction then determines the corresponding maneuver to reach this selected waypoint along a dynamically consistent ballistic arc, coupling learned transfer-geometry selection with classical astrodynamics. A terminal two-impulse reconstruction completes the rendezvous, supported by a linear maneuver-demand assessment used for reward shaping. Numerical studies characterize this interface on a two-body Earth-Mars benchmark. Across three independent training runs, RARL achieves a mean maneuver cost of 10.23 km/s, 1.72% above a validated local sequential convex programming reference. Training over dispersed initial states extends policy reuse across a departure family with fixed target state and transfer duration. Each of the three independently trained multi-state policies completes all 10,000 held-out Monte Carlo departures without impulse-cap violations, compared with a mean feasibility rate of 6.49% for single-state policies. This broader sampled feasibility is accompanied by a 0.61% increase in mean nominal maneuver cost, without further training across departures. These results demonstrate that a reachability-informed decision interface supports benchmark-quality trajectory construction and policy reuse across dispersed departure conditions.
Create a lesson
Related papers
Randomized Matvec Lower Bounds for Simplex-Based Matrix Games
Wendao Wu, Cong Fang
Optimal Stochastic Bilevel Optimization with First-Order Oracles
Linxuan Pan, Junchi Yang
Recognizing Signomial Convexity is Hard
Rui Zheng, Iosif Sakos, Antonios Varvitsiotis
Simplifying the computation of weak- second subderivatives of convex functionals via Γ-convergence
Gerd Wachsmuth
Routing in Line Networks with Handling Times
Gabriel Deza, Michal Tzur, Tal Raviv
Lower Bounds for Stochastic First-Order Algorithms with Variance Reduction in Nonconvex--Concave Minimax Optimization
Jiayi Song, Zi Xu