SDE Guided Monte Carlo Reinforcement Learning: A Stochastic Maximum Principle Approach for Robust Decision Making in Noisy Environments
Juncai Wang
Abstract
This work investigates whether the qualitative optimality conditions of the SMP can serve as guiding heuristics for designing robust tabular RL algorithms. We propose the SDE-MC-AC framework, which establishes a set of SMP-inspired correspondences: the potential field gradient is mapped to potential-based reward shaping, the diffusion coefficient to a softmax temperature, and the value function gradient magnitude to the episode-averaged absolute temporal-difference (TD) error. Crucially, we derive an SMP-motivated adaptive temperature schedule that scales exploration stochastically in response to local value uncertainty. The resulting Monte Carlo Actor-Critic agent integrates potential-shaped rewards, entropy regularization, and adaptive temperature control within episodic updates. We conduct a set of five principled experiments in a minimal, fixed maze testbed under three canonical noise regimes (perceptual, dynamic goal, action) to test five falsifiable hypotheses derived from the SMP. Our results provide evidence for (i) an inverted-U optimal stochasticity curve, (ii) the robust failure prevention of the adaptive schedule under non-stationary noise, (iii) convergence acceleration by potential field guidance, (iv) cross-noise generalization, and (v) the indispensability of the combined SDE components through systematic ablation. Trajectory visualizations reveal a characteristic ``macroscopically deterministic, microscopically stochastic'' navigation pattern consistent with the SDE model. The study demonstrates that even a heuristic transposition of the SMP into a discrete MDP can yield significant empirical gains, thereby opening a bridge between rigorous stochastic optimal control and practical robust reinforcement learning.
Create a lesson
Paper details
Categories: math.OC, cs.NA, math.NA, math.PR