Skip to content

SDE Guided Monte Carlo Reinforcement Learning: A Stochastic Maximum Principle Approach for Robust Decision Making in Noisy Environments

Juncai Wang

math.OCarXiv:2607.22541

Abstract

This work investigates whether the qualitative optimality conditions of the SMP can serve as guiding heuristics for designing robust tabular RL algorithms. We propose the SDE-MC-AC framework, which establishes a set of SMP-inspired correspondences: the potential field gradient is mapped to potential-based reward shaping, the diffusion coefficient to a softmax temperature, and the value function gradient magnitude to the episode-averaged absolute temporal-difference (TD) error. Crucially, we derive an SMP-motivated adaptive temperature schedule that scales exploration stochastically in response to local value uncertainty. The resulting Monte Carlo Actor-Critic agent integrates potential-shaped rewards, entropy regularization, and adaptive temperature control within episodic updates. We conduct a set of five principled experiments in a minimal, fixed maze testbed under three canonical noise regimes (perceptual, dynamic goal, action) to test five falsifiable hypotheses derived from the SMP. Our results provide evidence for (i) an inverted-U optimal stochasticity curve, (ii) the robust failure prevention of the adaptive schedule under non-stationary noise, (iii) convergence acceleration by potential field guidance, (iv) cross-noise generalization, and (v) the indispensability of the combined SDE components through systematic ablation. Trajectory visualizations reveal a characteristic ``macroscopically deterministic, microscopically stochastic'' navigation pattern consistent with the SDE model. The study demonstrates that even a heuristic transposition of the SMP into a discrete MDP can yield significant empirical gains, thereby opening a bridge between rigorous stochastic optimal control and practical robust reinforcement learning.

Create a lesson

Paper details

Categories: math.OC, cs.NA, math.NA, math.PR