Model-based Bootstrap for Offline Policy Evaluation in Tabular Reinforcement Learning
Weiwei Wang, Yuqiang Li, Xianyi Wu, Bingyi Jing
Abstract
Offline policy evaluation (OPE) is crucial in high-stakes reinforcement learning applications, where new policies must be assessed reliably before deployment. In such settings, point estimates alone are insufficient; principled uncertainty quantification, such as confidence intervals and variance estimates, is essential for safe and risk-aware decision-making. A comprehensive way to unify these tasks is to estimate the sampling distribution of the evaluation error. Existing approaches, however, often suffer from limited robustness, scalability, or finite-sample validity. In this paper, we propose a model-based bootstrap framework for uncertainty quantification of OPE in finite-horizon, time-inhomogeneous Markov decision processes (MDPs). Unlike classical bootstrap methods that rely on resampling complete episodes, the proposed method regenerates trajectories from an estimated MDP and can therefore accommodate a much broader range of offline data formats, including complete trajectories, transition-level observations, and trajectory fragments. This flexibility further improves finite-sample statistical efficiency. We establish bootstrap distributional consistency, asymptotically valid confidence intervals, and consistent variance estimation for the target policy value. Extensive simulations show that the proposed method accurately captures the sampling distribution of the OPE estimator, yielding tighter confidence intervals and more accurate variance estimates in most settings.
Create a lesson
Related papers
Prediction-Powered Smoothing and Validation for Disaggregated AI Evaluation
Sho Kawano, Zehang Richard Li, Paul A. Parker
TAP Accuracy Below the Fluctuation Scale and Universal Posterior Geometry in Spherical Linear Models
Jingbo Liu, Zhiyuan Yu
Online Supervised Dimension Reduction with Random Features: Diagnostics and Computational Trade-offs
Zhenlin Yao, Wei Xiong
Error bounds in Sobolev norms for approximations with norm constrained ReLU neural networks
Xianjun Li, Yunfei Yang
Next-token functional estimation
Milind Nakul, Vidya Muthukumar, Ashwin Pananjady
Null importance: Disentangling relevance for interpretable machine learning
Garvesh Raskutti, Kris Sankaran, Jiaxin Ye