A Finite Sample Analysis for Quantile Temporal Difference Learning in Distributional Reinforcement Learning
Zijie Cheng, Xiang Li, Yang Peng, Zhihua Zhang
Abstract
We establish a global finite-sample guarantee for synchronous quantile temporal-difference learning (QTD) in tabular distributional reinforcement learning. The proof separates two stability mechanisms. A global comparison argument, based on the order monotonicity of reward cumulative distribution functions and the W∞ contraction of the distributional Bellman operator, brings an arbitrarily initialized iterate into a local neighborhood. Inside that neighborhood, we linearize the QTD mean field. Its Jacobian is a nonsingular M-matrix, and the associated positive semigroup permits a variance-sensitive martingale analysis. For stepsizes αt=c(t+1)-a with a∈(1/2,1), the leading last-iterate fluctuation is of order O(T-a/2/1-γ) and has no polynomial dependence on the number of quantiles. The deterministic transient and the required burn-in can still depend on the smallest Bellman-target density, which is of order m-1 in the worst case. The result therefore distinguishes sharply between the local stochastic fluctuation and the global sample complexity.
Create a lesson
Related papers
A General Kernel Framework for Non-CND Distance Measures Using |D|-Dimensional Sparse Landmark Embeddings
Marcus M. Noack, Maher B. Alghalayini, Mark D. Risser
Fast Learning Rates for Physics-Informed Kernel Methods
Luc Brogat-Motte, Joachim Bona-Pellissier, Giacomo Meanti et al.
Rank and computation of the pathlifting Jacobian of a DAG ReLU network
Manon Verbockhaven
Preservation of Log-Concavity and Convergence of Wasserstein-Fisher-Rao Gradient Flows
Francesca Romana Crucinio, Sahani Pathiraja
Generalized DCCQ: From Binary Quotients to Multinomial Simplex Geometry and Critical-Strip Coordinates
Y. Kenan Yılmaz
Bracketing Uncertainty in Clustering Under the Manifold Hypothesis
Savik Kinger, Luciano Dyballa, Steven W. Zucker