From Relaxed Indexability to Exact Indexability: A t-Step Approach for Partially Observable Restless Bandits
Keqin Liu, Qizhen Jia
Abstract
Whittle index policies offer a scalable method for restless multi-armed bandits, but under partial observability even determining the indifference subsidy at a single belief requires solving an infinite-horizon belief-state problem with no closed-form value function. Liu [10] addresses this difficulty by linearizing the unknown decision boundary, leading to a linear system and a closed-form approximate Whittle index. However, the resulting threshold uses only a one-step active--passive comparison and does not account for longer-horizon continuation values. We extend this framework to a t-step lookahead threshold policy. For each subsidy m, the threshold is defined by the active-minus-passive advantage under t-step finite-horizon value iteration. At t=1, the threshold is m-independent and recovers the linear threshold of Liu [10]; for t>1, it becomes subsidy-dependent through the induced first-crossing structure and tracks the exact decision boundary more closely. The proposed algorithmic framework does not require indexability as an input and includes an indexability verification. Under the original Whittle indexability, we prove that the t-step approximate Whittle index converges geometrically to the exact Whittle index, \[ | Wt(ω)-W(ω)|=O(βt). \] Numerically, all 2,715 tested three-state instances are verified with computable priority index functions according to the proposed criterion. The P95 index error decreases from 2.18×10-2 at t=1 to 8.93×10-4 at t=8. In an exact-comparable instance with β=0.9999, t=2 already recovers the exact Whittle-index ordering. Moderate-depth threshold policies also outperform the one-step baseline and remain close to the optimal dynamic-programming benchmark, while runtime grows mildly with t.
Create a lesson
Related papers
How Model Growth, Recursion, and Boundary Operators Influence Scaling Exponents
Zixi Chen, Akshay Vegesna, Samip Dahal et al.
Evidence-Grounded Agentic Formulation Development in an Autonomous Laboratory
Michael M. Craig, Riley J. Hickman, Yingshan Ma et al.
Probabilistic Linear Explanations
Frederic Koriche, Jean-Marie Lagniez, Chi Tran
Double descent is the principle of least action
Congzhou M Sha
RLLBC-Lib: An Educational Code Library for Reinforcement Learning and Learning-Based Control
Bernd Frauenknecht, Emma Cramer, Artur Eisele et al.
Higher-order pruning of experts in mixture-of-experts language models
Alex M. Tseng, Prannay Kaul, Luca Zancato et al.