Skip to content

Feedback-Aware Tuning of Recursive Q-Learning

Masahiro Kojima

stat.MEarXiv:2609.12716

Abstract

Model choice in backward Q-learning is recursive because a later-stage choice changes the response supplied to an earlier regression and can alter its model-comparison statistic. Separate stagewise criteria do not directly assess the target-stage prediction risk of a completed Q-learning fit. We address this mismatch by treating the entire backward-fitting rule as the unit of comparison and propose feedback-aware soft tuning for sequential multiple assignment randomised trials. Each backward fit uses its own generated responses and is assessed at a common prediction target. The risk criterion retains downstream effects on upstream comparisons, while a separate correction accounts for estimating the final exponential weights from the same observations. For a fixed finite library of smooth recursive maps, we establish an exact risk identity under a Gaussian shift model and an oracle inequality with an explicit adaptation remainder. Under coordinate-representation and moment conditions, these guarantees transfer to prediction risk at any prespecified stage, with the number of stages fixed as sample size increases. A two-stage construction supplies an explicit observable implementation. Numerical studies examine risk estimation and finite-sample performance, and a simulated attention-deficit/hyperactivity-disorder trial illustrates the relation between comparison feedback and treatment recommendations.

Create a lesson