A penalized bandit algorithm
Damien Lamberton, Gilles Pagès
Abstract
We study a two armed-bandit algorithm with penalty. We show the convergence of the algorithm and establish the rate of convergence. For some choices of the parameters, we obtain a central limit theorem in which the limit distribution is characterized as the unique stationary distribution of a discontinuous Markov process.
Create a lesson
Related papers
Boolean Small-Ball Inequalities for Discrepancy Theory
Emrullah Akbas, Suvrit Sra
Markovian renormalisation for percolation in high-dimension: Semi-decidability of mean field behavior
Arthur Blanc-Renaudie
Point process convergence of large inradii of Poisson-Laguerre tessellations
Matthias Schulte, Martina Švarc Petráková
Interpolation of Gaussian Free Fields via Random Matrices
Gabriel Raposo
Almost-Uniform Bayesian Convergence to the Truth Is Not Characterized by Countable Additivity on Conditional Hitting Times
M. Ali Khan, Arthur Paul Pedersen, Maxwell B. Stinchcombe
The skeleton-blocks decomposition of Bienaymé trees, and applications to their local convergence
Marc Bernard, Robin Stephenson