Cross-Entropic Learning of a Machine for the Decision in a Partially Observable Universe
Frederic Dambreville
Abstract
Revision of the paper previously entitled "Learning a Machine for the Decision in a Partially Observable Markov Universe" In this paper, we are interested in optimal decisions in a partially observable universe. Our approach is to directly approximate an optimal strategic tree depending on the observation. This approximation is made by means of a parameterized probabilistic law. A particular family of hidden Markov models, with input and output, is considered as a model of policy. A method for optimizing the parameters of these HMMs is proposed and applied. This optimization is based on the cross-entropic principle for rare events simulation developed by Rubinstein.
Create a lesson
Related papers
UGM: A Unified Framework and New Perspectives for Accelerated Gradient Methods in Smooth and Strongly Convex Optimization
Danqing Zhou, Shiqian Ma, Junfeng Yang
When MILP Beats QP: Piecewise-Linear Reformulations of Sequentially Coupled Bilinear Programs
Quentin Ploussard, Maris Usis, Oluwabunmi Iwakin et al.
Marine Autonomous Vehicle Fleet Scheduling to Maximise Scientific Impact
Mehdi El Krari, Jonathan Smith, Maria Fox
Asymptotic consensus and flocking under decaying persistent excitation on rooted digraphs
Chiara Cicolani, Elisa Continelli, Cristina Pignotti
Co-Optimized Generation, Transmission, and Storage Expansion: System Value and Optimal Duration of Pumped-Storage Hydropower
Rafael Benchimol Klausner, Rafael Kelman
Randomized Quasi-Gauss--Newton Methods for Solving General Nonlinear Equations
Chengchang Liu, Luo Luo