The Complexity of Approximating the Value in Revealing POMDPs with Long-Run Average Objectives
Ali Asadi, Krishnendu Chatterjee, David Lurie
Abstract
We study partially observable Markov decision processes (POMDPs) with long-run average objectives, defined as the limit inferior of the expected average rewards. In general, the long-run average value of a POMDP is neither computable nor approximable. We therefore consider the subclass of revealing POMDPs, in which the current state is revealed to the controller with positive probability at each stage. First, we illustrate the practical relevance of this class through an application in control and optimization. Second, we establish that approximating the long-run average value of revealing POMDPs with long-run average objectives is EXPTIME-complete, thereby providing a tight computational complexity.
Create a lesson
Related papers
Gaussian-Restricted Barycenters for KL-Unbalanced Optimal Transport: Variational Theory and Fixed-Point Convergence
Jiaping Yang, Yunxin Zhang
Improved Gradient Descent Lower Bounds Beyond Nesterov
Yuhan Ye, Kaizhao Liu
Fundamental Limits of Adaptive Stabilization with an Unknown Growth Exponent
Zhaobo Liu
An Adaptive Projected-Gradient Algorithm for Sample-Average Approximations of Stochastic Multi-Objective Optimization
Yiyang Li, Lei Wang, Xiaojun Chen
Insensitizing Control Problems for Coupled Stochastic Parabolic Systems with State and Gradient Observations
Said Boulite, Abdellatif Elgrou, Abdelaziz Rhandi
Projected Subgradient Methods for a Class of Nonsmooth and Nonconvex Optimization Problems
Christian Kanzow, Jannis Krüger, Leo Lehmann