Skip to content

The Complexity of Approximating the Value in Revealing POMDPs with Long-Run Average Objectives

Ali Asadi, Krishnendu Chatterjee, David Lurie

math.OCarXiv:2609.01099

Abstract

We study partially observable Markov decision processes (POMDPs) with long-run average objectives, defined as the limit inferior of the expected average rewards. In general, the long-run average value of a POMDP is neither computable nor approximable. We therefore consider the subclass of revealing POMDPs, in which the current state is revealed to the controller with positive probability at each stage. First, we illustrate the practical relevance of this class through an application in control and optimization. Second, we establish that approximating the long-run average value of revealing POMDPs with long-run average objectives is EXPTIME-complete, thereby providing a tight computational complexity.

Create a lesson