Combinations and Mixtures of Optimal Policies in Unichain Markov Decision Processes are Optimal
Ronald Ortner
Abstract
We show that combinations of optimal (stationary) policies in unichain Markov decision processes are optimal. That is, let M be a unichain Markov decision process with state space S, action space A and policies πj*: S -> A (1≤ j≤ n) with optimal average infinite horizon reward. Then any combination πof these policies, where for each state i in S there is a j such that π(i)=πj*(i), is optimal as well. Furthermore, we prove that any mixture of optimal policies, where at each visit in a state i an arbitrary action πj*(i) of an optimal policy is chosen, yields optimal average reward, too.
Create a lesson
Related papers
Combinatorics of hyperplane arrangements and Witten zeta function at the origin
Kam Cheong Au, Kazuhiro Onodera
Proof of the Pach-Tardos conjecture
Lior Gishboliner, Xiangyu Li
Longest cycles intersect linearly in highly connected graphs
Jie Ma, Bo Ning, Ziyuan Zhao
Cutting a convex body into fat parts and approximating Euclidean distance by graph distances
János Pach, Gábor Tardos
A counterexample to the quantum Hedetniemi conjecture
Julius A. Zeiss
Schrijver-Delsarte rigidity in association schemes and undecidability of quantum graph homomorphism
Lorenzo Ciardo, Iris Hebbeker, Gideo Joubert et al.