Convergence and Loss Bounds for Bayesian Sequence Prediction
Marcus Hutter
Abstract
The probability of observing xt at time t, given past observations x1...xt-1 can be computed with Bayes' rule if the true generating distribution μ of the sequences x1x2x3... is known. If μ is unknown, but known to belong to a class M one can base ones prediction on the Bayes mix ξ defined as a weighted sum of distributions ν∈ M. Various convergence results of the mixture posterior ξt to the true posterior μt are presented. In particular a new (elementary) derivation of the convergence ξt/μt 1 is provided, which additionally gives the rate of convergence. A general sequence predictor is allowed to choose an action yt based on x1...xt-1 and receives loss xt yt if xt is the next symbol of the sequence. No assumptions are made on the structure of (apart from being bounded) and M. The Bayes-optimal prediction scheme Λξ based on mixture ξ and the Bayes-optimal informed prediction scheme Λμ are defined and the total loss Lξ of Λξ is bounded in terms of the total loss Lμ of Λμ. It is shown that Lξ is bounded for bounded Lμ and Lξ/Lμ 1 for Lμ ∞. Convergence of the instantaneous losses are also proven.
Create a lesson
Related papers
Calibrated RF-Fingerprinting Under Interference With Heterogeneous Transmission Protocols
Tariq Abdul-Quddoos, Xiangfang Li, Lijun Qian
Video DeltaNet: A Video-Native Hybrid Attention for Livestream Video Generation
Haocheng Xi, Yiming Xie, Hexu Zhao et al.
Don't Mask the Environment: Observation Supervision Changes How Agents Explore Under RL
Juzheng Zhang, Disha Makhija, Manoj Ghuhan Arivazhagan et al.
RISC-V and machine learning: a survey
Shriman Keshri, Apparna Singh, Chinmaya Kumar Palo et al.
Epidemiological Causal Graph Identification: Challenges, Identifiability and Algorithms
Sambit Mishra, Yingying Wang, Christine K. Johnson et al.
Multi-center Medical Data Mining with FL-Net - A One-stop Shop for Federated Learning
Simon Süwer, Julian Klemm, Elisa Acitelli et al.