Test-time Reinforcement Learning in Imperfect Information Games
Ondrej Kubicek, Viliam Lisy, Tuomas Sandholm
Abstract
Test-time reasoning has significantly improved performance in domains ranging from games to language models. However, test-time policy changes with formal guarantees on the performance of the resulting strategy remain a challenge in two-player zero-sum imperfect-information games. Existing solutions are limited to tabular methods or single gradient step updates. In this work, we investigate policy-gradient algorithms as a method for scalable test-time reasoning. We extend the concept of gadget game, tabular technique for test-time search, to the reinforcement learning setting. Unlike prior approaches, we represent the gadget game implicitly by modified sampling and neural policy rather then explicitly by constructing it, thereby removing constraints on subgame size. Furthermore, we formally prove that, unlike prior tabular algorithms, regularized policy-gradient algorithms limit possible strategy degradation caused by test-time reasoning, even without the gadget games. Our evaluation across small- and large-scale games confirms that additional test-time training often substantially improves performance relative to the blueprint strategy.
Create a lesson
Related papers
Constrained Fair Allocations via Partition Matroid Reductions
Benjamin Cookson, Nisarg Shah
Bidding Games with Rewards: Taming Infinite Configuration Space
Matan Pinkas
Reaching Fairness by Reallocating Goods
Robert Bredereck, Eva Deltl, Tanmay Inamdar et al.
LangBP: Language-Guided Reasoning and Acting for Joint Bidding and Pricing
Jiaqi Ding, Chuan Yang, Linghui Meng et al.
Mechanism Design for Facility Location Games Under a Prelocated Facility
Genjie Qin, Qizhi Fang, Wenjing Liu
The Exact MMS Guarantees of EFX and PMMS
Qinghua Qin