Online Learning in Stackelberg Security Games with Adaptive Attacker Sequences and Time-Varying Attack Intensities
Guanda Chen, Shiheng Zhang, Yue Wang, Yiding Ji
Abstract
This work studies no-regret online learning in Repeated Stackelberg Security Games with time-varying attack intensities. We formulate an extended security game in which an attacker may select multiple targets and derive an exact mixed-integer linear programming oracle under a optimistic tie-breaking rule. Under full-information feedback, the oracle is integrated with Follow-the-Perturbed-Leader and yields expected O(T) regret against non-anticipating sequences with time-varying follower numbers, attack intensities, and attacker types. Under bandit feedback, we consider multiple followers sharing a fixed attacker type and use a barycentric-spanner construction to reconstruct utility estimates from aggregate attack observations, obtaining expected O(T2/3) regret. Extensive simulations demonstrate the robustness and effectiveness of our approach under full and partial information feedback.
Create a lesson
Related papers
On the Role of Tie-Breaking Rules in the Convergence of Fictitious Play for Symmetric First-Price Auctions
Benjamin Heymann
Epsilon-Nash Equilibria in History-Dependent SA-MDPs
Brandon Gary Kaplowitz, Dominik Bohnet Zurcher, Akash Agrawal et al.
Core stability recognition for minimum-cost spanning tree games: Parameterized perspective
Michal Dvořák, Ioannis Kakatelis, Dušan Knop
Second-Best Gains from Trade in Matching Markets
Xiaohui Bei, Bo Li, Wenhao Wu et al.
Equilibria of Round-Robin: Computational Hardness and Fairness for Few Subadditive Agents
Paul W. Goldberg, Alexandros Hollender, Giannis Tyrovolas
Estimate then Predict: Convex Formulation for Travel Demand Forecasting
Youngseo Kim, Gioele Zardini, Samitha Samaranayake et al.