Projection-Free Bandit Online Optimization for Multi-Agent Systems with Dynamic Regret
Xia Jiang, Lu Liu, Gang Feng
Abstract
This paper investigates distributed online optimization for multi-agent dynamical systems with constrained inputs and time-varying cost functions. While online convex optimization offers a principal framework for sequential decision-making, existing online learning and optimization algorithms typically require accurate system models, limiting their applicability in practical settings. To overcome this challenge, we propose a distributed bandit online feedback optimization algorithm that relies solely on real-time input-output data. The algorithm employs a smoothing zeroth-order one-point estimator to construct local gradient approximations directly from cost evaluations. Additionally, to enforce input constraints effectively, we integrate a projection-free conditional gradient update, making the algorithm well-suited for online and large-scale settings. Furthermore, we establish a sublinear dynamic regret bound that depends on a temporal variation measure of system non-stationarity. Finally, numerical simulations demonstrate the effectiveness of the proposed algorithm.
Create a lesson
Related papers
A Smallest-Need-First Job Scheduling Framework with Adaptive Optimization of Idle Node Counts for Energy-Efficient HPC Systems
Reza Pulungan, Raka Satya Prasasta, Santana Yuda Pradata et al.
Bridging Agent Semantics with Spot Capacity: An Elastic and Recoverable Service Model
Minchen Yu
CLASP: Chained-Request-Aware Scaling and Operator Placement for Serverless Stream Processing
Tianyu Qi, Maria A. Rodriguez, Rajkumar Buyya
Performance Evaluation of RED-ONION: A High-Speed Disk-to-Disk Transfer System
Keichi Takahashi, Hiroaki Kataoka, Takeo Hosomi et al.
Memory-efficient GPU pipelines for real-time non-line-of-sight reconstruction
Alfonso López-Ruiz, Diego Royo
Great Expectations: Benchmarking the Real-World Performance of RVV 1.0 in HPC
Stepan Nassyr, Prateek Chawla, Daniel Seibel et al.