Skip to content

Zeroth-Order Nonsmooth Nonconvex Optimization with Convex Liftings and Its Application to State-Feedback H∞ Policy Optimization

Xuhao Wang, Yujie Tang

math.OCarXiv:2608.23178

Abstract

Direct policy optimization is widely used in reinforcement learning and control, but generally leads to nonconvex optimization problems. For state-feedback H∞ control, the policy objective is also nonsmooth, despite possessing a benign landscape whose hidden convexity can be revealed by the recently developed extended convex lifting framework. Motivated by recent advances in hidden convex optimization, we study zeroth-order optimization of nonsmooth, nonconvex problems admitting a convex lifting. We propose a zeroth-order proximal point algorithm: An inexact proximal-point outer loop constructs strongly convex subproblems, while an inner loop approximately solves each subproblem using only function evaluations. With probability at least 1-δ, our proposed algorithm returns an ε-optimal solution using O(dε-3) function evaluations, while all iterates remain feasible without explicit projection. Finally, we verify that the assumptions underlying our analysis hold for discrete-time state-feedback H∞ policy optimization, yielding an oracle complexity of O(nu nxε-3) for attaining a prescribed objective value gap, where nu× nx is the dimension of the feedback gain to be optimized over.

Create a lesson