The Sample Complexity of Distributionally Robust PAC Learning under Cressie--Read Divergences
Elad Aigner-Horev, Daniel Rosenberg, Roi Weiss
Abstract
We study distributionally robust PAC learning for the 0--1-loss, where adversarial perturbations of the data distribution are constrained by a Cressie--Read divergence of order k>1 and radius ρ≥ 0. For hypothesis classes with VC dimension d, we establish realizable and agnostic sample-complexity bounds tight up to constant and logarithmic factors, respectively; ordinary empirical risk minimization attains both rates up to logarithmic factors. For target accuracy ∈(0,1) and confidence δ∈(0,1), their respective orders are \[ \!\1, ρ 1k-1k \·(d+ δ-1) \!\12, ρ1k-1k 2 \·(d+ δ-1), \] where k=k/(k-1). For every fixed ρ>0, robustness changes the realizable -dependence from -1 to -k as 0. In the agnostic case, for 1<k<2, robustness changes the -dependence from -2 to -k, whereas for k≥2 the exponent remains the classical 2, with nontrivial ρ-dependence. Building on the known scalar reduction of robust 0--1 risk to ordinary classification error, our analysis reveals a scale-sensitive interaction between the statistical estimation of classification error and its amplification by robustness, sharply explaining the transition in the agnostic rate. We extend the previously studied χ2-divergence case to every Cressie--Read order k>1, close its upper--lower gaps, and recover standard PAC learning rates as ρ0, unlike previous bounds that fail to interpolate correctly in this limit.
Create a lesson
Related papers
How Model Growth, Recursion, and Boundary Operators Influence Scaling Exponents
Zixi Chen, Akshay Vegesna, Samip Dahal et al.
Evidence-Grounded Agentic Formulation Development in an Autonomous Laboratory
Michael M. Craig, Riley J. Hickman, Yingshan Ma et al.
Probabilistic Linear Explanations
Frederic Koriche, Jean-Marie Lagniez, Chi Tran
Double descent is the principle of least action
Congzhou M Sha
RLLBC-Lib: An Educational Code Library for Reinforcement Learning and Learning-Based Control
Bernd Frauenknecht, Emma Cramer, Artur Eisele et al.
Higher-order pruning of experts in mixture-of-experts language models
Alex M. Tseng, Prannay Kaul, Luca Zancato et al.