Skip to content

The Sample Complexity of Distributionally Robust PAC Learning under Cressie--Read Divergences

Elad Aigner-Horev, Daniel Rosenberg, Roi Weiss

cs.LGarXiv:2608.04686

Abstract

We study distributionally robust PAC learning for the 0--1-loss, where adversarial perturbations of the data distribution are constrained by a Cressie--Read divergence of order k>1 and radius ρ≥ 0. For hypothesis classes with VC dimension d, we establish realizable and agnostic sample-complexity bounds tight up to constant and logarithmic factors, respectively; ordinary empirical risk minimization attains both rates up to logarithmic factors. For target accuracy ∈(0,1) and confidence δ∈(0,1), their respective orders are \[ \!\1, ρ 1k-1k \·(d+ δ-1) \!\12, ρ1k-1k 2 \·(d+ δ-1), \] where k=k/(k-1). For every fixed ρ>0, robustness changes the realizable -dependence from -1 to -k as 0. In the agnostic case, for 1<k<2, robustness changes the -dependence from -2 to -k, whereas for k≥2 the exponent remains the classical 2, with nontrivial ρ-dependence. Building on the known scalar reduction of robust 0--1 risk to ordinary classification error, our analysis reveals a scale-sensitive interaction between the statistical estimation of classification error and its amplification by robustness, sharply explaining the transition in the agnostic rate. We extend the previously studied χ2-divergence case to every Cressie--Read order k>1, close its upper--lower gaps, and recover standard PAC learning rates as ρ0, unlike previous bounds that fail to interpolate correctly in this limit.

Create a lesson