Skip to content

Optimistic Rates for Multiclass PAC Learning

Xiaoyu Li, Andi Han, Jiaojiao Jiang, Junbin Gao

cs.LGarXiv:2608.10869

Abstract

Worst-case multiclass bounds do not become smaller when the best classifier is already nearly correct: what is missing is an optimistic rate, a guarantee whose fluctuation scales with the oracle risk itself. For a class of Natarajan dimension dN and Daniely-Shalev-Shwartz dimension dDS, the optimal excess risk is known at the two endpoints (dDS/n realizable, dN/n+dDS/n agnostic [HMZ24, CEH+26, Pab26]) and open in between. We close the gap: at every fixed oracle risk L, the optimal excess risk is Θ(L dN/n+dDS/n), uniformly in the alphabet size, attained by a learner that knows neither L nor the confidence level. The upper bound composes the cover-menu-compression architecture of [CEH+26], at the realizable rate of [Pab26], with a new comparator-facing relative compression theorem: a size-k compression rule that empirically dominates a comparator h has population risk at most L(h)+O(L(h)Γ+Γ) with Γ=(k n+(1/δ))/n, without stability; this transfers the comparison principle of the sharp binary theory [MQZ26] while discarding its Boolean-cube geometry, which does not lift to multiclass labels. The lower bound forces both terms using one class and one distribution at every fixed L, by a pair-Assouad scheme calibrated to L and a fiber argument on the pseudo-cubes underlying the Natarajan-versus-DS separation of [BCD+22]. Both theorems extend to list learning: against the best r-tuple of hypotheses, the same architecture and the same two engines yield an optimistic rate and a lower bound of the same shape, forcing the fluctuation term that [Pab26] expected to be necessary against list comparators, and removing the factor r from the known realizable list lower bound.

Create a lesson