Learning between the peaks: sharp asymptotics for kernel ridge regression under power-law anisotropy
Lorenzo Rizzi, Arie Wortsman Zurich, Bruno Loureiro
Abstract
We study kernel ridge regression under anisotropic Gaussian data, where the input covariance decays as a power law with exponent α≥ 0 for polynomial inner-product kernels. We derive asymptotically sharp expressions for the kernel spectrum and the generalization error in the polynomial high-dimensional regime n=Θ(dκ), revealing how anisotropy reshapes the learning curves. For weak anisotropy (0<α<1), the problem remains effectively high-dimensional and retains some features of the isotropic case, while departing from it in others: the variance still peaks at integer sample complexities κ∈N, but these peaks are progressively damped as α grows; meanwhile, for targets strongly aligned with the data's principal directions, the bias drops at fractional sample complexities, decoupling the bias transitions from the interpolation peaks. For strong anisotropy (α> 1), the effective dimension of the problem is constant, and the variance stops depending on sample size altogether, plateauing under ridgeless interpolation or vanishing at an explicit rate under fixed ridge penalty. The bias undergoes a sharp transition governed by the target's decay rate: below a threshold, learning is abrupt rather than gradual; above it, the bias decays as a power law that recovers the classical source and capacity rates. We finally specialize these results to single-index targets, showing how the alignment of the index with the data's principal directions determines the effect of anisotropy on learning. Together, our results clarify how the input geometry shapes the kernel features and fundamentally impacts its generalization properties.
Create a lesson
Related papers
A General Kernel Framework for Non-CND Distance Measures Using |D|-Dimensional Sparse Landmark Embeddings
Marcus M. Noack, Maher B. Alghalayini, Mark D. Risser
Fast Learning Rates for Physics-Informed Kernel Methods
Luc Brogat-Motte, Joachim Bona-Pellissier, Giacomo Meanti et al.
Rank and computation of the pathlifting Jacobian of a DAG ReLU network
Manon Verbockhaven
Preservation of Log-Concavity and Convergence of Wasserstein-Fisher-Rao Gradient Flows
Francesca Romana Crucinio, Sahani Pathiraja
Generalized DCCQ: From Binary Quotients to Multinomial Simplex Geometry and Critical-Strip Coordinates
Y. Kenan Yılmaz
Bracketing Uncertainty in Clustering Under the Manifold Hypothesis
Savik Kinger, Luciano Dyballa, Steven W. Zucker