Momentum in large-batch training: Polyak enlarges the critical batch size, Nesterov improves data efficiency
Jia-Nan Wang, Zixun Huang, Kairui Li, Lei Wu
Abstract
We study when and how momentum improves large-batch training in the one-pass regime, using power-law kernel regression as a tractable setting. We first characterize risk stability through the critical learning rate, defined as the largest learning rate for stable training, and obtain ηSGDcrit 1, ηPolyakcrit \1,B(1-ρ)\, and ηNesterovcrit \1,Bβ(1-ρ)\, where B is the batch size, ρ is the momentum factor, and β>1 is the capacity exponent. Within this admissible region, we derive scaling laws for the full risk dynamics, capturing the progression from an early transient, through power-law decay, to a noise floor. We then minimize the final-step risk over the admissible learning rates and momentum factors under a fixed data budget, yielding a three-regime batch-size phase diagram that reveals how the role of momentum changes with batch size. Notably, Polyak enlarges the critical batch size, the largest batch size preserving the best small-batch data-scaling exponent, thereby enabling greater parallelism without sacrificing data efficiency. In contrast, Nesterov achieves better data efficiency in the large-batch regime because its look-ahead mechanism suppresses noise accumulation. Numerical experiments validate the predicted stability boundaries, risk dynamics, and batch-size phase diagram.
Create a lesson
Related papers
Copula Transformations for Data-Consistent Inversion
Troy Butler, Tianyi Jiang, João Silva et al.
Full-Model Optimality for Tunable Linear Generative Priors in Compressed Sensing
Zhaoming Li, Paul Hand
A computational approach to maximum likelihood thresholds for colored Gaussian graphical models
Roser Homs, Olga Kuznetsova, Bernadette J. Stolz
From topology learning to graph generation: A unifying perspective
Xiaowen Dong, Hoi-To Wai, Siheng Chen et al.
Schrödinger Bridges on Lie Group Manifolds for Probabilistic Intrinsic Generation
Shizhe Zhang, Mingyang Zhao, Lei Ma
HyperMC: Multi-Fidelity Hyperparameter Tuning for Stochastic Gradient MCMC
Ming Tan, Xiyun Jiao