Skip to content

Momentum in large-batch training: Polyak enlarges the critical batch size, Nesterov improves data efficiency

Jia-Nan Wang, Zixun Huang, Kairui Li, Lei Wu

stat.MLarXiv:2609.02728

Abstract

We study when and how momentum improves large-batch training in the one-pass regime, using power-law kernel regression as a tractable setting. We first characterize risk stability through the critical learning rate, defined as the largest learning rate for stable training, and obtain ηSGDcrit 1, ηPolyakcrit \1,B(1-ρ)\, and ηNesterovcrit \1,Bβ(1-ρ)\, where B is the batch size, ρ is the momentum factor, and β>1 is the capacity exponent. Within this admissible region, we derive scaling laws for the full risk dynamics, capturing the progression from an early transient, through power-law decay, to a noise floor. We then minimize the final-step risk over the admissible learning rates and momentum factors under a fixed data budget, yielding a three-regime batch-size phase diagram that reveals how the role of momentum changes with batch size. Notably, Polyak enlarges the critical batch size, the largest batch size preserving the best small-batch data-scaling exponent, thereby enabling greater parallelism without sacrificing data efficiency. In contrast, Nesterov achieves better data efficiency in the large-batch regime because its look-ahead mechanism suppresses noise accumulation. Numerical experiments validate the predicted stability boundaries, risk dynamics, and batch-size phase diagram.

Create a lesson