Generalization, memorization, and overfitting for diffusion models trained in the lazy high-dimensional regime
Hugo Latourelle-Vigeant, Sinho Chewi, Aram-Alexandre Pooladian, John Sous, Theodor Misiakiewicz
Abstract
Modern score-based generative models have achieved remarkable empirical success in high-dimensional tasks such as image, audio, and video synthesis. These models reduce distribution learning to a sequence of regression problems that, if solved exactly on finite data, would ultimately reproduce the training samples. Their ability to generalize must therefore arise from the implicit or explicit regularization during training. In this work, we develop a generative counterpart to the theory of benign overfitting and algorithmic regularization for overparameterized neural networks in the supervised lazy-training regime. We study denoising score matching in a vector-valued reproducing kernel Hilbert space with an inner-product kernel. In the proportional high-dimensional regime n d, we derive exact risk trajectories under gradient flow training. These trajectories exhibit three phases governed by qualitatively distinct estimators: a spectral estimator that generalizes, a pure-noise score with localized peaks that interpolate the training objective, and an empirical Bayes estimator that memorizes the data. We then analyze how these estimators combine along the reverse-time SDE and characterize the distribution of the resulting samples. The analysis reveals familiar mechanisms from supervised learning, including kernel linearization and self-induced regularization from the nonlinear part of the kernel, but also reveals a distinct phenomenology specific to generative modeling.
Create a lesson
Related papers
A General Kernel Framework for Non-CND Distance Measures Using |D|-Dimensional Sparse Landmark Embeddings
Marcus M. Noack, Maher B. Alghalayini, Mark D. Risser
Fast Learning Rates for Physics-Informed Kernel Methods
Luc Brogat-Motte, Joachim Bona-Pellissier, Giacomo Meanti et al.
Rank and computation of the pathlifting Jacobian of a DAG ReLU network
Manon Verbockhaven
Preservation of Log-Concavity and Convergence of Wasserstein-Fisher-Rao Gradient Flows
Francesca Romana Crucinio, Sahani Pathiraja
Generalized DCCQ: From Binary Quotients to Multinomial Simplex Geometry and Critical-Strip Coordinates
Y. Kenan Yılmaz
Bracketing Uncertainty in Clustering Under the Manifold Hypothesis
Savik Kinger, Luciano Dyballa, Steven W. Zucker