Statistical physics and practical training of soft-committee machines
Martin Ahr, Michael Biehl, Robert Urbanczik
Abstract
Equilibrium states of large layered neural networks with differentiable activation function and a single, linear output unit are investigated using the replica formalism. The quenched free energy of a student network with a very large number of hidden units learning a rule of perfectly matching complexity is calculated analytically. The system undergoes a first order phase transition from unspecialized to specialized student configurations at a critical size of the training set. Computer simulations of learning by stochastic gradient descent from a fixed training set demonstrate that the equilibrium results describe quantitatively the plateau states which occur in practical training procedures at sufficiently small but finite learning rates.
Create a lesson
Related papers
Low-temperature magnetism and spin dynamics in the disordered triangular-lattice Yb3+ compound LiCaYb5(BO3)6
Monika Jawale, Saikat Nandi, Prashanta K. Mukharjee et al.
Neural Renormalization Group Flow for Percolation
Anaclara Alvez, Luca Camagna, Sergio Chibbaro et al.
Dynamical phase selection controls compute scaling in looped transformers
Gunn Kim
Semi-localized ground state in a 1D system with long-range hopping
Murod S. Bahovadinov, Faridun N. Jalolov, Vladimir E. Kravtsov et al.
Defect states in three-dimensional diamond photonic band gap crystals
Julia Rocha, Bart A. van Tiggelen, Ad Lagendijk et al.
Disorder-induced conducting edges on Kagomé lattice
A. Chmeruk, D. Jones, L. Chioncel