Every Layer Counts: An Exponential L2 Depth Hierarchy for ReLU Networks
Itay Safran
Abstract
We prove a depth hierarchy for ReLU neural networks in which every additional ReLU layer can save exponentially many neurons. For all k≥2, we construct a globally [0,1]-valued, 1-Lipschitz function realized by a depth-(k+1) network of width O(d4), whereas any depth-k network with unrestricted weights and width at most 2d2d(k-1) has squared L2 error at least 1/24 under an absolutely continuous distribution supported at exponential distance from the origin. To the best of our knowledge, this is the first exponential hierarchy across all adjacent fixed depths, and the first exponential separation for ReLU networks between two fixed depths whose shallower network has depth at least 3. The lower bound also immediately yields the corresponding hierarchy for exact computation. Moreover, the case k=2 gives a compactly supported separation between depths 3 and 2 with unrestricted shallow-network weights, answering a question raised by Safran, Eldan, and Shamir (2019). The distribution used in our construction nevertheless has all its mass at exponential radius, placing the hierarchy outside the regularity regime in which such a separation would imply major threshold-circuit lower bounds. We also prove an exact separation for a more regular target, which is globally [0,1]-valued and O( d)-Lipschitz and maps the unit hypercube onto [0,1]. It is computed by a polynomial-width depth-4 network, whereas any depth-3 network agreeing with it on the unit hypercube requires exponentially many first-layer neurons, even with unrestricted weights.
Create a lesson
Related papers
How Model Growth, Recursion, and Boundary Operators Influence Scaling Exponents
Zixi Chen, Akshay Vegesna, Samip Dahal et al.
Evidence-Grounded Agentic Formulation Development in an Autonomous Laboratory
Michael M. Craig, Riley J. Hickman, Yingshan Ma et al.
Probabilistic Linear Explanations
Frederic Koriche, Jean-Marie Lagniez, Chi Tran
Double descent is the principle of least action
Congzhou M Sha
RLLBC-Lib: An Educational Code Library for Reinforcement Learning and Learning-Based Control
Bernd Frauenknecht, Emma Cramer, Artur Eisele et al.
Higher-order pruning of experts in mixture-of-experts language models
Alex M. Tseng, Prannay Kaul, Luca Zancato et al.