High Probability Derivative Bounds for Random tanh Neural Networks on a Hypercube
Josef Dick, Michael Feischl, Fabian Zehetgruber
Abstract
We establish high-probability bounds for mixed input derivatives of wide random neural networks whose activation derivatives satisfy a factorial growth bound. Our main result specializes these estimates to networks with Xavier initialization. A direct deterministic analysis based on Euclidean operator norms of the weight matrices yields derivative bounds that generally grow exponentially with the depth. We show that this growth can be substantially improved for sufficiently wide Gaussian networks by isolating the term that is linear in the highest-order derivative and controlling the corresponding tangent directions by measurable finite nets. For scalar-output networks with Gaussian weights and Xavier initialization, we prove that there exist constants C,C0,C1>0 such that, whenever the common hidden width satisfies n ≥ C(L3n02(1+ n0)+L2(1+(L/η))), then, with probability at least 1-η, the estimate |DuRΦ(L)(x)| ≤ C0 |u|! (C1L)|u|-1Πj∈ uβj(η,n0) holds simultaneously for every non-empty u⊂eq[n0] and every x∈[0,1]n0. Thus, the first-order derivative bound is independent of the depth, while a square-free mixed derivative of order |u| grows at most polynomially as L|u|-1, apart from the coordinate factors. As consequences, we obtain high-probability bounds for the Euclidean Lipschitz constant and for weighted Sobolev norms of the network realization. The latter connect the derivative estimates to quasi-Monte Carlo integration and indicate how such regularity can enter the analysis of QMC-based training.
Create a lesson
Related papers
How Model Growth, Recursion, and Boundary Operators Influence Scaling Exponents
Zixi Chen, Akshay Vegesna, Samip Dahal et al.
Evidence-Grounded Agentic Formulation Development in an Autonomous Laboratory
Michael M. Craig, Riley J. Hickman, Yingshan Ma et al.
Probabilistic Linear Explanations
Frederic Koriche, Jean-Marie Lagniez, Chi Tran
Double descent is the principle of least action
Congzhou M Sha
RLLBC-Lib: An Educational Code Library for Reinforcement Learning and Learning-Based Control
Bernd Frauenknecht, Emma Cramer, Artur Eisele et al.
Higher-order pruning of experts in mixture-of-experts language models
Alex M. Tseng, Prannay Kaul, Luca Zancato et al.