Least-time Gradient Flow
Alessandro Betti, Marco Gori, Stefano Melacci, Jinwei Zhao
Abstract
Prescribing the speed of gradient flow on the risk itself, by the dynamics w=-u(E(w))∇ E(w)/∇ E(w)2, makes the risk e(t)=E(w(t)) obey e=-u(e) exactly, whatever the landscape~E; the time needed to reach zero risk from e0 is ∫0e0 e/u(e). Minimizing this time alone is ill posed, and we study the regularized problem ∈f\∫0e0(λ2u'2+1/u)\, e:\ u∈ H1(0,e0),\ u0,\ u(0)=0\, λ>0. We prove that the minimizer exists, is unique, and is a linearly scaled cycloid, and we show that the optimal rate behaves like u*(e)(9/(2λ))1/3e2/3 near zero risk: the exponent 2/3 is the one found in betti2026holder by a power-law ansatz, and it lies in the Hölder window (12,1) where the arrival is in finite time with vanishing weight speed. The proof follows the classical route: existence by the direct method, uniqueness by strict convexity, positivity of the minimizer away from the origin, and the explicit integration of the Euler-Lagrange equation.
Create a lesson
Related papers
TACO: Ternary Absolute-max Column-wise One-sparse Optimizer for LLM Fine-Tuning
Jichao Jiang, Cristian McGee, El Houcine Bergou et al.
FERPO: Forward Entropy-Regularized Policy Optimization
Sebastian Sanokowski, Alireza Sarmadi, Majid Khadiv
Cost-augmented Schrödinger bridges on graphs are exactly solvable: a Feynman-Kac tilt replaces learned control
Akshay Balsubramani
The Missing Primitive: Diagnosing and Repairing Mathematical Reasoning in Large Language Models
Shuo Xing, Zilin Dai, Chengyuan Qian et al.
Trust the Direction, Search the Step: Zero-and-First-Order Methods for LLM Fine-Tuning
Cristian McGee, El Houcine Bergou, Aritra Dutta
Generative modeling of intrinsically disordered protein regions by reinforcing sparse autoencoder features
Jason X. Liu, Sebastian Ibarraran, Frank Hu et al.