ResearchPod Summary
This paper investigates the behavior of gradient descent (GD) when the step size exceeds the classical stability threshold (the reciprocal of the Hessian's largest eigenvalue, or sharpness). While standard optimization theory requires small step sizes for convergence, deep learning models often operate in a regime where large step sizes are used. The authors extend previous work on isolated flat minima to a more general setting: overparameterized least-squares problems with vector-valued outputs and manifolds of flat minima. They employ tools from dynamical systems and differential geometry to derive a normal form for GD, providing a coordinate system that separates the dynamics into movement along the solution manifold and movement in the orthogonal directions.
The authors establish a normal form for GD in the neighborhood of a manifold of flat minima. This framework reveals that the dynamics decompose into three components:
Furthermore, the authors apply this framework to deep matrix factorization, proving that the flat minima form a fiber bundle over a product of spheres and that the sharpness is Morse-Bott along this manifold, providing a rigorous structural understanding of the loss landscape in these models.
This work provides a theoretical foundation for the "edge of stability" phenomenon observed in deep learning, where large step sizes do not lead to divergence but instead bias the model toward flatter, more stable minima. By generalizing the analysis to manifolds of flat minima and arbitrary codimension, the paper bridges the gap between abstract optimization theory and the complex, high-dimensional loss landscapes encountered in practical deep learning applications like matrix factorization.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.