A Dynamics Theory of RMSProp-Based Implicit Regularization in Deep Low-Rank Matrix Factorization.

Cao, Jian; Qian, Chen; Huang, Yihui; Chen, Dicheng; Gao, Yuncheng; Dong, Jiyang; Guo, Di; Qu, Xiaobo · IEEE Trans Neural Netw Learn Syst · 2025

basic_science · Level V

Where this comes from

Abstract

Implicit regularization induced by gradient optimization is an important way to understand generalization in neural networks. Recent theory explains implicit regularization over the deep matrix factorization (DMF) model and analyzes the trajectory of discrete gradient dynamics in the optimization process. These discrete gradient dynamics can mathematically characterize the practical learning rate of adaptive gradient (AdaGrad) optimization, such as root-mean-square propagation (RMSProp). Discrete gradient dynamics analysis has been successfully applied to shallow networks but encounters difficulty in complex computation for deep networks. In this work, we introduce another discrete gradient dynamics, landscape analysis, to theoretically and experimentally explain the implicit regularization of RMSProp-based deep networks. It mainly focuses on gradient regions like saddle points and local minima. We investigate the benefits of increasing learning rates in saddle point escaping (SPE) stages. We prove that, for a rank-R matrix reconstruction, DMF will converge to a second-order critical point after R stages of SPE. Besides, we analyze the time it takes to escape from the plateau of the SPE stage. These conclusions are further experimentally verified on low-rank matrix, image reconstruction, and Hankel matrix reconstruction problems. Our proof is also applicable to gradient descent (GD) and adaptive moment estimation (Adam) but cannot apply to AdaGrad, further showing experimentally that the implicit regularization capability of RMSProp is stronger than GD and AdaGrad and weaker than Adam.