Complexity control by gradient descent in deep networks.

Poggio, Tomaso; Liao, Qianli; Banburski, Andrzej · Nat Commun · 2020

basic_science · Level V

Where this comes from

Abstract

Overparametrized deep networks predict well, despite the lack of an explicit complexity control during training, such as an explicit regularization term. For exponential-type loss functions, we solve this puzzle by showing an effective regularization effect of gradient descent in terms of the normalized weights that are relevant for classification.