Complexity control by gradient descent in deep networks.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 32094327.
- Also identified by DOI 10.1038/s41467-020-14663-9 and PMC identifier 7039878.
- Licence recorded as CC BY.
- The licence permits redistribution, so the abstract is shown in full and the full text is available from the publisher.
Abstract
Overparametrized deep networks predict well, despite the lack of an explicit complexity control during training, such as an explicit regularization term. For exponential-type loss functions, we solve this puzzle by showing an effective regularization effect of gradient descent in terms of the normalized weights that are relevant for classification.