Neural network approximation: Three hidden layers are enough.

Shen, Zuowei; Yang, Haizhao; Zhang, Shijun · Neural Netw · 2021

basic_science · Level V

Where this comes from

Abstract

A three-hidden-layer neural network with super approximation power is introduced. This network is built with the floor function (⌊x⌋), the exponential function (2<sup>x</sup>), the step function (1<sub>x≥0</sub>), or their compositions as the activation function in each neuron and hence we call such networks as Floor-Exponential-Step (FLES) networks. For any width hyper-parameter N∈N<sup>+</sup>, it is shown that FLES networks with width max{d,N} and three hidden layers can uniformly approximate a Hölder continuous function f on [0,1]<sup>d</sup> with an exponential approximation rate 3λ(2d)<sup>α</sup>2<sup>-αN</sup>, where α∈(0,1] and λ>0 are the Hölder order and constant, respectively. More generally for an arbitrary continuous function f on [0,1]<sup>d</sup> with a modulus of continuity ω<sub>f</sub>(⋅), the constructive approximation rate is 2ω<sub>f</sub>(2d)2<sup>-N</sup>+ω<sub>f</sub>(2d2<sup>-N</sup>). Moreover, we extend such a result to general bounded continuous functions on a bounded set E⊆R<sup>d</sup>. As a consequence, this new class of networks overcomes the curse of dimensionality in approximation power when the variation of ω<sub>f</sub>(r) as r→0 is moderate (e.g., ω<sub>f</sub>(r)≲r<sup>α</sup> for Hölder continuous functions), since the major term to be concerned in our approximation rate is essentially d times a function of N independent of d within the modulus of continuity. Finally, we extend our analysis to derive similar approximation results in the L<sup>p</sup>-norm for p∈[1,∞) via replacing Floor-Exponential-Step activation functions by continuous activation functions.

Medical subject headings