Approximation rates in Besov norms and sample-complexity of Kolmogorov-Arnold networks with residual connections.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 41850012.
- Also identified by DOI 10.1016/j.neunet.2026.108797.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Inspired by the Kolmogorov-Arnold superposition theorem, Kolmogorov-Arnold Networks (KANs) have recently emerged as an improved backbone for most deep learning frameworks, promising more adaptivity than their multilayer perceptron (MLP) predecessor by allowing for trainable spline-based activation functions. In this paper, we probe the theoretical foundations of the KAN architecture by showing that it can optimally approximate any Besov function in B<sub>p,q</sub><sup>s</sup>(X) on a bounded open, or even fractal, domain X in R<sup>d</sup> at the optimal approximation rate with respect to any weaker Besov norm B<sub>p,q</sub><sup>α</sup>(X); where α < s. We complement our approximation result with a statistical guarantee by bounding the pseudodimension of the relevant class of Res-KANs. As an application of the latter, we directly deduce a dimension-free estimate on the sample complexity of a residual KAN model when learning a function of Besov regularity from N i.i.d. noiseless samples, showing that KANs can learn the smooth maps which they can approximate.
Medical subject headings
- Neural Networks, Computer
- Deep Learning