Fetching the paper…
Reading the bibliography…
Deep learning has exhibited superior performance for various tasks, especially for high-dimensional datasets, such as images.
Approximation of functions of several variables and imbedding theorems , volume 205
S. M. Nikol’skii · 1975
Earlier work this paper cites.
Optimal filtration of a function of many variables in white gaussian noise
M. Nyssbaum · 1983
Earlier work this paper cites.
More on the estimation of distribution densities
I. Ibragimov and R. Khas’minskii · 1984
Earlier work this paper cites.
Nonparametric estimation of a regression function that is smooth in a domain in ℝ k \mathbb{R}^{k}
M. Nyssbaum · 1987
Earlier work this paper cites.
Interpolation of Besov spaces
R. A. DeVore and V. A. Popov · 1988
Earlier work this paper cites.
Capabilities of three-layered perceptrons
B. Irie and S. Miyake · 1988
Earlier work this paper cites.
Approximation by superpositions of a sigmoidal function
G. Cybenko · 1989
Earlier work this paper cites.
Universal approximation of an unknown mapping and its derivatives using multilayer feedforward networks
K. Hornik, M. Stinchcombe, and H. White · 1990
Earlier work this paper cites.
Approximation capabilities of multilayer feedforward networks
K. Hornik · 1991
Earlier work this paper cites.
Density estimation in Besov spaces
G. Kerkyacharian and D. Picard · 1992
Earlier work this paper cites.
Wavelet compression and nonlinearn-widths
R. A. DeVore, G. Kyriazis, D. Leviatan, and V. M. Tikhomirov · 1993
Earlier work this paper cites.
Approximation of Periodic Functions
V. Temlyakov · 1993
Earlier work this paper cites.
Density estimation by wavelet thresholding
D. L. Donoho, I. M. Johnstone, G. Kerkyacharian, and D. Picard · 1996
Earlier work this paper cites.
Weak Convergence and Empirical Processes: With Applications to Statistics
A. W. van der Vaart and J. A. Wellner · 1996
Earlier work this paper cites.
Nonlinear approximation
R. A. DeVore · 1998
Earlier work this paper cites.
Minimax estimation via wavelet shrinkage
D. L. Donoho and I. M. Johnstone · 1998
Earlier work this paper cites.
A Wavelet Tour of Signal Processing
S. Mallat · 1999
Earlier work this paper cites.
Information-theoretic determination of minimax rates of convergence
Y. Yang and A. Barron · 1999
Earlier work this paper cites.
A global geometric framework for nonlinear dimensionality reduction
J. B. Tenenbaum, V. De Silva, and J. C. Langford · 2000
Earlier work this paper cites.
Nonlinear estimation in anisotropic multi-index denoising
G. Kerkyacharian, O. Lepski, and D. Picard · 2001
Cited alongside, same era.
Random rates in anisotropic regression (with a discussion and a rejoinder by the authors)
M. Hoffman and O. Lepski · 2002
Cited alongside, same era.
Wavelet threshold estimation of a regression function with random design
S. Zhang, M.-Y. Wong, and Z. Zheng · 2002
Cited alongside, same era.
Laplacian eigenmaps for dimensionality reduction and data representation
M. Belkin and P. Niyogi · 2003
Cited alongside, same era.
Nonlinear wavelet approximation in anisotropic besov spaces
C. Leisner · 2003
Cited alongside, same era.
Best practices for convolutional neural networks applied to visual document analysis
P. Y. Simard, D. Steinkraus, and J. C. Platt · 2003
Unsupervised Representation Learning with Deep Convolutional Generative Adversarial Networks
A. Radford, L. Metz, and S. Chintala · 2015
Later among the works it cites.
Minimax-optimal nonparametric regression in high dimensions
Y. Yang and S. T. Tokdar · 2015
Later among the works it cites.
Kolmogorov widths of the anisotropic besov classes of periodic functions of many variables
V. V. Myronyuk · 2016
Later among the works it cites.
Bayesian manifold regression
Y. Yang and D. B. Dunson · 2016
Later among the works it cites.
Widths of the anisotropic besov classes of periodic functions of several variables
V. V. Myronyuk · 2017
Later among the works it cites.
Neural network with unbounded activation functions is universal approximator
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Function spaces with dominating mixed smoothness
J. Vybiral · 2006
Cited alongside, same era.
Local polynomial regression on unknown manifolds
P. J. Bickel and B. Li · 2007
Cited alongside, same era.
Introduction to Nonparametric Estimation
A. B. Tsybakov · 2008
Cited alongside, same era.
Rectified linear units improve restricted boltzmann machines
V. Nair and G. E. Hinton · 2010
Cited alongside, same era.
Adaptive dimension reduction with a gaussian process prior
A. Bhattacharya, D. Pati, and D. B. Dunson · 2011
Cited alongside, same era.
Hyper-sparse optimal aggregation
S. Gaiffas and G. Lecue · 2011
Cited alongside, same era.
S. Sonoda and N. Murata · 2017
Later among the works it cites.
Error bounds for approximations with deep relu networks
D. Yarotsky · 2017
Later among the works it cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2018
Later among the works it cites.
Optimal learning with anisotropic gaussian svms
H. Hang and I. Steinwart · 2018
Later among the works it cites.
Nonparametric regression using deep neural networks with ReLU activation function
J. Schmidt-Hieber · 2018
Later among the works it cites.
Fast generalization error bound of deep learning from a kernel perspective
T. Suzuki · 2018
Later among the works it cites.
Nonparametric Regression on Low-Dimensional Manifolds using Deep ReLU Networks
M. Chen, H. Jiang, W. Liao, and T. Zhao · 2019
Closest in time.
Efficient approximation of deep relu networks for functions on low dimensional manifolds
M. Chen, H. Jiang, W. Liao, and T. Zhao · 2019
Closest in time.
S. Hayakawa and T. Suzuki · 2019
Closest in time.
Deep ReLU network approximation of functions on a manifold
J. Schmidt-Hieber · 2019
Closest in time.
Adaptivity of deep ReLU network for learning in Besov and mixed smooth Besov spaces: optimal rate and curse of dimensionality
T. Suzuki · 2019
Closest in time.
Rapid convergence of the unadjusted langevin algorithm: Isoperimetry suffices
S. Vempala and A. Wibisono · 2019
Closest in time.
Adaptive approximation and generalization of deep neural network with intrinsic dimensionality
R. Nakada and M. Imaizumi · 2020
Closest in time.
Benefit of deep learning with non-convex noisy gradient descent: Provable excess risk bound and superiority to kernel methods
T. Suzuki and S. Akiyama · 2021
Closest in time.