Fetching the paper…
Reading the bibliography…
We introduce a new theoretical framework to analyze deep learning optimization with connection to its generalization error.
A generalization theory of gradient descent for learning over-parameterized deep ReLU networks
Y. Cao and Q. Gu · 1902
Earlier work this paper cites.
Generalization bounds of stochastic gradient descent for wide and deep neural networks
Y. Cao and Q. Gu · 1905
Earlier work this paper cites.
A. Nitanda and T. Suzuki · 1905
Earlier work this paper cites.
Singular Integrals and Differentiability Properties of Functions
E. M. Stein · 1970
Earlier work this paper cites.
The Brunn-Minkowski inequality in gauss space
C. Borell · 1975
Earlier work this paper cites.
Theory of Function Spaces
H. Triebel · 1983
Earlier work this paper cites.
Interpolation of Operators
C. Bennett and R. Sharpley · 1988
Earlier work this paper cites.
Strong Feller property for semilinear stochastic evolution equations and applications
B. Maslowski · 1989
Earlier work this paper cites.
Non-explosion, boundedness and ergodicity for stochastic semilinear equations
G. Da Prato and J. Zabczyk · 1992
Earlier work this paper cites.
A simple weight decay can improve generalization
A. Krogh and J. A. Hertz · 1992
Earlier work this paper cites.
Large deviations for the invariant measure of a reaction-diffusion equation with non-Gaussian perturbations
R. Sowers · 1992
Earlier work this paper cites.
Besov spaces on domains in ℝ d \mathbb{R}^{d}
R. A. DeVore and R. C. Sharpley · 1993
Earlier work this paper cites.
Ergodicité d’une classe d’équations aux dérivées partielles stochastiques
S. Jacquot and G. Royer · 1995
Earlier work this paper cites.
New concentration inequalities in product spaces
M. Talagrand · 1996
Earlier work this paper cites.
Geometric ergodicity for stochastic PDEs
T. Shardlow · 1999
Earlier work this paper cites.
Information-theoretic determination of minimax rates of convergence
Y. Yang and A. Barron · 1999
Earlier work this paper cites.
Convergence rates of posterior distributions
S. Ghosal, J. K. Ghosh, and A. W. van der Vaart · 2000
Earlier work this paper cites.
A Bennett concentration inequality and its application to suprema of empirical process
O. Bousquet · 2002
Earlier work this paper cites.
Exponential mixing properties of stochastic PDEs through asymptotic coupling
M. Hairer · 2002
Earlier work this paper cites.
Improving the sample complexity using global data
S. Mendelson · 2002
Earlier work this paper cites.
Optimal aggregation of classifiers in statistical learning
A. B. Tsybakov et al · 2004
Earlier work this paper cites.
Local Rademacher complexities
P. Bartlett, O. Bousquet, and S. Mendelson · 2005
Earlier work this paper cites.
Exponential convergence rates in classification
V. Koltchinskii and O. Beznosova · 2005
Earlier work this paper cites.
Empirical minimization
P. L. Bartlett and S. Mendelson · 2006
Earlier work this paper cites.
Concentration inequalities and asymptotic results for ratio type empirical processes
E. Giné and V. Koltchinskii · 2006
Cited alongside, same era.
Local Rademacher complexities and oracle inequalities in risk minimization
V. Koltchinskii · 2006
Cited alongside, same era.
Fast learning rates for plug-in classifiers
J.-Y. Audibert, A. B. Tsybakov, et al · 2007
Cited alongside, same era.
Support Vector Machines
I. Steinwart and A. Christmann · 2008
Cited alongside, same era.
Rates of contraction of posterior distributions based on Gaussian process priors
A. W. van der Vaart and J. H. van Zanten · 2008
Cited alongside, same era.
Optimal rates for regularized least squares regression
I. Steinwart, D. Hush, and C. Scovel · 2009
Cited alongside, same era.
A note on lazy training in supervised differentiable programming
L. Chizat and F. Bach · 2018
Later among the works it cites.
Global non-convex optimization with discretized diffusions
M. A. Erdogdu, L. Mackey, and O. Shamir · 2018
Later among the works it cites.
Neural tangent kernel: Convergence and generalization in neural networks
A. Jacot, F. Gabriel, and C. Hongler · 2018
Later among the works it cites.
A mean field view of the landscape of two-layer neural networks
S. Mei, A. Montanari, and P.-M. Nguyen · 2018
Later among the works it cites.
Generalization bounds of SGLD for non-convex learning: Two theoretical viewpoints
W. Mou, L. Wang, X. Zhai, and K. Zheng · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Information rates of nonparametric gaussian process methods
A. W. van der Vaart and J. H. van Zanten · 2011
Cited alongside, same era.
Bayesian learning via stochastic gradient Langevin dynamics
M. Welling and Y.-W. Teh · 2011
Cited alongside, same era.
Foundations of machine learning
M. Mohri, A. Rostamizadeh, and A. Talwalkar · 2012
Cited alongside, same era.
Concentration Inequalities: A Nonasymptotic Theory of Independence
S. Boucheron, G. Lugosi, and P. Massart · 2013
Cited alongside, same era.
Dropout training as adaptive regularization
S. Wager, S. Wang, and P. S. Liang · 2013
Cited alongside, same era.
Stochastic Equations in Infinite Dimensions
G. Da Prato and J. Zabczyk · 2014
Cited alongside, same era.
J. Sirignano and K. Spiliopoulos · 2018
Later among the works it cites.
A convergence theory for deep learning via over-parameterization
Z. Allen-Zhu, Y. Li, and Z. Song · 2019
Later among the works it cites.
Fine-grained analysis of optimization and generalization for overparameterized two-layer neural networks
S. Arora, S. Du, W. Hu, Z. Li, and R. Wang · 2019
Later among the works it cites.
Deep neural networks learn non-smooth functions effectively
M. Imaizumi and K. Fukumizu · 2019
Later among the works it cites.
Z. Ji and M. Telgarsky · 2019
Later among the works it cites.
Mean-field theory of two-layers neural networks: dimension-free bounds and kernel limit
S. Mei, T. Misiakiewicz, and A. Montanari · 2019
Later among the works it cites.
Convergence types and rates in generic karhunen-loève expansions with applications to sample path properties
I. Steinwart · 2019
Later among the works it cites.
Adaptivity of deep reLU network for learning in besov and mixed smooth besov spaces: optimal rate and curse of dimensionality
T. Suzuki · 2019
Later among the works it cites.
Deep learning is adaptive to intrinsic dimensionality of model smoothness in anisotropic besov space
T. Suzuki and A. Nitanda · 2019
Later among the works it cites.
High-Dimensional Statistics: A Non-Asymptotic Viewpoint
M. Wainwright · 2019
Later among the works it cites.
A comparative analysis of optimization and generalization properties of two-layer neural network and random feature models under gradient descent dynamics
E. Weinan, C. Ma, and L. Wu · 2019
Later among the works it cites.
An improved analysis of training over-parameterized deep neural networks
D. Zou and Q. Gu · 2019
Later among the works it cites.
Influence of the regularity of the test functions for weak convergence in numerical discretization of spdes
C.-E. Bréhier · 2020
Closest in time.
On the minimax optimality and superiority of deep neural network learning over sparse parameter spaces
S. Hayakawa and T. Suzuki · 2020
Closest in time.
Dimension-free convergence rates for gradient Langevin dynamics in RKHS
B. Muzellec, K. Sato, M. Massias, and T. Suzuki · 2020
Closest in time.
Towards moderate overparameterization: global convergence guarantees for training shallow neural networks
S. Oymak and M. Soltanolkotabi · 2020
Closest in time.
PAC-Bayes analysis beyond the usual bounds
O. Rivasplata, I. Kuzborskij, C. Szepesvári, and J. Shawe-Taylor · 2020
Closest in time.
Nonparametric regression using deep neural networks with ReLU activation function
J. Schmidt-Hieber · 2020
Closest in time.
Compression based bound for non-compressed network: Unified generalization error analysis of large compressible deep neural network
T. Suzuki, H. Abe, and T. Nishimura · 2020
Closest in time.