Fetching the paper…
Reading the bibliography…
The purpose of this article is to review the achievements made in the last few years towards the understanding of the reasons behind the success and subtleties of neural network-based machine learning.
Theory of reproducing kernels
Nachman Aronszajn · 1950
Earlier work this paper cites.
Approximation by superpositions of a sigmoidal function
George Cybenko · 1989
Earlier work this paper cites.
Neural net approximation
Andrew R Barron · 1992
Earlier work this paper cites.
Universal approximation bounds for superpositions of a sigmoidal function
Andrew R. Barron · 1993
Earlier work this paper cites.
Multilayer feedforward networks with a nonpolynomial activation function can approximate any function
Moshe Leshno, Vladimir Ya Lin, Allan Pinkus, and Shimon Schocken · 1993
Earlier work this paper cites.
Flat minima
Sepp Hochreiter and Jurgen Schmidhuber · 1997
Earlier work this paper cites.
A convexity principle for interacting gases
Robert J McCann · 1997
Earlier work this paper cites.
Efficient noise-tolerant learning from statistical queries
Michael Kearns · 1998
Earlier work this paper cites.
Uniform approximation by neural networks
Y Makovoz · 1998
Earlier work this paper cites.
On best approximation by ridge functions
VE Maiorov · 1999
Earlier work this paper cites.
Lower bounds for approximation by MLP neural networks
Vitaly Maiorov and Allan Pinkus · 1999
Earlier work this paper cites.
Approximation theory of the mlp model in neural networks
Allan Pinkus · 1999
Earlier work this paper cites.
Rademacher and gaussian complexities: Risk bounds and structural results
Peter L. Bartlett and Shahar Mendelson · 2002
Earlier work this paper cites.
Stability and generalization
Olivier Bousquet and André Elisseeff · 2002
Earlier work this paper cites.
Local Rademacher complexities
Peter L Bartlett, Olivier Bousquet, Shahar Mendelson, et al · 2005
Earlier work this paper cites.
Optimal transport: old and new
Cédric Villani · 2008
Earlier work this paper cites.
Efficient backprop
Yann A LeCun, Léon Bottou, Genevieve B Orr, and Klaus-Robert Müller · 2012
Earlier work this paper cites.
Lecture 6.5—RmsProp: Divide the gradient by a running average of its recent magnitude
T. Tieleman and G. Hinton · 2012
Earlier work this paper cites.
Random matrices and complexity of spin glasses
Antonio Auffinger, Gérard Ben Arous, and Jiří Černỳ · 2013
Earlier work this paper cites.
Invariant scattering convolution networks
Joan Bruna and Stéphane Mallat · 2013
Earlier work this paper cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Andrew M Saxe, James L McClelland, and Surya Ganguli · 2013
Earlier work this paper cites.
On the computational efficiency of training neural networks
Roi Livni, Shai Shalev-Shwartz, and Ohad Shamir · 2014
Earlier work this paper cites.
Understanding machine learning: From theory to algorithms
Shai Shalev-Shwartz and Shai Ben-David · 2014
Earlier work this paper cites.
The loss surfaces of multilayer networks
Anna Choromanska, Mikael Henaff, Michael Mathieu, Gérard Ben Arous, and Yann LeCun · 2015
Earlier work this paper cites.
On the rate of convergence in Wasserstein distance of the empirical measure
Nicolas Fournier and Arnaud Guillin · 2015
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
The power of depth for feedforward neural networks
Ronen Eldan and Ohad Shamir · 2016
Earlier work this paper cites.
Deep learning without poor local minima
Kenji Kawaguchi · 2016
Earlier work this paper cites.
Risk bounds for high-dimensional ridge function combinations including neural networks
Jason M Klusowski and Andrew R Barron · 2016
Earlier work this paper cites.
Local minima in training of deep networks
Grzegorz Swirszcz, Wojciech Marian Czarnecki, and Razvan Pascanu · 2016
Earlier work this paper cites.
Multiscale adaptive representation of signals: I. the basic framework
Cheng Tai and Weinan E · 2016
Cited alongside, same era.
Statistical physics of inference: Thresholds and algorithms
Lenka Zdeborová and Florent Krzakala · 2016
Cited alongside, same era.
High-dimensional dynamics of generalization error in neural networks
Madhu S Advani and Andrew M Saxe · 2017
Cited alongside, same era.
Breaking the curse of dimensionality with convex neural networks
Francis Bach · 2017
Cited alongside, same era.
On the equivalence between kernel quadrature rules and random feature expansions
Francis Bach · 2017
Cited alongside, same era.
SGD learns the conjugate kernel class of the network
Amit Daniely · 2017
Theoretical insights into the optimization landscape of over-parameterized shallow neural networks
Mahdi Soltanolkotabi, Adel Javanmard, and Jason D Lee · 2018
Later among the works it cites.
How SGD selects the global minima in over-parameterized learning: A dynamical stability perspective
Lei Wu, Chao Ma, and Weinan E · 2018
Later among the works it cites.
On exact computation with an infinitely wide neural net
Sanjeev Arora, Simon S Du, Wei Hu, Zhiyuan Li, Ruslan Salakhutdinov, and Ruosong Wang · 2019
Later among the works it cites.
Optimal errors and phase transitions in high-dimensional generalized linear models
Jean Barbier, Florent Krzakala, Nicolas Macris, Léo Miolane, and Lenka Zdeborová · 2019
Later among the works it cites.
Reconciling modern machine-learning practice and the classical bias–variance trade-off
Mikhail Belkin, Daniel Hsu, Siyuan Ma, and Soumik Mandal · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Computing nonvacuous generalization bounds for deep (stochastic) neural networks with many more parameters than training data
Gintare Karolina Dziugaite and Daniel M. Roy · 2017
Cited alongside, same era.
Densely connected convolutional networks
Gao Huang, Zhuang Liu, Laurens Van Der Maaten, and Kilian Q Weinberger · 2017
Cited alongside, same era.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish S. Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping T.P. Tang · 2017
Cited alongside, same era.
Maximum principle based algorithms for deep learning
Qianxiao Li, Long Chen, Cheng Tai, and Weinan E · 2017
Cited alongside, same era.
Learning ReLUs via gradient descent
Mahdi Soltanolkotabi · 2017
Cited alongside, same era.
Yuandong Tian · 2017
Cited alongside, same era.
A deep neural network’s loss surface contains every low-dimensional pattern
Wojciech Marian Czarnecki, Simon Osindero, Razvan Pascanu, and Max Jaderberg · 2019
Later among the works it cites.
Gradient descent provably optimizes over-parameterized neural networks
Simon S. Du, Xiyu Zhai, Barnabas Poczos, and Aarti Singh · 2019
Later among the works it cites.
A priori estimates of the population risk for residual networks
Weinan E, Chao Ma, and Qingcan Wang · 2019
Later among the works it cites.
Barron spaces and the compositional function spaces for neural network models
Weinan E, Chao Ma, and Lei Wu · 2019
Later among the works it cites.
Weinan E, Chao Ma, and Lei Wu · 2019
Later among the works it cites.
Machine learning from a continuous viewpoint
Weinan E, Chao Ma, and Lei Wu · 2019
Later among the works it cites.
Mean-field Langevin dynamics and energy landscape of neural networks
Kaitong Hu, Zhenjie Ren, David Siska, and Lukasz Szpruch · 2019
Later among the works it cites.
Analysis of a two-layer neural network via displacement convexity
Adel Javanmard, Marco Mondelli, and Andrea Montanari · 2019
Later among the works it cites.
Explaining landscape connectivity of low-cost solutions for multilayer nets
Rohith Kuditipudi, Xiang Wang, Holden Lee, Yi Zhang, Zhiyuan Li, Wei Hu, Rong Ge, and Sanjeev Arora · 2019
Later among the works it cites.
Loss landscape sightseeing with multi-point optimization
Ivan Skorokhodov and Mikhail Burtsev · 2019
Later among the works it cites.
Global convergence of gradient descent for deep linear residual networks
Lei Wu, Qingcan Wang, and Chao Ma · 2019
Later among the works it cites.
Poly-time universality and limitations of deep learning
Emmanuel Abbe and Colin Sandon · 2020
Closest in time.
A dynamical central limit theorem for shallow neural networks
Zhengdao Chen, Grant Rotskoff, Joan Bruna, and Eric Vanden-Eijnden · 2020
Closest in time.
Implicit bias of gradient descent for wide two-layer neural networks trained with the logistic loss
Lenaic Chizat and Francis Bach · 2020
Closest in time.
Weinan E and Stephan Wojtowytsch · 2020
Closest in time.
On the Banach spaces associated with multi-layer ReLU networks of infinite width
Weinan E and Stephan Wojtowytsch · 2020
Closest in time.
Representation formulas and pointwise properties for barron functions
Weinan E and Stephan Wojtowytsch · 2020
Closest in time.
The gaussian equivalence of generative models for learning with two-layer neural networks
Sebastian Goldt, Galen Reeves, Marc Mézard, Florent Krzakala, and Lenka Zdeborová · 2020
Closest in time.
Complexity measures for neural networks with general activation functions using path-based norms
Zhong Li, Chao Ma, and Lei Wu · 2020
Closest in time.
A qualitative study of the dynamic behavior of adaptive gradient algorithms
Chao Ma, Lei Wu, and Weinan E · 2020
Closest in time.
Chao Ma, Lei Wu, and Weinan E · 2020
Closest in time.
The slow deterioration of the generalization error of the random feature model
Chao Ma, Lei Wu, and Weinan E · 2020
Closest in time.
Optimization and generalization of shallow neural networks with quadratic activation functions
Stefano Sarao Mannelli, Eric Vanden-Eijnden, and Lenka Zdeborová · 2020
Closest in time.
A rigorous framework for the mean field limit of multilayer neural networks
Phan-Minh Nguyen and Huy Tuan Pham · 2020
Closest in time.
Stephan Wojtowytsch · 2020
Closest in time.