Fetching the paper…
Reading the bibliography…
Neural network training is usually accomplished by solving a non-convex optimization problem using stochastic gradient descent.
On the approximate realization of continuous mappings by neural networks
K.-I. Funahashi · 1989
Earlier work this paper cites.
Error Bounds for Approximation with Neural Networks
M. Burger and A. Neubauer · 2001
Earlier work this paper cites.
Neural Network Learning: Theoretical Foundations
M. Anthony and P. Bartlett · 2009
Earlier work this paper cites.
Partial Differential Equations (second edition)
L. C. Evans · 2010
Earlier work this paper cites.
Improving neural networks by preventing co-adaptation of feature detectors
G. E. Hinton, N. Srivastava, A. Krizhevsky, I. Sutskever, and R. R. Salakhutdinov · 2012
Earlier work this paper cites.
Stochastic gradient descent for non-smooth optimization: Convergence results and optimal averaging schemes
O. Shamir and T. Zhang · 2013
Earlier work this paper cites.
Qualitatively characterizing neural network optimization problems
I. J. Goodfellow, O. Vinyals, and A. M. Saxe · 2014
Earlier work this paper cites.
The loss surfaces of multilayer networks
A. Choromanska, M. Henaff, M. Mathieu, G. B. Arous, and Y. LeCun · 2015
Earlier work this paper cites.
Measure Theory and Fine Properties of Functions, Revised Edition
L. C. Evans and R. F. Gariepy · 2015
Earlier work this paper cites.
Deep learning without poor local minima
K. Kawaguchi · 2016
Earlier work this paper cites.
Regularizing cnns with locally constrained decorrelations
P. Rodríguez, J. Gonzalez, G. Cucurull, J. M. Gonfaus, and X. Roca · 2016
Earlier work this paper cites.
On the quality of the initial basin in overspecified neural networks
I. Safran and O. Shamir · 2016
Earlier work this paper cites.
Optimal approximation with sparsely connected deep neural networks
H. Bölcskei, P. Grohs, G. Kutyniok, and P. Petersen · 2017
Earlier work this paper cites.
Sobolev training for neural networks
W. M. Czarnecki, S. Osindero, M. Jaderberg, G. Swirszcz, and R. Pascanu · 2017
Cited alongside, same era.
Size-independent sample complexity of neural networks
N. Golowich, A. Rakhlin, and O. Shamir · 2017
Cited alongside, same era.
Convergence analysis of two-layer neural networks with ReLU activation
Y. Li and Y. Yuan · 2017
Cited alongside, same era.
Exploring generalization in deep learning
B. Neyshabur, S. Bhojanapalli, D. McAllester, and N. Srebro · 2017
Cited alongside, same era.
The loss surface of deep and wide neural networks
Q. Nguyen and M. Hein · 2017
Cited alongside, same era.
Geometry of neural network loss surfaces via random matrix theory
J. Pennington and Y. Bahri · 2017
Cited alongside, same era.
Visualizing the loss landscape of neural nets
H. Li, Z. Xu, G. Taylor, C. Studer, and T. Goldstein · 2018
Later among the works it cites.
Learning overparameterized neural networks via stochastic gradient descent on structured data
Y. Li and Y. Liang · 2018
Later among the works it cites.
Darts: Differentiable architecture search
H. Liu, K. Simonyan, and Y. Yang · 2018
Later among the works it cites.
A mean field view of the landscape of two-layer neural networks
S. Mei, A. Montanari, and P.-M. Nguyen · 2018
Later among the works it cites.
The universal approximation power of finite-width deep ReLU networks
D. Perekrestenko, P. Grohs, D. Elbrächter, and H. Bölcskei · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Optimal approximation of piecewise smooth functions using deep ReLU neural networks
P. Petersen and F. Voigtlaender · 2017
Cited alongside, same era.
Error bounds for approximations with deep ReLU networks
D. Yarotsky · 2017
Cited alongside, same era.
A Convergence Theory for Deep Learning via Over-Parameterization
Z. Allen-Zhu, Y. Li, and Z. Song · 2018
Cited alongside, same era.
Stronger generalization bounds for deep nets via a compression approach
S. Arora, R. Ge, B. Neyshabur, and Y. Zhang · 2018
Cited alongside, same era.
Can we gain more from orthogonality regularizations in training deep networks?
N. Bansal, X. Chen, and Z. Wang · 2018
Cited alongside, same era.
J. Berner, P. Grohs, and A. Jentzen · 2018
Cited alongside, same era.
P. Petersen, M. Raslan, and F. Voigtlaender · 2018
Later among the works it cites.
Provable approximation properties for deep neural networks
U. Shaham, A. Cloninger, and R. R. Coifman · 2018
Later among the works it cites.
Learning transferable architectures for scalable image recognition
B. Zoph, V. Vasudevan, J. Shlens, and Q. V. Le · 2018
Later among the works it cites.
Towards a regularity theory for ReLU networks–chain rule and global error estimates
J. Berner, D. Elbrächter, P. Grohs, and A. Jentzen · 2019
Closest in time.
Approximation spaces of deep neural networks
R. Gribonval, G. Kutyniok, M. Nielsen, and F. Voigtlaender · 2019
Closest in time.
Error bounds for approximations with deep ReLU neural networks in W s , p {W^{s,p}} norms
I. Gühring, G. Kutyniok, and P. Petersen · 2019
Closest in time.
Chapter 15 - evolving deep neural networks
R. Miikkulainen, J. Liang, E. Meyerson, A. Rawal, D. Fink, O. Francon, B. Raju, H. Shahrzad, A. Navruzyan, N. Duffy, and B. Hodjat · 2019
Closest in time.