Fetching the paper…
Reading the bibliography…
In this paper, we study the trainability of rectified linear unit (ReLU) networks.
A stochastic approximation method
H. Robbins and S. Monro · 1951
Earlier work this paper cites.
Learning internal representations by error propagation
D. E. Rumelhart, G. E. Hinton, and R. J. Williams · 1985
Earlier work this paper cites.
Approximation by superpositions of a sigmoidal function
G. Cybenko · 1989
Earlier work this paper cites.
Approximation capabilities of multilayer feedforward networks
K. Hornik · 1991
Earlier work this paper cites.
A limited memory algorithm for bound constrained optimization
R. H. Byrd, P. Lu, J. Nocedal, and C. Zhu · 1995
Earlier work this paper cites.
Efficient backprop
Y. LeCun, L. Bottou, G. B. Orr, and K. R. Müller · 1998
Earlier work this paper cites.
Distributing points on the sphere: partitions, separation, quadrature and energy
P. C. Leopardi · 2007
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
X. Glorot and Y. Bengio · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
J. Duchi, E. Hazan, and Y. Singer · 2011
Earlier work this paper cites.
Deep neural networks for acoustic modeling in speech recognition
G. Hinton, L. Deng, D. Yu, G. Dahl, A. Mohamed, N. Jaitly, A. Senior, V. Vanhoucke, P. Nguyen, B. Kingsbury, et al · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. Hinton · 2012
Earlier work this paper cites.
Overview of mini-batch gradient descent
G. Hinton · 2014
Cited alongside, same era.
On the computational efficiency of training neural networks
R. Livni, S. Shalev-Shwartz, and O. Shamir · 2014
Cited alongside, same era.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
A. M. Saxe, J. L. McClelland, and S. Ganguli · 2014
Cited alongside, same era.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
K. He, X. Zhang, S. Ren, and J. Sun · 2015
Cited alongside, same era.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
S. Ioffe and C. Szegedy · 2015
Cited alongside, same era.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2015
Cited alongside, same era.
Mastering the game of go with deep neural networks and tree search
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. Van Den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, et al · 2016
Later among the works it cites.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Y. Wu, M. Schuster, Z. Chen, Q. V. Le, M. Norouzi, W. Macherey, M. Krikun, Y. Cao, Q. Gao, K. Macherey, et al · 2016
Later among the works it cites.
The loss surface of deep and wide neural networks
Q. Nguyen and M. Hein · 2017
Later among the works it cites.
A convergence theory for deep learning via over-parameterization
Z. Allen-Zhu, Y. Li, and Z. Song · 2018
Later among the works it cites.
Learning overparameterized neural networks via stochastic gradient descent on structured data
Y. Li and Y. Liang · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Data-dependent initializations of convolutional neural networks
P. Krähenbühl, C. Doersch, J. Donahue, and T. Darrell · 2015
Cited alongside, same era.
All you need is a good init
D. Mishkin and J. Matas · 2016
Cited alongside, same era.
On the quality of the initial basin in overspecified neural networks
I. Safran and O. Shamir · 2016
Cited alongside, same era.
Weight normalization: A simple reparameterization to accelerate training of deep neural networks
T. Salimans and D. P. Kingma · 2016
Cited alongside, same era.
Gradient descent finds global minima of deep neural networks
S. S. Du, J. D. Lee, H. Li, L. Wang, and X. Zhai
Cited in the paper.
Gradient descent provably optimizes over-parameterized neural networks
S. S. Du, X. Zhai, B. Poczos, and A. Singh
Cited in the paper.
Stochastic gradient descent optimizes over-parameterized deep relu networks
D. Zou, Y. Cao, D. Zhou, and Q. Gu · 2018
Later among the works it cites.
Dying ReLU and initialization: Theory and numerical examples
L. Lu, Y. Shin, Y. Su, and G. E. Karniadakis · 2019
Closest in time.
S. Oymak and M. Soltanolkotabi · 2019
Closest in time.
On the convergence of adam and beyond
S. J. Reddi, S. Kale, and S. Kumar · 2019
Closest in time.
Theoretical insights into the optimization landscape of over-parameterized shallow neural networks
M. Soltanolkotabi, A. Javanmard, and J. D. Lee · 2019
Closest in time.