Fetching the paper…
Reading the bibliography…
The dying ReLU refers to the problem when ReLU neurons become inactive and only output 0 for any input.
A limited memory algorithm for bound constrained optimization
R. H. Byrd, P. Lu, J. Nocedal, and C. Zhu · 1995
Earlier work this paper cites.
Efficient backprop
Y. LeCun, L. Bottou, G. B. Orr, and K.-R. Müller · 1998
Earlier work this paper cites.
Local minima and plateaus in hierarchical structures of multilayer perceptrons
K. Fukumizu and S. Amari · 2000
Earlier work this paper cites.
Singularities affect dynamics of learning in neuromanifolds
S. Amari, H. Park, and T. Ozeki · 2006
Earlier work this paper cites.
Numerical Optimization
J. Nocedal and S. J. Wright · 2006
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
X. Glorot and Y. Bengio · 2010
Earlier work this paper cites.
Rectified linear units improve restricted boltzmann machines
V. Nair and G. E. Hinton · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
J. Duchi, E. Hazan, and Y. Singer · 2011
Earlier work this paper cites.
Deep sparse rectifier neural networks
X. Glorot, A. Bordes, and Y. Bengio · 2011
Earlier work this paper cites.
Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups
G. Hinton, L. Deng, D. Yu, G. E. Dahl, A. Mohamed, N. Jaitly, A. Senior, V. Vanhoucke, P. Nguyen, T. N. Sainath, et al · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Earlier work this paper cites.
Adadelta: an adaptive learning rate method
M. D. Zeiler · 2012
Earlier work this paper cites.
Rectifier nonlinearities improve neural network acoustic models
A. L. Maas, A. Y. Hannun, and A. Y. Ng · 2013
Earlier work this paper cites.
Introductory lectures on convex optimization: A basic course
Y. Nesterov · 2013
Earlier work this paper cites.
On the importance of initialization and momentum in deep learning
I. Sutskever, J. Martens, G. Dahl, and G. Hinton · 2013
Earlier work this paper cites.
Overview of mini-batch gradient descent
G. Hinton · 2014
Earlier work this paper cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
A. M. Saxe, J. L. McClelland, and S. Ganguli · 2014
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov · 2014
Earlier work this paper cites.
Fast and accurate deep network learning by exponential linear units (elus)
D. A. Clevert, T. Unterthiner, and S. Hochreiter · 2015
Earlier work this paper cites.
Escaping from saddle points—online stochastic gradient for tensor decomposition
R. Ge, F. Huang, C. Jin, and Y. Yuan · 2015
Cited alongside, same era.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
K. He, X. Zhang, S. Ren, and J. Sun · 2015
Cited alongside, same era.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
S. Ioffe and C. Szegedy · 2015
Cited alongside, same era.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2015
Cited alongside, same era.
Deep learning
Y. LeCun, Y. Bengio, and G. Hinton · 2015
Cited alongside, same era.
Training very deep networks
R. K. Srivastava, K. Greff, and J. Schmidhuber · 2015
Cited alongside, same era.
Approximating continuous functions by relu nets of minimal width
B. Hanin and M. Sellke · 2017
Later among the works it cites.
How to escape saddle points efficiently
C. Jin, R. Ge, P. Netrapalli, S. M. Kakade, and M. I. Jordan · 2017
Later among the works it cites.
Self-normalizing neural networks
G. Klambauer, T. Unterthiner, A. Mayr, and S. Hochreiter · 2017
Later among the works it cites.
Searching for activation functions
P. Ramachandran, B. Zoph, and Q. V. Le · 2017
Later among the works it cites.
Parametric exponential linear unit for deep convolutional neural networks
L. Trottier, P. Gigu, B. Chaib-draa, et al · 2017
Later among the works it cites.
Error bounds for approximations with deep relu networks
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Deeply learned face representations are sparse, selective, and robust
Y. Sun, X. Wang, and X. Tang · 2015
Cited alongside, same era.
Going deeper with convolutions
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich · 2015
Cited alongside, same era.
J. L. Ba, J. R. Kiros, and G. E. Hinton · 2016
Cited alongside, same era.
Matrix completion has no spurious local minimum
R. Ge, J. D. Lee, and T. Ma · 2016
Cited alongside, same era.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Cited alongside, same era.
Deep learning without poor local minima
K. Kawaguchi · 2016
Cited alongside, same era.
D. Yarotsky · 2017
Later among the works it cites.
Critical points of neural networks: Analytical forms and landscape properties
Y. Zhou and Y. Liang · 2017
Later among the works it cites.
Deep learning using rectified linear units (relu)
A. F. Agarap · 2018
Later among the works it cites.
A convergence analysis of gradient descent for deep linear neural networks
S. Arora, N. Cohen, N. Golowich, and W. Hu · 2018
Later among the works it cites.
Dynamical isometry and a mean field theory of rnns: Gating enables signal propagation in recurrent neural networks
M. Chen, J. Pennington, and S. Schoenholz · 2018
Later among the works it cites.
Gradient descent finds global minima of deep neural networks
S. S. Du, J. D. Lee, H. Li, L. Wang, and X. Zhai · 2018
Later among the works it cites.
Gradient descent learns one-hidden-layer cnn: Don’t be afraid of spurious local minima
S. S Du, J. D. Lee, Y. Tian, A Singh, and B Poczos · 2018
Later among the works it cites.
Which neural net architectures give rise to exploding and vanishing gradients?
B. Hanin · 2018
Later among the works it cites.
Optimal approximation of piecewise smooth functions using deep relu neural networks
P. Petersen and F. Voigtlaender · 2018
Later among the works it cites.
Spurious local minima are common in two-layer relu neural networks
I. Safran and O. Shamir · 2018
Later among the works it cites.
No spurious local minima in a two hidden unit relu network
C. Wu, J. Luo, and J. Lee · 2018
Later among the works it cites.
Group normalization
Y. Wu and K. He · 2018
Later among the works it cites.
Small nonlinearities in activation functions create bad local minima in neural networks
C. Yun, S. Sra, and Jadbabaie A · 2018
Later among the works it cites.
ReLU deep neural networks and linear finite elements
J. He, L. Li, J. Xu, and C. Zheng · 2020
Closest in time.