Fetching the paper…
Reading the bibliography…
Recent theoretical work has demonstrated that deep neural networks have superior performance over shallow networks, but their training is more difficult, e.g., they suffer from the vanishing gradient problem.
Approximation by superpositions of a sigmoidal function
G. Cybenko · 1989
Earlier work this paper cites.
Multilayer feedforward networks are universal approximators
K. Hornik, M. Stinchcombe, and H. White · 1989
Earlier work this paper cites.
Universal approximation bounds for superpositions of a sigmoidal function
A. R. Barron · 1993
Earlier work this paper cites.
A limited memory algorithm for bound constrained optimization
R. H. Byrd, P. Lu, J. Nocedal, and C. Zhu · 1995
Earlier work this paper cites.
Neural networks for optimal approximation of smooth and analytic functions
H. Mhaskar · 1996
Earlier work this paper cites.
Efficient backprop
Y. LeCun, L. Bottou, G. B. Orr, and K.-R. Müller · 1998
Earlier work this paper cites.
Local minima and plateaus in hierarchical structures of multilayer perceptrons
K. Fukumizu and S. Amari · 2000
Earlier work this paper cites.
Training a single sigmoidal neuron is hard
J. Šíma · 2002
Earlier work this paper cites.
Singularities affect dynamics of learning in neuromanifolds
S. Amari, H. Park, and T. Ozeki · 2006
Earlier work this paper cites.
Numerical Optimization
J. Nocedal and S. J. Wright · 2006
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
X. Glorot and Y. Bengio · 2010
Earlier work this paper cites.
Shallow vs. deep sum-product networks
O. Delalleau and Y. Bengio · 2011
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
J. Duchi, E. Hazan, and Y. Singer · 2011
Earlier work this paper cites.
Deep sparse rectifier neural networks
X. Glorot, A. Bordes, and Y. Bengio · 2011
Earlier work this paper cites.
Adadelta: an adaptive learning rate method
M. D. Zeiler · 2012
Earlier work this paper cites.
Rectifier nonlinearities improve neural network acoustic models
A. L. Maas, A. Y. Hannun, and A. Y. Ng · 2013
Earlier work this paper cites.
On the importance of initialization and momentum in deep learning
I. Sutskever, J. Martens, G. Dahl, and G. Hinton · 2013
Earlier work this paper cites.
Overview of mini-batch gradient descent
G. Hinton · 2014
Cited alongside, same era.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
A. M. Saxe, J. L. McClelland, and S. Ganguli · 2014
Cited alongside, same era.
Dropout: a simple way to prevent neural networks from overfitting
N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov · 2014
Cited alongside, same era.
Fast and accurate deep network learning by exponential linear units (elus)
D.-A. Clevert, T. Unterthiner, and S. Hochreiter · 2015
Cited alongside, same era.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
K. He, X. Zhang, S. Ren, and J. Sun · 2015
Cited alongside, same era.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Gradient descent learns one-hidden-layer cnn: Don’t be afraid of spurious local minima
S. Du, J. Lee, Y. Tian, B. Poczos, and A. Singh · 2017
Later among the works it cites.
Approximating continuous functions by relu nets of minimal width
B. Hanin and M. Sellke · 2017
Later among the works it cites.
Self-normalizing neural networks
G. Klambauer, T. Unterthiner, A. Mayr, and S. Hochreiter · 2017
Later among the works it cites.
Why deep neural networks for function approximation?
S. Liang and R. Srikant · 2017
Later among the works it cites.
The expressive power of neural networks: A view from the width
Z. Lu, H. Pu, F. Wang, Z. Hu, and L. Wang · 2017
Later among the works it cites.
When and why are deep networks better than shallow ones?
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
S. Ioffe and C. Szegedy · 2015
Cited alongside, same era.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2015
Cited alongside, same era.
Training very deep networks
R. K. Srivastava, K. Greff, and J. Schmidhuber · 2015
Cited alongside, same era.
J. L. Ba, J. R. Kiros, and G. E. Hinton · 2016
Cited alongside, same era.
The power of depth for feedforward neural networks
R. Eldan and O. Shamir · 2016
Cited alongside, same era.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Cited alongside, same era.
Deep learning without poor local minima
K. Kawaguchi · 2016
Cited alongside, same era.
H. Mhaskar, Q. Liao, and T. A. Poggio · 2017
Later among the works it cites.
Why and when can deep-but not shallow-networks avoid the curse of dimensionality: a review
T. Poggio, H. Mhaskar, L. Rosasco, B. Miranda, and Q. Liao · 2017
Later among the works it cites.
Spurious local minima are common in two-layer relu neural networks
I. Safran and O. Shamir · 2017
Later among the works it cites.
Error bounds for approximations with deep relu networks
D. Yarotsky · 2017
Later among the works it cites.
Critical points of neural networks: Analytical forms and landscape properties
Y. Zhou and Y. Liang · 2017
Later among the works it cites.
How to start training: The effect of initialization and architecture
B. Hanin and D. Rolnick · 2018
Closest in time.
Neural networks should be wide enough to learn disconnected decision regions
Q. Nguyen, M. Mukkamala, and M. Hein · 2018
Closest in time.
Optimal approximation of piecewise smooth functions using deep relu neural networks
P. Petersen and F. Voigtlaender · 2018
Closest in time.
No spurious local minima in a two hidden unit relu network
C. Wu, J. Luo, and J. Lee · 2018
Closest in time.
Y. Wu and K. He · 2018
Closest in time.
Small nonlinearities in activation functions create bad local minima in neural networks
C. Yun, S. Sra, and Jadbabaie A · 2018
Closest in time.
On the effect of the activation function on the distribution of hidden nodes in a deep network, 2019
P. M. Long and H. Sedghi · 2019
Closest in time.