Fetching the paper…
Reading the bibliography…
Deep neural network is difficult to train and this predicament becomes worse as the depth increases.
Learning long-term dependencies with gradient descent is difficult
Y. Bengio, P. Simard, and P. Frasconi · 1994
Earlier work this paper cites.
Regularization theory and neural networks architectures
F. Girosi, M. Jones, and T. Poggio · 1995
Earlier work this paper cites.
Combinatorial Group Theory, Applications to Geometry
D. J. Collins, R. I. Grigorchuk, P. F. Kurchanov, P. M. Cohn, and G. (Mathematik) · 1998
Earlier work this paper cites.
Efficient backprop
Y. Lecun, L. Bottou, G. B. Orr, and K. R. M¨¹ller · 2000
Earlier work this paper cites.
Learning multiple layers of features from tiny images
A. Krizhevsky · 2009
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
X. Glorot and Y. Bengio · 2010
Earlier work this paper cites.
Rectified linear units improve restricted boltzmann machines
V. Nair and G. E. Hinton · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
J. Duchi, E. Hazan, and Y. Singer · 2011
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Earlier work this paper cites.
Rmsprop: Divide the gradient by a running average of its recent magnitude
T. Tieleman and G. Hinton · 2012
Earlier work this paper cites.
Adadelta: An adaptive learning rate method
M. Zeiler · 2012
Earlier work this paper cites.
Learning hierarchical category structure in deep neural networks
A. Saxe, J. McClelland, and S. Ganguli · 2013
Earlier work this paper cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
A. M. Saxe, J. L. McClelland, and S. Ganguli · 2013
Earlier work this paper cites.
On the importance of initialization and momentum in deep learning
I. Sutskever, J. Martens, G. Dahl, and G. Hinton · 2013
Earlier work this paper cites.
Linear algebra and matrix analysis for statistics
S. Banerjee and A. Roy · 2014
Cited alongside, same era.
Overfeat: Integrated recognition, localization and detection using convolutional networks
P. Sermanet, D. Eigen, X. Zhang, M. Mathieu, R. Fergus, and Y. Lecun · 2014
Cited alongside, same era.
Random walk initialization for training very deep feedforward networks
D. Sussillo and L. F. Abbott · 2014
Cited alongside, same era.
Visualizing and understanding convolutional networks
M. D. Zeiler and R. Fergus · 2014
Cited alongside, same era.
Convolutional neural networks at constrained time cost
K. He and J. Sun · 2015
Cited alongside, same era.
R. K. Srivastava, K. Greff, and J. Schmidhuber · 2015
Later among the works it cites.
R. K. Srivastava, K. Greff, and J. Schmidhuber · 2015
Later among the works it cites.
Going deeper with convolutions
C. Szegedy, W. Liu, Y. Jia, and P. Sermanet · 2015
Later among the works it cites.
Normalization propagation: A parametric technique for removing internal covariate shift in deep networks
D. Arpit, Y. Zhou, B. U. Kota, and V. Govindaraju · 2016
Later among the works it cites.
J. L. Ba, J. R. Kiros, and G. E. Hinton · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
K. He, X. Zhang, S. Ren, and J. Sun · 2015
Cited alongside, same era.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
K. He, X. Zhang, S. Ren, and J. Sun · 2015
Cited alongside, same era.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
S. Ioffe and C. Szegedy · 2015
Cited alongside, same era.
Adam: A method for stochastic optimization
D. Kingma and J. Ba · 2015
Cited alongside, same era.
Data-dependent initializations of convolutional neural networks
P. Krähenbühl, C. Doersch, J. Donahue, and T. Darrell · 2015
Cited alongside, same era.
Fully convolutional networks for semantic segmentation
J. Long, E. Shelhamer, and T. Darrell · 2015
Cited alongside, same era.
All you need is a good init
D. Mishkin and J. Matas · 2015
Cited alongside, same era.
L.-C. Chen, G. Papandreou, I. Kokkinos, K. Murphy, and A. L. Yuille · 2016
Later among the works it cites.
Optimization on submanifolds of convolution kernels in cnns
O. Mete and O. Takayuki · 2016
Later among the works it cites.
Learning to refine object segments
P. O. Pinheiro, T.-Y. Lin, R. Collobert, and P. Dollár · 2016
Later among the works it cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
S. Ren, K. He, R. Girshick, and J. Sun · 2016
Later among the works it cites.
Weight normalization: A simple reparameterization to accelerate training of deep neural networks
T. Salimans and D. P. Kingma · 2016
Later among the works it cites.
Understanding and improving convolutional neural networks via concatenated rectified linear units
W. Shang, K. Sohn, D. Almeida, and H. Lee · 2016
Later among the works it cites.
Training region-based object detectors with online hard example mining
A. Shrivastava, A. Gupta, and R. B. Girshick · 2016
Later among the works it cites.
Residual networks are exponential ensembles of relatively shallow networks
A. Veit, M. Wilber, and S. Belongie · 2016
Later among the works it cites.
B. Yang, J. Yan, Z. Lei, and S. Z. Li · 2016
Later among the works it cites.