Fetching the paper…
Reading the bibliography…
We propose a new layer design by adding a linear gating mechanism to shortcut connections.
Learning long-term dependencies with gradient descent is difficult
Y. Bengio, P. Simard, and P Frasconi · 1994
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann Lecun, Léon Bottou, Yoshua Bengio, and Patrick Haffner · 1998
Earlier work this paper cites.
Greedy layer-wise training of deep networks
Y. Bengio, P. Lamblin, D Popovici, and H Larochelle · 2007
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky · 2009
Earlier work this paper cites.
Exploring strategies for training deep neural networks
Hugo Larochelle, Yoshua Bengio, Jérôme Louradour, and Pascal Lamblin · 2009
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
X. Glorot and Y. Bengio · 2010
Earlier work this paper cites.
Rectified linear units improve restricted boltzmann machines
Vinod Nair and Geoffrey E. Hinton · 2010
Earlier work this paper cites.
Deep sparse rectifier neural networks
Xavier Glorot, Antoine Bordes, and Yoshua Bengio · 2011
Earlier work this paper cites.
On the complexity of neural network classifiers: A comparison between shallow and deep architectures
Monica Bianchini and Franco Scarselli · 2013
Cited alongside, same era.
Adam: A Method for Stochastic Optimization
D. Kingma and J. Ba · 2014
Cited alongside, same era.
On the Number of Linear Regions of Deep Neural Networks
G. Montúfar, R. Pascanu, K. Cho, and Y. Bengio · 2014
Cited alongside, same era.
Going deeper with convolutions
Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott E. Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich · 2014
Cited alongside, same era.
The Power of Depth for Feedforward Neural Networks
R. Eldan and O. Shamir · 2015
Cited alongside, same era.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Rupesh Kumar Srivastava, Klaus Greff, and Jürgen Schmidhuber · 2015
Later among the works it cites.
Identity Mappings in Deep Residual Networks
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Closest in time.
Deep Networks with Stochastic Depth
G. Huang, Y. Sun, Z. Liu, D. Sedra, and K. Weinberger · 2016
Closest in time.
FractalNet: Ultra-Deep Neural Networks without Residuals
G. Larsson, M. Maire, and G. Shakhnarovich · 2016
Closest in time.
Benefits of depth in neural networks
M. Telgarsky · 2016
Closest in time.
Revise Saturated Activation Functions
B. Xu, R. Huang, and M. Li · 2016
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
S. Ioffe and C. Szegedy · 2015
Cited alongside, same era.
ImageNet Large Scale Visual Recognition Challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li Fei-Fei · 2015
Cited alongside, same era.
Incorporating nesterov momentum into adam
Timothy Dozat
Cited in the paper.
Delving Deep into Rectifiers: Surpassing Human-Level Performance on ImageNet Classification
K. He, X. Zhang, S. Ren, and J. Sun
Cited in the paper.
Deep Residual Learning for Image Recognition
K. He, X. Zhang, S. Ren, and J. Sun
Cited in the paper.
Sergey Zagoruyko and Nikos Komodakis · 2016
Closest in time.