Fetching the paper…
Reading the bibliography…
Skip connections made the training of very deep networks possible and have become an indispensable component in a variety of neural architectures.
The eigenvalues of mega-dimensional matrices
J. Skilling · 1989
Earlier work this paper cites.
Untersuchungen zu dynamischen neuronalen Netzen
S. Hochreiter · 1991
Earlier work this paper cites.
Learning long-term dependencies with gradient descent is difficult
Y. Bengio, P. Simard, and P. Frasconi · 1994
Earlier work this paper cites.
Fast exact multiplication by the hessian
B.A. Pearlmutter · 1994
Earlier work this paper cites.
On-line learning in soft committee machines
D. Saad and S.A. Solla · 1995
Earlier work this paper cites.
Singularities affect dynamics of learning in neuromanifolds
S. Amari, H. Park, and T. Ozeki · 2006
Earlier work this paper cites.
Dynamics of learning near singularities in layered networks
H. Wei, J. Zhang, F. Cousseau, T. Ozeki, and S. Amari · 2008
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
X. Glorot and Y. Bengio · 2010
Earlier work this paper cites.
Random search for hyper-parameter optimization
J. Bergstra and Y. Bengio · 2012
Cited alongside, same era.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
A.M. Saxe, J.M. McClelland, and S. Ganguli · 2013
Cited alongside, same era.
Adam: a method for stochastic optimization
D.P. Kingma and J.L. Ba · 2014
Cited alongside, same era.
Reducing overfitting in deep networks by decorrelating representations
M. Cogswell, F. Ahmed, R. Girshick, L. Zitnick, and D. Batra · 2015
Cited alongside, same era.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2015
Cited alongside, same era.
Identity matters in deep learning
M. Hardt and T. Ma · 2016
Later among the works it cites.
Identity mappings in deep residual networks
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Later among the works it cites.
Densely connected convolutional networks
G. Huang, Z. Liu, K.Q. Weinberger, and L. van der Maaten · 2016
Later among the works it cites.
S. Li, J. Jiao, Y. Han, and T. Weissman · 2016
Later among the works it cites.
Learning deep parsimonious representations
R. Liao, A.G. Schwing, R.S. Zemel, and R. Urtasun · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Batch normalization: Accelerating deep network training by reducing internal covariate shift
S. Ioffe and C. Szegedy · 2015
Cited alongside, same era.
Training very deep networks
R.K. Srivastava, K. Greff, and J. Schmidhuber · 2015
Cited alongside, same era.
Efficient approaches for escaping higher order saddle points in non-convex optimization
A. Anandkumar and R. Ge · 2016
Cited alongside, same era.
Aggregated residual transformations for deep neural networks
S. Xie, R. Girshick, P. Dollar, Z. Tu, and K. He · 2016
Later among the works it cites.
The shattered gradients problem: If resnets are the answer, then what is the question?
D. Balduzzi, M. Frean, L. Leary, J.P. Lewis, K.W.-D. Ma, and B. McWilliams · 2017
Closest in time.