Fetching the paper…
Reading the bibliography…
Augmenting neural networks with skip connections, as introduced in the so-called ResNet architecture, surprised the community by enabling the training of networks of more than 1,000 layers with significant performance gains.
N. J. Higham, “Stable iterations for the matrix square root,” Numerical Algorithms
1997
Earlier work this paper cites.
Oxford University Press on Demand, 2004
J. C. Gower, G. B. Dijksterhuis, and others, Procrustes problems · 2004
Earlier work this paper cites.
A. Krizhevsky and G. Hinton, “Learning multiple layers of features from tiny images.(2009),” tech. rep., 2009
2009
Earlier work this paper cites.
X. Glorot and Y. Bengio, “Understanding the difficulty of training deep feedforward neural networks,” 3 2010
2010
Earlier work this paper cites.
G. F. Montufar, R. Pascanu, K. Cho, and Y. Bengio, “On the Number of Linear Regions of Deep Neural Networks,” in Advances in Neural Information Processing Systems 27
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Delving Deep into Rectifiers: Surpassing Human-Level Performance on ImageNet Classification,” in 2015 IEEE International Conference on Computer Vision (ICCV)
2015
Earlier work this paper cites.
S. Ioffe and C. Szegedy, “Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift,” 6 2015
2015
Earlier work this paper cites.
R. K. Srivastava, K. Greff, and J. Schmidhuber, “Training Very Deep Networks,” 2015
2015
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep Residual Learning for Image Recognition,” in 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)
2016
Cited alongside, same era.
K. He, X. Zhang, S. Ren, and J. Sun, “Identity mappings in deep residual networks,” in Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)
2016
Cited alongside, same era.
K. Kawaguchi, “Deep Learning without Poor Local Minima,” 2016
2016
Cited alongside, same era.
A. Veit, M. J. Wilber, and S. Belongie, “Residual Networks Behave Like Ensembles of Relatively Shallow Networks,” in Advances in Neural Information Processing Systems 29
2016
Cited alongside, same era.
D. Silver, J. Schrittwieser, K. Simonyan, I. Antonoglou, A. Huang, A. Guez, T. Hubert, L. Baker, M. Lai, A. Bolton, Y. Chen, T. Lillicrap, F. Hui, L. Sifre, G. van den Driessche, T. Graepel, and D. Hassabis, “Mastering the game of Go without human knowledge,” Nature
A. N. Gomez, M. Ren, R. Urtasun, and R. B. Grosse, “The reversible residual network: Backpropagation without storing activations,” in Advances in Neural Information Processing Systems
2017
Later among the works it cites.
E. Hoffer, I. Hubara, and D. Soudry, “Train longer, generalize better: Closing the generalization gap in large batch training of neural networks,” in Advances in Neural Information Processing Systems
2017
Later among the works it cites.
2017
Later among the works it cites.
A. E. Orhan and X. Pitkow, “Skip connections eliminate singularities,” in 6th International Conference on Learning Representations, ICLR 2018 - Conference Track Proceedings
2018
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2017
Cited alongside, same era.
G. Huang, Z. Liu, L. v. d. Maaten, and K. Q. Weinberger, “Densely Connected Convolutional Networks,” in 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)
2017
Cited alongside, same era.
M. Hardt and T. Ma, “Identity matters in deep learning,” in 5th International Conference on Learning Representations, ICLR 2017 - Conference Track Proceedings
2017
Cited alongside, same era.
D. Balduzzi, M. Frean, L. Leary, J. P. Lewis, K. W.-D. Ma, and B. McWilliams, “The Shattered Gradients Problem: If resnets are the answer, then what is the question?,” 7 2017
2017
Cited alongside, same era.
L. Dinh, J. Sohl-Dickstein Google, B. Samy, and B. Google Brain, “Density estimation using Real NVP,” in ICLR
2017
Cited alongside, same era.
2018
Closest in time.
S. Gunasekar, J. D. Lee, N. Srebro, and D. Soudry, “Implicit bias of gradient descent on linear convolutional networks,” in Advances in Neural Information Processing Systems
2018
Closest in time.
J. Behrmann, W. Grathwohl, R. T. Chen, D. Duvenaud, and J. H. Jacobsen, “Invertible residual networks,” in 36th International Conference on Machine Learning, ICML 2019
2019
Closest in time.
K. Kawaguchi and Y. Bengio, “Depth with nonlinearity creates no bad local minima in ResNets,” Neural Networks
2019
Closest in time.
H. Sedghi, V. Gupta, and P. M. Long, “The Singular Values of Convolutional Layers,” in International Conference on Learning Representations
2019
Closest in time.