Fetching the paper…
Reading the bibliography…
A residual-networks family with hundreds or even thousands of layers dominates major image recognition tasks, but building a network by simply stacking residual blocks inevitably limits its optimization ability.
2006
Earlier work this paper cites.
A. Krizhenvshky, and G. Hinton, “Learning multiple layers of features from tiny images,” M.Sc. thesis, Dept. of Comput. Sci., Univ. of Toronto, Toronto, ON, Canada, 2009
2009
Earlier work this paper cites.
X. Glorot, and Y. Bengio, “Understanding the difficulty of training deep feedforward neural networks,” in Proc. Conf. Art. Intell. Stat. , 2010, pp. 249–256
2010
Earlier work this paper cites.
D. Erhan, Y. Bengio, A. Courville, P. A. Manzagol, P. Vincent, and S. Bengio, “Why does unsupervised pre-training help deep learning?,” The Journal of Machine Learning Research , vol. 11, pp. 625–660, Mar. 2010
2010
Earlier work this paper cites.
V. Nair, and G. Hinton, “Rectified linear units improve restricted Boltzmann machines,” in Proc. ICML , 2010, pp. 807–814
2010
Earlier work this paper cites.
Y. Netzer, T. Wang, A. Coates, A. Bissacco, B. Wu, and A. Y. Ng, “Reading digits in natural images with unsupervised feature learning,” in Proc. NIPS Workshop Deep Learning and Unsupervised feature learning. , 2011, pp. 1–9
2011
Earlier work this paper cites.
A. Krizhenvshky, I. Sutskever, and G. Hinton, “Imagenet classification with deep convolutional networks,” in Proc. Adv. Neural Inf. Process. Syst. , 2012, pp. 1097–1105
2012
Earlier work this paper cites.
2012
Earlier work this paper cites.
M. Lin, Q. Chen, and S. Yan, “Network in network,” arXiv preprint arXiv:1312.4400 , 2013
2013
Earlier work this paper cites.
2013
Earlier work this paper cites.
L. Wan, M. Zeiler, S. Zhang, Y. L. Cun, and R. Fergus, “Regularization of neural networks using dropconnect,” in Proc. ICML. , 2013, pp. 1058–1066
2013
Earlier work this paper cites.
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
2014
Cited alongside, same era.
2014
Cited alongside, same era.
2014
Cited alongside, same era.
Y. Bengio, P. Simard, and P. Frasconi, “Learning long-term dependencies with gradient descent is difficult,” IEEE Trans. Neural Networks , vol. 5, no. 2, pp. 157–166, Aug. 2014
2014
Cited alongside, same era.
K. He, and J. Sun, “Convolutional neural networks at constrained time cost,” in Proc. IEEE Conf. Comput. Vis. Pattern Recog. , 2015, pp. 5353–5360
2015
Later among the works it cites.
2015
Later among the works it cites.
2016
Closest in time.
S. Zagoruyko, and N. Komodakis, “Wide residual networks,” arXiv preprint arXiv:1605.07146 , 2016
2016
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov, “Dropout: a simple way to prevent neural networks from overfitting,” The Journal of Machine Learning Research , vol. 15, pp. 1929–1958, Jun. 2014
2014
Cited alongside, same era.
Y. LeCun, Y. Bengio, and G. Hinton, “Deep learning,” Nature , vol. 521, no. 7553, pp. 436–444, May. 2015
2015
Cited alongside, same era.
C. -Y. Lee, S. Xie, P. Gallagher, Z. Zhang, and Z. Tu, “Deeply-supervised nets,” in Proc. AISTATS , 2015, pp. 562–570
2015
Cited alongside, same era.
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich, “Going deeper with convolutions,” in Proc. IEEE Conf. Comput. Vis. Pattern Recog. , 2015, pp. 1–9
2015
Cited alongside, same era.
2015
Cited alongside, same era.
2015
Cited alongside, same era.
2015
Cited alongside, same era.
2015
Cited alongside, same era.
2016
Closest in time.
2016
Closest in time.
2016
Closest in time.
2016
Closest in time.
2016
Closest in time.
2016
Closest in time.
2016
Closest in time.