Fetching the paper…
Reading the bibliography…
Deep residual networks (ResNets) have significantly pushed forward the state-of-the-art on image classification, increasing in performance as networks grow both deeper and wider.
Automatic differentiation: Techniques and applications
L. B. Rall · 1981
Earlier work this paper cites.
Learning representations by back-propagating errors
D. Rumelhart, G. Hinton, and R. Williams · 1986
Earlier work this paper cites.
A learning algorithm for continually running fully recurrent neural networks
R. J. Williams and D. Zipser · 1989
Earlier work this paper cites.
Handwritten digit recognition with a back-propagation network
Y. LeCun, B. E. Boser, J. S. Denker, D. Henderson, R. E. Howard, W. E. Hubbard, and L. D. Jackel · 1990
Earlier work this paper cites.
Learning long-term dependencies with gradient descent is difficult
Y. Bengio, P. Simard, and P. Frasconi · 1994
Earlier work this paper cites.
Learning multiple layers of features from tiny images
A. Krizhevsky and G. Hinton · 2009
Earlier work this paper cites.
Rectified linear units improve restricted Boltzmann machines
V. Nair and G. E. Hinton · 2010
Earlier work this paper cites.
Large scale distributed deep networks
J. Dean, G. Corrado, R. Monga, K. Chen, M. Devin, M. Mao, A. Senior, P. Tucker, K. Yang, Q. V. Le, et al · 2012
Earlier work this paper cites.
ImageNet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Earlier work this paper cites.
Training deep and recurrent networks with Hessian-free optimization
J. Martens and I. Sutskever · 2012
Earlier work this paper cites.
Microsoft COCO: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Cited alongside, same era.
Rigid-motion scattering for image classification
L. Sifre · 2014
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2014
Cited alongside, same era.
NICE: Non-linear independent components estimation
L. Dinh, D. Krueger, and Y. Bengio · 2015
Cited alongside, same era.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
S. Ioffe and C. Szegedy · 2015
Cited alongside, same era.
Gradient-based hyperparameter optimization through reversible learning
D. Maclaurin, D. K. Duvenaud, and R. P. Adams · 2015
Theano: A Python framework for fast computation of mathematical expressions
R. Al-Rfou, G. Alain, A. Almahairi, C. Angermueller, D. Bahdanau, N. Ballas, F. Bastien, J. Bayer, A. Belikov, A. Belopolsky, et al · 2016
Later among the works it cites.
Xception: Deep learning with depthwise separable convolutions
F. Chollet · 2016
Later among the works it cites.
Memory-efficient backpropagation through time
A. Gruslys, R. Munos, I. Danihelka, M. Lanctot, and A. Graves · 2016
Later among the works it cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Later among the works it cites.
Decoupled neural interfaces using synthetic gradients
M. Jaderberg, W. M. Czarnecki, S. Osindero, O. Vinyals, A. Graves, and K. Kavukcuoglu · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
ImageNet large scale visual recognition challenge
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, et al · 2015
Cited alongside, same era.
R. K. Srivastava, K. Greff, and J. Schmidhuber · 2015
Cited alongside, same era.
Going deeper with convolutions
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. E. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich · 2015
Cited alongside, same era.
TensorFlow: Large-scale machine learning on heterogeneous distributed systems
M. Abadi, A. Agarwal, P. Barham, E. Brevdo, Z. Chen, C. Citro, G. S. Corrado, A. Davis, J. Dean, M. Devin, et al · 2016
Cited alongside, same era.
Revisiting distributed synchronous sgd
J. Chen, R. Monga, S. Bengio, and R. Jozefowicz
Cited in the paper.
Training deep nets with sublinear memory cost
T. Chen, B. Xu, C. Zhang, and C. Guestrin
Cited in the paper.
Neural machine translation in linear time
N. Kalchbrenner, L. Espeholt, K. Simonyan, A. v. d. Oord, A. Graves, and K. Kavukcuoglu · 2016
Later among the works it cites.
Density estimation using real NVP
L. Dinh, J. Sohl-Dickstein, and S. Bengio · 2017
Closest in time.
Pyramid scene parsing network
H. Zhao, J. Shi, X. Qi, X. Wang, and J. Jia · 2017
Closest in time.
Unpaired image-to-image translation using cycle-consistent adversarial networks
J.-Y. Zhu, T. Park, P. Isola, and A. A. Efros · 2017
Closest in time.