Fetching the paper…
Reading the bibliography…
The ability to train large-scale neural networks has resulted in state-of-the-art performance in many areas of computer vision.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Large-scale deep unsupervised learning using graphics processors
R. Raina, A. Madhavan, and A. Y. Ng · 2009
Earlier work this paper cites.
Theano: a cpu and gpu math expression compiler
J. Bergstra, O. Breuleux, F. Bastien, P. Lamblin, R. Pascanu, G. Desjardins, J. Turian, D. Warde-Farley, and Y. Bengio · 2010
Earlier work this paper cites.
Rectified linear units improve restricted boltzmann machines
V. Nair and G. E. Hinton · 2010
Earlier work this paper cites.
Parallelized stochastic gradient descent
M. Zinkevich, M. Weimer, L. Li, and A. J. Smola · 2010
Earlier work this paper cites.
Torch7: A matlab-like environment for machine learning
R. Collobert, K. Kavukcuoglu, and C. Farabet · 2011
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
J. Duchi, E. Hazan, and Y. Singer · 2011
Earlier work this paper cites.
Building high-level features using large scale unsupervised learning
Q. V. Le, M. Ranzato, R. Monga, M. Devin, K. Chen, G. S. Corrado, J. Dean, and A. Y. Ng · 2011
Cited alongside, same era.
Hogwild!: A lock-free approach to parallelizing stochastic gradient descent
F. Niu, B. Recht, C. Ré, and S. J. Wright · 2011
Cited alongside, same era.
Distributed delayed stochastic optimization
A. Agarwal and J. C. Duchi · 2012
Cited alongside, same era.
Multi-column deep neural networks for image classification
D. Ciresan, U. Meier, and J. Schmidhuber · 2012
Cited alongside, same era.
Large scale distributed deep networks
J. Dean, G. Corrado, R. Monga, K. Chen, M. Devin, Q. Le, M. Mao, M. Ranzato, A. Senior, P. Tucker, K. Yang, and A. Ng · 2012
Cited alongside, same era.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. Hinton · 2012
Later among the works it cites.
Adadelta: An adaptive learning rate method
M. D. Zeiler · 2012
Later among the works it cites.
Deep learning with cots hpc systems
A. Coates, B. Huval, T. Wang, D. Wu, B. Catanzaro, and N. Andrew · 2013
Closest in time.
Decaf: A deep convolutional activation feature for generic visual recognition
J. Donahue, Y. Jia, O. Vinyals, J. Hoffman, N. Zhang, E. Tzeng, and T. Darrell · 2013
Closest in time.
Rich feature hierarchies for accurate object detection and semantic segmentation
R. Girshick, J. Donahue, T. Darrell, and J. Malik · 2013
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
G. E. Hinton, N. Srivastava, A. Krizhevsky, I. Sutskever, and R. R. Salakhutdinov · 2012
Cited alongside, same era.
Visualizing and Understanding Convolutional Networks
M. D. Zeiler and R. Fergus · 2013
Closest in time.