H. Robbins and S. Monro, “A stochastic approximation method,” The Annals of Mathematical Statistics , 1951
1951
Earlier work this paper cites.
J. J. Dongarra et al. , “Top500 supercomputer sites,” 1994
1994
Earlier work this paper cites.
Y. LeCun and C. Cortes, “The MNIST database of handwritten digits,” 1998
1998
Earlier work this paper cites.
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner, “Gradient-based learning applied to document recognition,” Proc. IEEE , 1998
1998
Earlier work this paper cites.
J. S. Vetter and M. O. McCracken, “Statistical scalability analysis of communication operations in distributed applications,” in Proceedings of the Eighth ACM SIGPLAN Symposium on Principles and Practices of Parallel Programming , ser. PPoPP ’01. New York, NY, USA: ACM, 2001, pp. 123–132
2001
Earlier work this paper cites.
R. Collobert, S. Bengio, and J. Mariéthoz, “Torch: a modular machine learning software library,” Idiap, Tech. Rep., 2002
2002
Earlier work this paper cites.
K. Chellapilla, S. Puri, and P. Simard, “High Performance Convolutional Neural Networks for Document Processing,” in ICFHR , 2006
2006
Earlier work this paper cites.
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “ImageNet: A Large-Scale Hierarchical Image Database,” in CVPR , 2009
2009
Earlier work this paper cites.
A. Krizhevsky, “Learning multiple layers of features from tiny images,” Master’s thesis, 2009
2009
Earlier work this paper cites.
J. Bergstra, F. Bastien, O. Breuleux, P. Lamblin, R. Pascanu, O. Delalleau, G. Desjardins, D. Warde-Farley, I. Goodfellow, A. Bergeron et al. , “Theano: Deep learning on GPUs with Python,” in NIPS, BigLearning Workshop, Granada, Spain , 2011
2011
Earlier work this paper cites.
B. Recht, C. Re, S. Wright, and F. Niu, “Hogwild: A lock-free approach to parallelizing stochastic gradient descent,” in NIPS , 2011
2011
Earlier work this paper cites.
J. Dean, G. Corrado, R. Monga, K. Chen, M. Devin, M. Mao, A. Senior, P. Tucker, K. Yang, Q. V. Le et al. , “Large scale distributed deep networks,” in Advances in neural information processing systems , 2012
2012
Earlier work this paper cites.
A. Krizhevsky, I. Sutskever, and G. Hinton, “Imagenet classification with deep convolutional neural networks,” in Advances in Neural Information Processing Systems 25 , 2012, pp. 1097–1105
2012
Earlier work this paper cites.
Y. Jia, E. Shelhamer, J. Donahue, S. Karayev, J. Long, R. Girshick, S. Guadarrama, and T. Darrell, “Caffe: Convolutional architecture for fast feature embedding,” in ACMMM , 2014
2014
Earlier work this paper cites.
D. Yu, A. Eversole, M. Seltzer, K. Yao, Z. Huang, B. Guenter, O. Kuchaiev, Y. Zhang, F. Seide, H. Wang et al. , “An introduction to computational networks and the computational network toolkit,” Microsoft Technical Report MSR-TR-2014–112 , 2014
2014
Earlier work this paper cites.
T. Chen, M. Li, Y. Li, M. Lin, N. Wang, M. Wang, T. Xiao, B. Xu, C. Zhang, and Z. Zhang, “Mxnet: A flexible and efficient machine learning library for heterogeneous distributed systems,” arXiv:1512.01274 , 2015
Original
2015
Earlier work this paper cites.
S. Tokui, K. Oono, S. Hido, and J. Clayton, “Chainer: a next-generation open source framework for deep learning,” in LearningSys @NIPS , 2015
2015
Earlier work this paper cites.