A. Petrowski, G. Dreyfus, and C. Girault, “Performance analysis of a pipelined backpropagation parallel algorithm,” IEEE Transactions on Neural Networks , vol. 4, no. 6, pp. 970–981, 1993
1993
Earlier work this paper cites.
N. Qian, “On the momentum term in gradient descent learning algorithms,” Neural networks , vol. 12, no. 1, pp. 145–151, 1999
1999
Earlier work this paper cites.
H. Lee, P. Pham, Y. Largman, and A. Y. Ng, “Unsupervised feature learning for audio classification using convolutional deep belief networks,” in Advances in neural information processing systems , 2009, pp. 1096–1104
2009
Earlier work this paper cites.
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” in 2009 IEEE conference on computer vision and pattern recognition . Ieee, 2009, pp. 248–255
2009
Earlier work this paper cites.
A. Krizhevsky and G. Hinton, “Learning multiple layers of features from tiny images,” Citeseer, Tech. Rep., 2009
2009
Earlier work this paper cites.
J. Dean, G. Corrado, R. Monga, K. Chen, M. Devin, M. Mao, M. Ranzato, A. Senior, P. Tucker, K. Yang et al. , “Large scale distributed deep networks,” in Advances in neural information processing systems , 2012, pp. 1223–1231
2012
Earlier work this paper cites.
X. Chen, A. Eversole, G. Li, D. Yu, and F. Seide, “Pipelined back-propagation for context-dependent deep neural networks,” in Thirteenth Annual Conference of the International Speech Communication Association , 2012
2012
Earlier work this paper cites.
T. Tieleman and G. Hinton, “Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude,” COURSERA: Neural networks for machine learning , vol. 4, no. 2, pp. 26–31, 2012
2012
Earlier work this paper cites.
M. Kamruzzaman, S. Swanson, and D. M. Tullsen, “Load-balanced pipeline parallelism,” in SC’13: Proceedings of the International Conference on High Performance Computing, Networking, Storage and Analysis . IEEE, 2013, pp. 1–12
2013
Earlier work this paper cites.
A. Karpathy, G. Toderici, S. Shetty, T. Leung, R. Sukthankar, and L. Fei-Fei, “Large-scale video classification with convolutional neural networks,” in Proceedings of the IEEE conference on Computer Vision and Pattern Recognition , 2014, pp. 1725–1732
2014
Earlier work this paper cites.
N. Kalchbrenner, E. Grefenstette, and P. Blunsom, “A convolutional neural network for modelling sentences,” arXiv preprint arXiv:1404.2188 , 2014
Original
2014
Earlier work this paper cites.
Y. Taigman, M. Yang, M. Ranzato, and L. Wolf, “Deepface: Closing the gap to human-level performance in face verification,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2014, pp. 1701–1708
2014
Earlier work this paper cites.