Fetching the paper…
Reading the bibliography…
Asynchronous methods are widely used in deep learning, but have limited theoretical justification when applied to non-convex problems.
B. T. Polyak, “Some methods of speeding up the convergence of iteration methods,” USSR Computational Mathematics and Mathematical Physics , vol. 4, no. 5, pp. 1–17, 1964
1964
Earlier work this paper cites.
W. A. Gardner, “Learning characteristics of stochastic-gradient-descent algorithms: A general study, analysis, and critique,” Signal Processing , vol. 6, no. 2, pp. 113–133, 1984
1984
Earlier work this paper cites.
S.-i. Amari, “Backpropagation and stochastic gradient descent method,” Neurocomputing , vol. 5, no. 4-5, pp. 185–196, 1993
1993
Earlier work this paper cites.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural computation , vol. 9, no. 8, pp. 1735–1780, 1997
1997
Earlier work this paper cites.
T. Zhang, “Solving large scale linear prediction problems using stochastic gradient descent algorithms,” in Proceedings of the twenty-first international conference on Machine learning . ACM, 2004, p. 116
2004
Earlier work this paper cites.
L. Bottou, “Large-scale machine learning with stochastic gradient descent,” in Proceedings of COMPSTAT’2010 . Springer, 2010, pp. 177–186
2010
Earlier work this paper cites.
F. Niu, B. Recht, C. Re, and S. Wright, “Hogwild: A lock-free approach to parallelizing stochastic gradient descent,” in Advances in Neural Information Processing Systems , 2011, pp. 693–701
2011
Earlier work this paper cites.
J. Dean, G. Corrado, R. Monga, K. Chen, M. Devin, M. Mao, A. Senior, P. Tucker, K. Yang, Q. V. Le et al. , “Large scale distributed deep networks,” in Advances in neural information processing systems , 2012, pp. 1223–1231
2012
Earlier work this paper cites.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,” in Advances in neural information processing systems , 2012, pp. 1097–1105
2012
Earlier work this paper cites.
I. Sutskever, J. Martens, G. Dahl, and G. Hinton, “On the importance of initialization and momentum in deep learning,” in Proceedings of the 30th international conference on machine learning (ICML-13) , 2013, pp. 1139–1147
2013
Cited alongside, same era.
T. Chilimbi, Y. Suzue, J. Apacible, and K. Kalyanaraman, “Project adam: Building an efficient and scalable deep learning training system,” in 11th USENIX Symposium on Operating Systems Design and Implementation (OSDI 14) , 2014, pp. 571–582
2014
Cited alongside, same era.
C. Zhang and C. Re, “Dimmwitted: A study of main-memory statistical analytics,” PVLDB , vol. 7, no. 12, pp. 1283–1294, 2014. [Online]. Available: http://www.vldb.org/pvldb/vol7/p1283-zhang.pdf
2014
Cited alongside, same era.
2015
Cited alongside, same era.
2016
Closest in time.
H. Cui, H. Zhang, G. R. Ganger, P. B. Gibbons, and E. P. Xing, “Geeps: Scalable deep learning on distributed gpus with a gpu-specialized parameter server,” in Proc. of the Eleventh European Conference on Computer Systems . ACM, 2016, p. 4
2016
Closest in time.
2016
Closest in time.
“Caffe solver documentation,” http://caffe.berkeleyvision.org/tutorial/solver.html , accessed: 2016-09-29
2016
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
S. Chaturapruek, J. C. Duchi, and C. Ré, “Asynchronous stochastic convex optimization: the noise is in the noise and sgd don’t care,” in NIPS , 2015, pp. 1531–1539
2015
Cited alongside, same era.
C. M. De Sa, C. Zhang, K. Olukotun, and C. Ré, “Taming the wild: A unified analysis of hogwild-style algorithms,” in Advances in Neural Information Processing Systems , 2015, pp. 2674–2682
2015
Cited alongside, same era.
2015
Cited alongside, same era.
2016
Cited alongside, same era.
2016
Closest in time.
2016
Closest in time.
“Caffe solver for CIFAR,” https://github.com/BVLC/caffe/blob/master/examples/cifar10/cifar10_quick_solver.prototxt , accessed: 2016-09-28
2016
Closest in time.
“CaffeNet, solver for ImageNet,” https://github.com/BVLC/caffe/tree/master/models/bvlc_reference_caffenet , accessed: 2016-09-28
2016
Closest in time.