Fetching the paper…
Reading the bibliography…
Understanding the convergence performance of asynchronous stochastic gradient descent method (Async-SGD) has received increasing attention in recent years due to their foundational role in machine learning.
H. Robbins and S. Monro, “A stochastic approximation method,” The annals of mathematical statistics , pp. 400–407, 1951
1951
Earlier work this paper cites.
J. Kiefer and J. Wolfowitz, “Stochastic estimation of the maximum of a regression function,” The Annals of Mathematical Statistics , pp. 462–466, 1952
1952
Earlier work this paper cites.
T. P. Minka, “Old and new matrix algebra useful for statistics,” See www. stat. cmu. edu/minka/papers/matrix. html , 2000
2000
Earlier work this paper cites.
J.-P. Zhang, Z.-W. Li, and J. Yang, “A parallel svm training algorithm on large-scale classification problems,” in Machine Learning and Cybernetics, 2005. Proceedings of 2005 International Conference on , vol. 3. IEEE, 2005, pp. 1637–1641
2005
Earlier work this paper cites.
S. Gratton, A. Sartenaer, and P. L. Toint, “Recursive trust-region methods for multiscale nonlinear optimization,” SIAM Journal on Optimization , vol. 19, no. 1, pp. 414–444, 2008
2008
Earlier work this paper cites.
A. Nemirovski, A. Juditsky, G. Lan, and A. Shapiro, “Robust stochastic approximation approach to stochastic programming,” SIAM Journal on optimization , vol. 19, no. 4, pp. 1574–1609, 2009
2009
Earlier work this paper cites.
C. Cartis, N. I. Gould, and P. L. Toint, “On the complexity of steepest descent, newton’s and regularized newton’s methods for nonconvex unconstrained optimization problems,” Siam journal on optimization , vol. 20, no. 6, pp. 2833–2852, 2010
2010
Earlier work this paper cites.
L. Balzano, R. Nowak, and B. Recht, “Online identification and tracking of subspaces from highly incomplete information,” in Communication, Control, and Computing (Allerton), 2010 48th Annual Allerton Conference on . IEEE, 2010, pp. 704–711
2010
Earlier work this paper cites.
B. Recht, C. Re, S. Wright, and F. Niu, “Hogwild: A lock-free approach to parallelizing stochastic gradient descent,” in Advances in neural information processing systems , 2011, pp. 693–701
2011
Earlier work this paper cites.
E. Moulines and F. R. Bach, “Non-asymptotic analysis of stochastic approximation algorithms for machine learning,” in Advances in Neural Information Processing Systems , 2011, pp. 451–459
2011
Earlier work this paper cites.
A. Agarwal and J. C. Duchi, “Distributed delayed stochastic optimization,” in Advances in Neural Information Processing Systems , 2011, pp. 873–881
2011
Earlier work this paper cites.
H.-F. Yu, C.-J. Hsieh, S. Si, and I. Dhillon, “Scalable coordinate descent approaches to parallel matrix factorization for recommender systems,” in Data Mining (ICDM), 2012 IEEE 12th International Conference on . IEEE, 2012, pp. 765–774
2012
Cited alongside, same era.
J. Dean, G. Corrado, R. Monga, K. Chen, M. Devin, M. Mao, A. Senior, P. Tucker, K. Yang, Q. V. Le et al. , “Large scale distributed deep networks,” in Advances in neural information processing systems , 2012, pp. 1223–1231
2012
Cited alongside, same era.
2013
Cited alongside, same era.
R. Johnson and T. Zhang, “Accelerating stochastic gradient descent using predictive variance reduction,” in Advances in neural information processing systems , 2013, pp. 315–323
2013
Cited alongside, same era.
S. J. Reddi, A. Hefny, S. Sra, B. Poczos, and A. J. Smola, “On variance reduction in stochastic gradient descent and its asynchronous variants,” in Advances in Neural Information Processing Systems , 2015, pp. 2647–2655
2015
Later among the works it cites.
2015
Later among the works it cites.
P. L. Combettes and J.-C. Pesquet, “Stochastic quasi-fejér block-coordinate fixed point iterations with random sweeping,” SIAM Journal on Optimization , vol. 25, no. 2, pp. 1221–1248, 2015
2015
Later among the works it cites.
M. Abadi, P. Barham, J. Chen, Z. Chen, A. Davis, J. Dean, M. Devin, S. Ghemawat, G. Irving, M. Isard et al. , “Tensorflow: A system for large-scale machine learning.” in OSDI , vol. 16, 2016, pp. 265–283
2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
S. Ghadimi and G. Lan, “Stochastic first-and zeroth-order methods for nonconvex stochastic programming,” SIAM Journal on Optimization , vol. 23, no. 4, pp. 2341–2368, 2013
2013
Cited alongside, same era.
H.-F. Yu, C.-J. Hsieh, S. Si, and I. S. Dhillon, “Parallel matrix factorization for recommender systems,” Knowledge and Information Systems , vol. 41, no. 3, pp. 793–819, 2014
2014
Cited alongside, same era.
2014
Cited alongside, same era.
M. Li, D. G. Andersen, J. W. Park, A. J. Smola, A. Ahmed, V. Josifovski, J. Long, E. J. Shekita, and B.-Y. Su, “Scaling distributed machine learning with the parameter server.” in OSDI , vol. 1, no. 10.4, 2014, p. 3
2014
Cited alongside, same era.
2014
Cited alongside, same era.
D. Poeter. (2015) Gordon Moore Predicts 10 More Years for Moore’s Law. [Online]. Available: https://www.pcmag.com/article2/0,2817,2484098,00.asp
2015
Cited alongside, same era.
X. Lian, Y. Huang, Y. Li, and J. Liu, “Asynchronous parallel stochastic gradient for nonconvex optimization,” in Advances in Neural Information Processing Systems , 2015, pp. 2737–2745
2015
Cited alongside, same era.
S. J. Reddi, A. Hefny, S. Sra, B. Poczos, and A. Smola, “Stochastic variance reduction for nonconvex optimization,” in International conference on machine learning , 2016, pp. 314–323
2016
Later among the works it cites.
2016
Later among the works it cites.
S. Zheng, Q. Meng, T. Wang, W. Chen, N. Yu, Z.-M. Ma, and T.-Y. Liu, “Asynchronous stochastic gradient descent with delay compensation,” in International Conference on Machine Learning , 2017, pp. 4120–4129
2017
Later among the works it cites.
M. Schmidt, N. Le Roux, and F. Bach, “Minimizing finite sums with the stochastic average gradient,” Mathematical Programming , vol. 162, no. 1-2, pp. 83–112, 2017
2017
Later among the works it cites.
T. Sun, R. Hannah, and W. Yin, “Asynchronous coordinate descent under more realistic assumptions,” in Advances in Neural Information Processing Systems , 2017, pp. 6183–6191
2017
Later among the works it cites.
2018
Closest in time.
Z. Huo and H. Huang, “Asynchronous mini-batch gradient descent with variance reduction for non-convex optimization.” in AAAI , 2017, pp. 2043–2049
2049
Closest in time.