Fetching the paper…
Reading the bibliography…
In this paper, we propose a distributed algorithm for stochastic smooth, non-convex optimization.
Y. Nesterov, “Introductory lectures on convex programming volume i: Basic course,” 1998
1998
Earlier work this paper cites.
A. Nemirovski, A. Juditsky, G. Lan, and A. Shapiro, “Robust stochastic approximation approach to stochastic programming,” SIAM Journal on optimization , vol. 19, no. 4, pp. 1574–1609, 2009
2009
Earlier work this paper cites.
E. Moulines and F. R. Bach, “Non-asymptotic analysis of stochastic approximation algorithms for machine learning,” in Advances in Neural Information Processing Systems , 2011, pp. 451–459
2011
Earlier work this paper cites.
O. Dekel, R. Gilad-Bachrach, O. Shamir, and L. Xiao, “Optimal distributed online prediction using mini-batches,” Journal of Machine Learning Research , vol. 13, no. Jan, pp. 165–202, 2012
2012
Earlier work this paper cites.
T. Léauté and B. Faltings, “Protecting privacy through distributed computation in multi-agent decision making,” Journal of Artificial Intelligence Research , vol. 47, pp. 649–695, 2013
2013
Earlier work this paper cites.
S. Ghadimi and G. Lan, “Stochastic first-and zeroth-order methods for nonconvex stochastic programming,” SIAM Journal on Optimization , vol. 23, no. 4, pp. 2341–2368, 2013
2013
Earlier work this paper cites.
R. Johnson and T. Zhang, “Accelerating stochastic gradient descent using predictive variance reduction,” in Advances in neural information processing systems , 2013, pp. 315–323
2013
Earlier work this paper cites.
2014
Earlier work this paper cites.
M. Li, D. G. Andersen, A. J. Smola, and K. Yu, “Communication efficient distributed machine learning with the parameter server,” in Advances in Neural Information Processing Systems , 2014, pp. 19–27
2014
Earlier work this paper cites.
A. Defazio, F. Bach, and S. Lacoste-Julien, “SAGA: A fast incremental gradient method with support for non-strongly convex composite objectives,” in Advances in neural information processing systems , 2014, pp. 1646–1654
2014
Earlier work this paper cites.
O. Shamir, N. Srebro, and T. Zhang, “Communication-efficient distributed optimization using an approximate newton-type method,” in International conference on machine learning , 2014, pp. 1000–1008
2014
Earlier work this paper cites.
E. P. Xing, Q. Ho, W. Dai, J. K. Kim, J. Wei, S. Lee, X. Zheng, P. Xie, A. Kumar, and Y. Yu, “Petuum: A new platform for distributed machine learning on big data,” IEEE Transactions on Big Data , vol. 1, no. 2, pp. 49–67, June 2015
2015
Earlier work this paper cites.
E. P. Xing, Q. Ho, P. Xie, and D. Wei, “Strategies and principles of distributed machine learning on big data,” Engineering , vol. 2, no. 2, pp. 179–195, 2016
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
N. Dryden, T. Moon, S. A. Jacobs, and B. Van Essen, “Communication quantization for data-parallel training of deep neural networks,” in 2016 2nd Workshop on Machine Learning in HPC Environments (MLHPC) . IEEE, 2016, pp. 1–8
2016
Cited alongside, same era.
2016
Cited alongside, same era.
S. J. Reddi, A. Hefny, S. Sra, B. Poczos, and A. Smola, “Stochastic Variance Reduction for Nonconvex Optimization,” in Int. Conf. Machine Learn. , 2016, pp. 314–323
2016
Cited alongside, same era.
Z. Allen-Zhu and E. Hazan, “Variance reduction for faster non-convex optimization,” in International conference on machine learning , 2016, pp. 699–707
2016
Cited alongside, same era.
P. Jiang and G. Agrawal, “A linear speedup analysis of distributed deep learning with sparse and quantized communication,” in Advances in Neural Information Processing Systems , 2018, pp. 2525–2536
2018
Later among the works it cites.
T. Chen, G. Giannakis, T. Sun, and W. Yin, “LAG: Lazily aggregated gradient for communication-efficient distributed learning,” in Advances in Neural Information Processing Systems , 2018, pp. 5050–5060
2018
Later among the works it cites.
2018
Later among the works it cites.
C. Fang, C. J. Li, Z. Lin, and T. Zhang, “SPIDER: Near-optimal non-convex optimization via stochastic path-integrated differential estimator,” in Advances in Neural Information Processing Systems , 2018, pp. 689–699
2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2016
Cited alongside, same era.
2016
Cited alongside, same era.
W. Wen, C. Xu, F. Yan, C. Wu, Y. Wang, Y. Chen, and H. Li, “TernGrad: Ternary gradients to reduce communication in distributed deep learning,” in Advances in neural information processing systems , 2017, pp. 1509–1519
2017
Cited alongside, same era.
D. Alistarh, D. Grubic, J. Li, R. Tomioka, and M. Vojnovic, “QSGD: Communication-efficient SGD via gradient quantization and encoding,” in Advances in Neural Information Processing Systems , 2017, pp. 1709–1720
2017
Cited alongside, same era.
2017
Cited alongside, same era.
J. D. Lee, Q. Lin, T. Ma, and T. Yang, “Distributed stochastic variance reduced gradient methods by sampling extra data with replacement,” The Journal of Machine Learning Research , vol. 18, no. 1, pp. 4404–4446, 2017
2017
Cited alongside, same era.
2017
Cited alongside, same era.
L. M. Nguyen, J. Liu, K. Scheinberg, and M. Takáč, “SARAH: A novel method for machine learning problems using stochastic recursive gradient,” in Proceedings of the 34th International Conference on Machine Learning-Volume 70 . JMLR. org, 2017, pp. 2613–2621
2017
Cited alongside, same era.
2018
Later among the works it cites.
D. Zhou, P. Xu, and Q. Gu, “Stochastic nested variance reduced gradient descent for nonconvex optimization,” in Adv. Neural Inf. Process. Systems , 2018, pp. 3921–3932
2018
Later among the works it cites.
2019
Closest in time.
2019
Closest in time.
H. Yu, S. Yang, and S. Zhu, “Parallel restarted SGD with faster convergence and less communication: Demystifying why model averaging works for deep learning,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 33, 2019, pp. 5693–5700
2019
Closest in time.
2019
Closest in time.
H. Yu and R. Jin, “On the computation and communication complexity of parallel SGD with dynamic batch sizes for stochastic non-convex optimization,” in International Conference on Machine Learning , 2019, pp. 7174–7183
2019
Closest in time.
F. Haddadpour, M. M. Kamani, M. Mahdavi, and V. Cadambe, “Trading redundancy for communication: Speeding up distributed SGD for non-convex optimization,” in International Conference on Machine Learning , 2019, pp. 2545–2554
2019
Closest in time.
2019
Closest in time.