Stochastic first-and zeroth-order methods for nonconvex stochastic programming
S. Ghadimi and G. Lan · 2013
Earlier work this paper cites.
Communication-efficient distributed optimization using an approximate newton-type method
O. Shamir, N. Srebro, and T. Zhang · 2014
Earlier work this paper cites.
Deep learning with elastic averaging sgd
S. Zhang, A. E. Choromanska, and Y. LeCun · 2015
Earlier work this paper cites.
Tensorflow: A system for large-scale machine learning
M. Abadi, P. Barham, J. Chen, Z. Chen, A. Davis, J. Dean, M. Devin, S. Ghemawat, G. Irving, M. Isard, M. K. Kudlur, J. Levenberg, R. Monga, S. Moore, D. G. Murray, B. Steiner, P. Tucker, V. Vasudevan, P. Warden, M. Wicke, Y. Yu, and X. Zheng · 2016
Earlier work this paper cites.
Aide: Fast and communication efficient distributed optimization
Original
S. J. Reddi, J. Konečnỳ, P. Richtárik, B. Póczós, and A. Smola · 2016
Earlier work this paper cites.
Communication-efficient learning of deep networks from decentralized data
H. B. McMahan, E. Moore, D. Ramage, S. Hampson, and B. A. y Arcas · 2017
Earlier work this paper cites.