Fetching the paper…
Reading the bibliography…
Recent developments on large-scale distributed machine learning applications, e.g., deep neural networks, benefit enormously from the advances in distributed non-convex optimization techniques, e.g., distributed Stochastic Gradient Descent (SGD).
Some methods of speeding up the convergence of iteration methods
Polyak, B. T · 1964
Earlier work this paper cites.
Matrix Analysis
Horn, R. A. and Johnson, C. R · 1985
Earlier work this paper cites.
Heavy-ball method in nonconvex optimization problems
Zavriev, S. and Kostyuk, F · 1993
Earlier work this paper cites.
Introductory Lectures on Convex Optimization: A Basic Course
Nesterov, Y · 2004
Earlier work this paper cites.
ImageNet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A. and Hinton, G · 2009
Earlier work this paper cites.
Large scale distributed deep networks
Dean, J., Corrado, G., Monga, R., Chen, K., Devin, M., Mao, M., Senior, A., Tucker, P., Yang, K., Le, Q. V., et al · 2012
Earlier work this paper cites.
Optimal distributed online prediction using mini-batches
Dekel, O., Gilad-Bachrach, R., Shamir, O., and Xiao, L · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G. E · 2012
Earlier work this paper cites.
Stochastic first-and zeroth-order methods for nonconvex stochastic programming
Ghadimi, S. and Lan, G · 2013
Earlier work this paper cites.
On the importance of initialization and momentum in deep learning
Sutskever, I., Martens, J., Dahl, G., and Hinton, G · 2013
Earlier work this paper cites.
Global convergence of the heavy-ball method for convex optimization
Ghadimi, E., Feyzmahdavian, H. R., and Johansson, M · 2014
Earlier work this paper cites.
Communication efficient distributed machine learning with the parameter server
Li, M., Andersen, D. G., Smola, A. J., and Yu, K · 2014
Cited alongside, same era.
1-bit stochastic gradient descent and its application to data-parallel distributed training of speech DNNs
Seide, F., Fu, H., Droppo, J., Li, G., and Yu, D · 2014
Cited alongside, same era.
Asynchronous parallel stochastic gradient for nonconvex optimization
Lian, X., Huang, Y., Li, Y., and Liu, J · 2015
Cited alongside, same era.
Scalable distributed DNN training using commodity GPU cloud computing
Strom, N · 2015
Cited alongside, same era.
Scalable training of deep learning machines by incremental block training with intra-block parallel optimization and blockwise model-update filtering
Chen, K. and Huo, Q · 2016
Cited alongside, same era.
Deep residual learning for image recognition
Perturbed iterate analysis for asynchronous stochastic optimization
Mania, H., Pan, X., Papailiopoulos, D., Recht, B., Ramchandran, K., and Jordan, M. I · 2017
Later among the works it cites.
Communication-efficient learning of deep networks from decentralized data
McMahan, H. B., Moore, E., Ramage, D., Hampson, S., et al · 2017
Later among the works it cites.
Terngrad: Ternary gradients to reduce communication in distributed deep learning
Wen, W., Xu, C., Yan, F., Wu, C., Wang, Y., Chen, Y., and Li, H · 2017
Later among the works it cites.
Stochastic gradient push for distributed deep learning
Assran, M., Loizou, N., Ballas, N., and Rabbat, M · 2018
Later among the works it cites.
A linear speedup analysis of distributed deep learning with sparse and quantized communication
Jiang, P. and Agrawal, G · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Cited alongside, same era.
CNTK: Microsoft’s open-source deep-learning toolkit
Seide, F. and Agarwal, A · 2016
Cited alongside, same era.
Sparse communication for distributed gradient descent
Aji, A. F. and Heafield, K · 2017
Cited alongside, same era.
QSGD: Communication-efficient SGD via gradient quantization and encoding
Alistarh, D., Grubic, D., Li, J., Tomioka, R., and Vojnovic, M · 2017
Cited alongside, same era.
Accurate, large minibatch SGD: training imagenet in 1 hour
Goyal, P., Dollár, P., Girshick, R., Noordhuis, P., Wesolowski, L., Kyrola, A., Tulloch, A., Jia, Y., and He, K · 2017
Cited alongside, same era.
Collaborative deep learning in fixed topology networks
Jiang, Z., Balu, A., Hegde, C., and Sarkar, S · 2017
Cited alongside, same era.
Can decentralized algorithms outperform centralized algorithms? A case study for decentralized parallel stochastic gradient descent
Lian, X., Zhang, C., Zhang, H., Hsieh, C.-J., Zhang, W., and Liu, J · 2017
Cited alongside, same era.
Lin, T., Stich, S. U., and Jaggi, M · 2018
Later among the works it cites.
Local SGD converges fast and communicates little
Stich, S. U · 2018
Later among the works it cites.
Experiments on parallel training of deep neural network using model averaging
Su, H., Chen, H., and Xu, H · 2018
Later among the works it cites.
D 2 {D}^{2} : Decentralized training over decentralized data
Tang, H., Lian, X., Yan, M., Zhang, C., and Liu, J · 2018
Later among the works it cites.
A unified analysis of stochastic momentum methods for deep learning
Yan, Y., Yang, T., Li, Z., Lin, Q., and Yang, Y · 2018
Later among the works it cites.
Yu, H., Yang, S., and Zhu, S · 2018
Later among the works it cites.
Zhou, F. and Cong, G · 2018
Later among the works it cites.