Fetching the paper…
Reading the bibliography…
Distributed synchronous stochastic gradient descent (S-SGD) has been widely used in training large-scale deep neural networks (DNNs), but it typically requires very high communication bandwidth between computational workers (e.g., GPUs) to exchange gradients iteratively.
1901
Earlier work this paper cites.
M. P. Marcus, M. A. Marcinkiewicz, and B. Santorini, “Building a large annotated corpus of English: The Penn Treebank,” Computational linguistics , vol. 19, no. 2, pp. 313–330, 1993
1993
Earlier work this paper cites.
S. Sarvotham, R. Riedi, and R. Baraniuk, “Connection-level analysis and modeling of network traffic,” in The 1st ACM SIGCOMM Workshop on Internet Measurement . ACM, 2001, pp. 99–103
2001
Earlier work this paper cites.
E. Chan, M. Heimlich, A. Purkayastha, and R. Van De Geijn, “Collective communication: theory, practice, and experience,” Concurrency and Computation: Practice and Experience , vol. 19, no. 13, pp. 1749–1783, 2007
2007
Earlier work this paper cites.
J. Pješivac-Grbović, T. Angskun, G. Bosilca, G. E. Fagg, E. Gabriel, and J. J. Dongarra, “Performance analysis of MPI collective operations,” Cluster Computing , vol. 10, no. 2, pp. 127–143, 2007
2007
Earlier work this paper cites.
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “ImageNet: A large-scale hierarchical image database,” in Computer Vision and Pattern Recognition . IEEE, 2009, pp. 248–255
2009
Earlier work this paper cites.
T. Hoefler, W. Gropp, R. Thakur, and J. L. Träff, “Toward performance models of MPI implementations for understanding application scaling issues,” in European MPI Users’ Group Meeting . Springer, 2010, pp. 21–30
2010
Earlier work this paper cites.
A. Krizhevsky, V. Nair, and G. Hinton, “Cifar-10 (canadian institute for advanced research),” URL http://www.cs.toronto.edu/kriz/cifar.html , 2010
2010
Earlier work this paper cites.
J. Dean, G. Corrado, R. Monga, K. Chen, M. Devin, M. Mao, A. Senior, P. Tucker, K. Yang, Q. V. Le et al. , “Large scale distributed deep networks,” in Advances in Neural Information Processing Systems , 2012, pp. 1223–1231
2012
Earlier work this paper cites.
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
X. Mei, K. Zhao, C. Liu, and X. Chu, “Benchmarking the memory hierarchy of modern gpus,” in IFIP International Conference on Network and Parallel Computing . Springer, 2014, pp. 144–156
2014
Earlier work this paper cites.
F. Seide, H. Fu, J. Droppo, G. Li, and D. Yu, “1-bit stochastic gradient descent and its application to data-parallel distributed training of speech DNNs,” in Fifteenth Annual Conference of the International Speech Communication Association , 2014
2014
Cited alongside, same era.
Y. LeCun, Y. Bengio, and G. Hinton, “Deep learning,” Nature , vol. 521, no. 7553, p. 436, 2015
2015
Cited alongside, same era.
2016
Cited alongside, same era.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Computer Vision and Pattern Recognition , 2016, pp. 770–778
2016
Cited alongside, same era.
C.-Y. Chen, J. Choi, D. Brand, A. Agrawal, W. Zhang, and K. Gopalakrishnan, “Adacomp: Adaptive residual gradient compression for data-parallel distributed training,” in The 32nd AAAI Conference on Artificial Intelligence , 2018
2018
Later among the works it cites.
Y. Lin, S. Han, H. Mao, Y. Wang, and W. J. Dally, “Deep gradient compression: Reducing the communication bandwidth for distributed training,” in International Conference on Learning Representations , 2018
2018
Later among the works it cites.
J. Wu, W. Huang, J. Huang, and T. Zhang, “Error compensated quantized SGD and its applications to large-scale distributed optimization,” International Conference on Machine Learning , 2018
2018
Later among the works it cites.
J. Bernstein, Y.-X. Wang, K. Azizzadenesheli, and A. Anandkumar, “SIGNSGD: Compressed optimisation for non-convex problems,” in International Conference on Machine Learning , 2018, pp. 559–568
2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2017
Cited alongside, same era.
I. Hubara, M. Courbariaux, D. Soudry, R. El-Yaniv, and Y. Bengio, “Quantized neural networks: Training neural networks with low precision weights and activations,” The Journal of Machine Learning Research , vol. 18, no. 1, pp. 6869–6898, 2017
2017
Cited alongside, same era.
A. F. Aji and K. Heafield, “Sparse communication for distributed gradient descent,” in The 2017 Conference on Empirical Methods in Natural Language Processing , 2017, pp. 440–445
2017
Cited alongside, same era.
W. Wen, C. Xu, F. Yan, C. Wu, Y. Wang, Y. Chen, and H. Li, “Terngrad: Ternary gradients to reduce communication in distributed deep learning,” in Advances in Neural Information Processing Systems , 2017, pp. 1509–1519
2017
Cited alongside, same era.
S. Shi, W. Qiang, and X. Chu, “Performance modeling and evaluation of distributed deep learning frameworks on GPUs,” in The Fourth International Conference on Big Data Intelligence and Computing . IEEE, 2018
2018
Cited alongside, same era.
S. Shi, X. Chu, and B. Li, “A DAG model of synchronous stochastic gradient descent in distributed deep learning,” in The 24th International Conference on Parallel and Distributed Systems . IEEE, 2018
2018
Cited alongside, same era.
Y. You, Z. Zhang, C.-J. Hsieh, J. Demmel, and K. Keutzer, “ImageNet training in minutes,” in The 47th International Conference on Parallel Processing . ACM, 2018
2018
Cited alongside, same era.
X. Jia, S. Song, S. Shi, W. He, Y. Wang, H. Rong, F. Zhou, L. Xie, Z. Guo, Y. Yang, L. Yu, T. Chen, G. Hu, and X. Chu, “Highly scalable deep learning training system with mixed-precision: Training ImageNet in four minutes,” Workshop on Systems for ML and Open Source Software, collocated with NeurIPS 2018 , 2018
2018
Cited alongside, same era.
D. Alistarh, T. Hoefler, M. Johansson, N. Konstantinov, S. Khirirat, and C. Renggli, “The convergence of sparsified gradient methods,” in Advances in Neural Information Processing Systems , 2018, pp. 5973–5983
2018
Later among the works it cites.
S. U. Stich, J.-B. Cordonnier, and M. Jaggi, “Sparsified SGD with memory,” in Advances in Neural Information Processing Systems , 2018, pp. 4452–4463
2018
Later among the works it cites.
P. Jiang and G. Agrawal, “A linear speedup analysis of distributed deep learning with sparse and quantized communication,” in Advances in Neural Information Processing Systems , 2018, pp. 2530–2541
2018
Later among the works it cites.
2018
Later among the works it cites.
J. Wangni, J. Wang, J. Liu, and T. Zhang, “Gradient sparsification for communication-efficient distributed optimization,” in Advances in Neural Information Processing Systems , 2018, pp. 1299–1309
2018
Later among the works it cites.
2018
Later among the works it cites.
A. Shanbhag, H. Pirk, and S. Madden, “Efficient Top-K query processing on massively parallel hardware,” in The 2018 International Conference on Management of Data . ACM, 2018, pp. 1557–1570
2018
Later among the works it cites.
W. Wang and N. Srebro, “Stochastic nonconvex optimization with large minibatches,” The 30th International Conference on Algorithmic Learning Theory , 2019
2019
Closest in time.