Fetching the paper…
Reading the bibliography…
Large-scale distributed optimization is of great importance in various applications.
A method for the construction of minimum-redundancy codes
Huffman, D. A · 1952
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A · 2009
Earlier work this paper cites.
LIBSVM: A library for support vector machines
Chang, C.-C. and Lin, C.-J · 2011
Earlier work this paper cites.
Hogwild: A lock-free approach to parallelizing stochastic gradient descent
Recht, B., Re, C., Wright, S., and Niu, F · 2011
Earlier work this paper cites.
Large scale distributed deep networks
Dean, J., Corrado, G., Monga, R., Chen, K., Devin, M., Mao, M., aurelio Ranzato, M., Senior, A., Tucker, P., Yang, K., Le, Q. V., and Ng, A. Y · 2012
Earlier work this paper cites.
1-bit stochastic gradient descent and its application to data-parallel distributed training of speech DNNs
Seide, F., Fu, H., Droppo, J., Li, G., and Yu, D · 2014
Earlier work this paper cites.
Asynchronous stochastic convex optimization: the noise is in the noise and sgd don’t care
Chaturapruek, S., Duchi, J. C., and Ré, C · 2015
Earlier work this paper cites.
ImageNet large scale visual recognition challenge
Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M., Berg, A. C., and Fei-Fei, L · 2015
Cited alongside, same era.
Scalable distributed DNN training using commodity GPU cloud computing
Strom, N · 2015
Cited alongside, same era.
Performance modeling and scalability optimization of distributed deep learning systems
Yan, F., Ruwase, O., He, Y., and Chilimbi, T · 2015
Cited alongside, same era.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Cited alongside, same era.
Fast asynchronous parallel stochastic gradient descent: A lock-free approach with convergence guarantee
Zhao, S.-Y. and Li, W.-J · 2016
Cited alongside, same era.
DoReFa-Net: Training low bitwidth convolutional neural networks with low bitwidth gradients
QSGD: Communication-efficient SGD via gradient quantization and encoding
Alistarh, D., Grubic, D., Li, J., Tomioka, R., and Vojnovic, M · 2017
Later among the works it cites.
Understanding and optimizing asynchronous low-precision stochastic gradient descent
De Sa, C., Feldman, M., Ré, C., and Olukotun, K · 2017
Later among the works it cites.
Gradient sparsification for communication-efficient distributed optimization
Wangni, J., Wang, J., Liu, J., and Zhang, T · 2017
Later among the works it cites.
TernGrad: Ternary gradients to reduce communication in distributed deep learning
Wen, W., Xu, C., Yan, F., Wu, C., Wang, Y., Chen, Y., and Li, H · 2017
Later among the works it cites.
ZipML: Training linear models with end-to-end low precision, and a little bit of deep learning
Zhang, H., Li, J., Kara, K., Alistarh, D., Liu, J., and Zhang, C · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zhou, S., Ni, Z., Zhou, X., Wen, H., Wu, Y., and Zou, Y · 2016
Cited alongside, same era.
Sparse communication for distributed gradient descent
Aji, A. F. and Heafield, K · 2017
Cited alongside, same era.
Asynchronous stochastic gradient descent with delay compensation
Zheng, S., Meng, Q., Wang, T., Chen, W., Yu, N., Ma, Z.-M., and Liu, T.-Y · 2017
Later among the works it cites.
Deep gradient compression: Reducing the communication bandwidth for distributed training
Lin, Y., Han, S., Mao, H., Wang, Y., and Dally, W. J · 2018
Closest in time.