Fetching the paper…
Reading the bibliography…
We study gradient compression methods to alleviate the communication bottleneck in data-parallel distributed optimization.
A Stochastic Approximation Method
Robbins, H. and Monro, S · 1951
Earlier work this paper cites.
Methods of simultaneous iteration for calculating eigenvectors of matrices
Stewart, G. and Miller, J · 1975
Earlier work this paper cites.
Simultaneous iteration for computing invariant subspaces of non-Hermitian matrices
Stewart, G · 1976
Earlier work this paper cites.
Simplified neuron model as a principal component analyzer
Oja, E · 1982
Earlier work this paper cites.
Spectral regularization algorithms for learning large incomplete matrices
Mazumder, R., Hastie, T., and Tibshirani, R · 2010
Earlier work this paper cites.
Large scale distributed deep networks
Dean, J., Corrado, G., Monga, R., Chen, K., Devin, M., Mao, M., Senior, A., Tucker, P., Yang, K., Le, Q. V., et al · 2012
Earlier work this paper cites.
1-bit stochastic gradient descent and its application to data-parallel distributed training of speech dnns
Seide, F., Fu, H., Droppo, J., Li, G., and Yu, D · 2014
Earlier work this paper cites.
Stochastic Spectral Descent for Restricted Boltzmann Machines
Carlson, D., Cevher, V., and Carin, L · 2015
Earlier work this paper cites.
Iandola, F. N., Ashraf, K., Moskewicz, M. W., and Keutzer, K · 2015
Earlier work this paper cites.
Lecture notes on solving large scale eigenvalue problems
Arbenz, P · 2016
Earlier work this paper cites.
Accelerated gradient methods for nonconvex nonlinear and stochastic programming
Ghadimi, S. and Lan, G · 2016
Earlier work this paper cites.
Federated learning: Strategies for improving communication efficiency
Konečnỳ, J., McMahan, H. B., Yu, F. X., Richtárik, P., Suresh, A. T., and Bacon, D · 2016
Earlier work this paper cites.
QSGD: Communication-efficient sgd via gradient quantization and encoding
Alistarh, D., Grubic, D., Li, J., Tomioka, R., and Vojnovic, M · 2017
Earlier work this paper cites.
Accurate, large minibatch SGD: training imagenet in 1 hour
Goyal, P., Dollar, P., Girshick, R., Noordhuis, P., Wesolowski, L., Kyrola, A., Tulloch, A., Jia, Y., and He, K · 2017
Cited alongside, same era.
Terngrad: Ternary gradients to reduce communication in distributed deep learning
Wen, W., Xu, C., Yan, F., Wu, C., Wang, Y., Chen, Y., and Li, H · 2017
Cited alongside, same era.
Spectral norm regularization for improving the generalizability of deep learning
Yoshida, Y. and Miyato, T · 2017
Cited alongside, same era.
Sketchy decisions: Convex low-rank matrix optimization with optimal storage
Yurtsever, A., Udell, M., Tropp, J. A., and Cevher, V · 2017
Cited alongside, same era.
Stronger generalization bounds for deep nets via a compression approach
Arora, S., Ge, R., Neyshabur, B., and Zhang, Y · 2018
Cited alongside, same era.
Measuring the effects of data parallelism on neural network training
Shallue, C. J., Lee, J., Antognini, J., Sohl-Dickstein, J., Frostig, R., and Dahl, G. E · 2018
Later among the works it cites.
Sparsified SGD with memory
Stich, S. U., Cordonnier, J.-B., and Jaggi, M · 2018
Later among the works it cites.
ATOMO: Communication-efficient learning via atomic sparsification
Wang, H., Sievert, S., Liu, S., Charles, Z., Papailiopoulos, D., and Wright, S · 2018
Later among the works it cites.
Gradient sparsification for communication-efficient distributed optimization
Wangni, J., Wang, J., Liu, J., and Zhang, T · 2018
Later among the works it cites.
Gradiveq: Vector quantization for bandwidth-efficient gradient aggregation in distributed CNN training
Yu, M., Lin, Z., Narra, K., Li, S., Li, Y., Kim, N. S., Schwing, A. G., Annavaram, M., and Avestimehr, S · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Optimized broadcast for deep learning workloads on dense-GPU infiniband clusters: MPI or NCCL?
Awan, A. A., Chu, C.-H., Subramoni, H., and Panda, D. K · 2018
Cited alongside, same era.
signSGD: compressed optimisation for non-convex problems
Bernstein, J., Wang, Y.-X., Azizzadenesheli, K., and Anandkumar, A · 2018
Cited alongside, same era.
Expanding the reach of federated learning by reducing client resource requirements
Caldas, S., Konecný, J., McMahan, H. B., and Talwalkar, A · 2018
Cited alongside, same era.
Detecting memorization in ReLU networks
Collins, E., Bigdeli, S. A., and Süsstrunk, S · 2018
Cited alongside, same era.
Characterizing implicit bias in terms of optimization geometry
Gunasekar, S., Lee, J., Soudry, D., and Srebro, N · 2018
Cited alongside, same era.
Algorithmic regularization in over-parameterized matrix sensing and neural networks with quadratic activations
Li, Y., Ma, T., and Zhang, H · 2018
Cited alongside, same era.
Deep gradient compression: Reducing the communication bandwidth for distributed training
Lin, Y., Han, S., Mao, H., Wang, Y., and Dally, W. J · 2018
Cited alongside, same era.
Adaptive input representations for neural language modeling
Baevski, A. and Auli, M · 2019
Closest in time.
signSGD with majority vote is communication efficient and fault tolerant
Bernstein, J., Zhao, J., Azizzadenesheli, K., and Anandkumar, A · 2019
Closest in time.
Error feedback fixes SignSGD and other gradient compression schemes
Karimireddy, S. P., Rebjock, Q., Stich, S. U., and Jaggi, M · 2019
Closest in time.
fairseq: A fast, extensible toolkit for sequence modeling
Ott, M., Edunov, S., Baevski, A., Fan, A., Gross, S., Ng, N., Grangier, D., and Auli, M · 2019
Closest in time.
High performance distributed deep learning: A beginner’s guide
Panda, D. K. D., Subramoni, H., and Awan, A. A · 2019
Closest in time.
Stich, S. U. and Karimireddy, S. P · 2019
Closest in time.
signSGD with majority vote
Zhao, J · 2019
Closest in time.