Fetching the paper…
Reading the bibliography…
Recently, researchers proposed various low-precision gradient compression, for efficient communication in large-scale distributed optimization.
Minimization methods for nonsmooth convex and quasiconvex functions
Nesterov, Y · 1984
Earlier work this paper cites.
Numerical optimization
Wright, S. and Nocedal, J · 1999
Earlier work this paper cites.
Solving large scale linear prediction problems using stochastic gradient descent algorithms
Zhang, T · 2004
Earlier work this paper cites.
Introduction to algorithms
Cormen, T. H., Leiserson, C. E., Rivest, R. L., and Stein, C · 2009
Earlier work this paper cites.
Large-scale machine learning with stochastic gradient descent
Bottou, L · 2010
Earlier work this paper cites.
Computer vision: algorithms and applications
Szeliski, R · 2010
Earlier work this paper cites.
Distributed delayed stochastic optimization
Agarwal, A. and Duchi, J. C · 2011
Earlier work this paper cites.
Better mini-batch algorithms via accelerated gradient methods
Cotter, A., Shamir, O., Srebro, N., and Sridharan, K · 2011
Earlier work this paper cites.
Communication/computation tradeoffs in consensus-based distributed optimization
Tsianos, K., Lawlor, S., and Rabbat, M. G · 2012
Earlier work this paper cites.
Communication-efficient algorithms for statistical optimization
Zhang, Y., Wainwright, M. J., and Duchi, J. C · 2012
Earlier work this paper cites.
More effective distributed ml via a stale synchronous parallel parameter server
Ho, Q., Cipar, J., Cui, H., Lee, S., Kim, J. K., Gibbons, P. B., Gibson, G. A., Ganger, G., and Xing, E. P · 2013
Earlier work this paper cites.
Accelerating stochastic gradient descent using predictive variance reduction
Johnson, R. and Zhang, T · 2013
Earlier work this paper cites.
Introductory lectures on convex optimization: A basic course , volume 87
Nesterov, Y · 2013
Earlier work this paper cites.
Communication-efficient distributed optimization using an approximate newton-type method
Shamir, O., Srebro, N., and Zhang, T · 2014
Earlier work this paper cites.
Asynchronous distributed admm for consensus optimization
Zhang, R. and Kwok, J · 2014
Earlier work this paper cites.
Beyond convexity: Stochastic quasi-convex optimization
Hazan, E., Levy, K., and Shalev-Shwartz, S · 2015
Cited alongside, same era.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S. and Szegedy, C · 2015
Cited alongside, same era.
Path-SGD: Path-normalized optimization in deep neural networks
Neyshabur, B., Salakhutdinov, R. R., and Srebro, N · 2015
Cited alongside, same era.
Disco: Distributed optimization for self-concordant empirical loss
Zhang, Y. and Lin, X · 2015
Cited alongside, same era.
A stochastic quasi-Newton method for large-scale optimization
Byrd, R. H., Hansen, S. L., Nocedal, J., and Singer, Y · 2016
Cited alongside, same era.
Weight normalization: A simple reparameterization to accelerate training of deep neural networks
Salimans, T. and Kingma, D. P · 2016
Can decentralized algorithms outperform centralized algorithms? a case study for decentralized parallel stochastic gradient descent
Lian, X., Zhang, C., Zhang, H., Hsieh, C.-J., Zhang, W., and Liu, J · 2017
Later among the works it cites.
Minimizing finite sums with the stochastic average gradient
Schmidt, M., Le Roux, N., and Bach, F · 2017
Later among the works it cites.
Improved optimization of finite sums with minibatch stochastic variance reduced proximal iterations
Wang, J. and Zhang, T · 2017
Later among the works it cites.
Terngrad: Ternary gradients to reduce communication in distributed deep learning
Wen, W., Xu, C., Yan, F., Wu, C., Wang, Y., Chen, Y., and Li, H · 2017
Later among the works it cites.
The convergence of sparsified gradient methods
Alistarh, D., Hoefler, T., Johansson, M., Konstantinov, N., Khirirat, S., and Renggli, C · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
On the convergence of decentralized gradient descent
Yuan, K., Ling, Q., and Yin, W · 2016
Cited alongside, same era.
Dorefa-net: Training low bitwidth convolutional neural networks with low bitwidth gradients
Zhou, S., Wu, Y., Ni, Z., Zhou, X., Wen, H., and Zou, Y · 2016
Cited alongside, same era.
Sparse communication for distributed gradient descent
Aji, A. F. and Heafield, K · 2017
Cited alongside, same era.
QSGD: Communication-efficient SGD via gradient quantization and encoding
Alistarh, D., Grubic, D., Li, J., Tomioka, R., and Vojnovic, M · 2017
Cited alongside, same era.
Accurate, large minibatch SGD: training imagenet in 1 hour
Goyal, P., Dollár, P., Girshick, R., Noordhuis, P., Wesolowski, L., Kyrola, A., Tulloch, A., Jia, Y., and He, K · 2017
Cited alongside, same era.
How to escape saddle points efficiently
Jin, C., Ge, R., Netrapalli, P., Kakade, S. M., and Jordan, M. I · 2017
Cited alongside, same era.
SignSGD: compressed optimisation for non-convex problems
Bernstein, J., Wang, Y.-X., Azizzadenesheli, K., and Anandkumar, A · 2018
Later among the works it cites.
Optimization methods for large-scale machine learning
Bottou, L., Curtis, F. E., and Nocedal, J · 2018
Later among the works it cites.
An alternative view: When does SGD escape local minima?
Kleinberg, R., Li, Y., and Yuan, Y · 2018
Later among the works it cites.
SGD and Hogwild! convergence without the bounded gradients assumption
Nguyen, L. M., Nguyen, P. H., van Dijk, M., Richtárik, P., Scheinberg, K., and Takáč, M · 2018
Later among the works it cites.
Sparsified SGD with memory
Stich, S. U., Cordonnier, J.-B., and Jaggi, M · 2018
Later among the works it cites.
Atomo: Communication-efficient learning via atomic sparsification
Wang, H., Sievert, S., Liu, S., Charles, Z., Papailiopoulos, D., and Wright, S · 2018
Later among the works it cites.
Wang, J. and Joshi, G · 2018
Later among the works it cites.
Gradient sparsification for communication-efficient distributed optimization
Wangni, J., Wang, J., Liu, J., and Zhang, T · 2018
Later among the works it cites.
Error compensated quantized SGD and its applications to large-scale distributed optimization
Wu, J., Huang, W., Huang, J., and Zhang, T · 2018
Later among the works it cites.