Fetching the paper…
Reading the bibliography…
We investigate fast and communication-efficient algorithms for the classic problem of minimizing a sum of strongly convex and smooth functions that are distributed among $n$ different nodes, which can communicate using a limited number of bits.
NUQSGD: improved communication efficiency for data-parallel SGD via nonuniform quantization
Ramezani-Kebrya, A., Faghri, F., and Roy, D. M · 1908
Earlier work this paper cites.
Cubic regularization of newton method and its global performance
Nesterov, Y. and Polyak, B · 1927
Earlier work this paper cites.
Communication complexity of convex optimization
Tsitsiklis, J. N. and Luo, Z · 1986
Earlier work this paper cites.
LIBSVM: A library for support vector machines
Chang, C.-C. and Lin, C.-J · 2011
Earlier work this paper cites.
Hogwild!: A lock-free approach to parallelizing stochastic gradient descent
Niu, F., Recht, B., Ré, C., and Wright, S. J · 2011
Earlier work this paper cites.
Information-theoretic lower bounds for distributed statistical estimation with communication constraints
Zhang, Y., Duchi, J., Jordan, M. I., and Wainwright, M. J · 2013
Earlier work this paper cites.
Communication-efficient distributed dual coordinate ascent
Jaggi, M., Smith, V., Takáč, M., Terhorst, J., Krishnan, S., Hofmann, T., and Jordan, M. I · 2014
Earlier work this paper cites.
Scaling distributed machine learning with the parameter server
Li, M., Andersen, D. G., Park, J. W., Smola, A. J., Ahmed, A., Josifovski, V., Long, J., Shekita, E. J., and Su, B.-Y · 2014
Earlier work this paper cites.
Fundamental limits of online and distributed algorithms for statistical learning and estimation
Shamir, O · 2014
Earlier work this paper cites.
Communication-efficient distributed optimization using an approximate newton-type method
Shamir, O., Srebro, N., and Zhang, T · 2014
Earlier work this paper cites.
Communication complexity of distributed convex learning and optimization
Arjevani, Y. and Shamir, O · 2015
Earlier work this paper cites.
Disco: Distributed optimization for self-concordant empirical loss
Zhang, Y. and Lin, X · 2015
Earlier work this paper cites.
Qsgd: Communication-efficient sgd via gradient quantization and encoding
Alistarh, D., Grubic, D., Li, J., Tomioka, R., and Vojnovic, M · 2016
Cited alongside, same era.
Aide: Fast and communication efficient distributed optimization
Reddi, S. J., Konečnỳ, J., Richtárik, P., Póczós, B., and Smola, A · 2016
Cited alongside, same era.
Optimal algorithms for smooth and strongly convex distributed optimization in networks
Scaman, K., Bach, F., Bubeck, S., Lee, Y. T., and Massoulié, L · 2017
Cited alongside, same era.
Distributed mean estimation with limited communication
Suresh, A. T., Yu, F. X., Kumar, S., and McMahan, H. B · 2017
Cited alongside, same era.
The convergence of sparsified gradient methods
Alistarh, D., Hoefler, T., Johansson, M., Khirirat, S., Konstantinov, N., and Renggli, C · 2018
Gradient methods for unconstrained problems
Chen, Y · 2019
Later among the works it cites.
Improved communication lower bounds for distributed optimisation
Alistarh, D. and Korhonen, J. H · 2020
Later among the works it cites.
Communication-efficient variance-reduced stochastic gradient descent
Ghadikolaei, H. S. and Magnússon, S · 2020
Later among the works it cites.
Statistically preconditioned accelerated gradient method for distributed optimization
Hendrikx, H., Xiao, L., Bubeck, S., Bach, F., and Massoulie, L · 2020
Later among the works it cites.
On maintaining linear convergence of distributed learning and optimization under limited communication
Magnússon, S., Shokri-Ghadikolaei, H., and Li, N · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Communication-efficient distributed statistical inference
Jordan, M. I., Lee, J. D., and Yang, Y · 2018
Cited alongside, same era.
Distributed learning with compressed gradients
Khirirat, S., Feyzmahdavian, H. R., and Johansson, M · 2018
Cited alongside, same era.
Sgd and hogwild! convergence without the bounded gradients assumption
Nguyen, L., Nguyen, P. H., Dijk, M., Richtárik, P., Scheinberg, K., and Takác, M · 2018
Cited alongside, same era.
Giant: Globally improved approximate newton method for distributed optimization
Wang, S., Roosta, F., Xu, P., and Mahoney, M. W · 2018
Cited alongside, same era.
Communication-computation efficient gradient coding
Ye, M. and Abbe, E · 2018
Cited alongside, same era.
Demystifying parallel and distributed deep learning: An in-depth concurrency analysis
Ben-Nun, T. and Hoefler, T · 2019
Cited alongside, same era.
Randomized block-diagonal preconditioning for parallel learning
Mendler-Dünner, C. and Lucchi, A · 2020
Later among the works it cites.
The communication complexity of optimization
Vempala, S. S., Wang, R., and Woodruff, D. P · 2020
Later among the works it cites.
Distributed adaptive newton methods with globally superlinear convergence
Zhang, J., You, K., and Başar, T · 2020
Later among the works it cites.
New bounds for distributed mean estimation and variance reduction
Davies, P., Gurunanthan, V., Moshrefi, N., Ashkboos, S., and Alistarh, D · 2021
Closest in time.
Distributed second order methods with fast rates and compressed communication
Islamov, R., Qian, X., and Richtárik, P · 2021
Closest in time.
Fednl: Making newton-type methods applicable to federated learning, 2021
Safaryan, M., Islamov, R., Qian, X., and Richtárik, P · 2021
Closest in time.