Fetching the paper…
Reading the bibliography…
Decentralized stochastic optimization methods have gained a lot of attention recently, mainly because of their cheap per iteration cost, data locality, and their communication-efficiency.
Unified optimal analysis of the (stochastic) gradient method
Stich, S. U · 1907
Earlier work this paper cites.
Algebraic connectivity of graphs
Fiedler, M · 1973
Earlier work this paper cites.
Problem complexity and method efficiency in optimization
Nemirovsky, A. S. and Yudin, D. B · 1983
Earlier work this paper cites.
Problems in decentralized decision making and computation
Tsitsiklis, J. N · 1984
Earlier work this paper cites.
Feddane: A federated newton-type method
Li, T., Sahu, A. K., Zaheer, M., Sanjabi, M., Talwalkar, A., and Smith, V · 2001
Earlier work this paper cites.
Gossip-based computation of aggregate information
Kempe, D., Dobra, A., and Gehrke, J · 2003
Earlier work this paper cites.
Introductory Lectures on Convex Optimization , volume 87 of Springer Science & Business Media
Nesterov, Y · 2004
Earlier work this paper cites.
Fast linear iterations for distributed averaging
Xiao, L. and Boyd, S · 2004
Earlier work this paper cites.
Randomized gossip algorithms
Boyd, S., Ghosh, A., Prabhakar, B., and Shah, D · 2006
Earlier work this paper cites.
Distributed subgradient methods for multi-agent optimization
Nedić, A. and Ozdaglar, A · 2009
Earlier work this paper cites.
A randomized incremental subgradient method for distributed optimization in networked systems
Johansson, B., Rabi, M., and Johansson, M · 2010
Earlier work this paper cites.
Asynchronous gossip algorithm for stochastic optimization: Constant stepsize analysis
Ram, S. S., Nedić, A., and Veeravalli, V. V · 2010
Earlier work this paper cites.
Parallelized stochastic gradient descent
Zinkevich, M., Weimer, M., Li, L., and Smola, A. J · 2010
Earlier work this paper cites.
Non-asymptotic analysis of stochastic approximation algorithms for machine learning
Bach, F. R. and Moulines, E · 2011
Earlier work this paper cites.
Dual averaging for distributed optimization: Convergence analysis and network scaling
Duchi, J. C., Agarwal, A., and Wainwright, M. J · 2011
Earlier work this paper cites.
Large scale distributed deep networks
Dean, J., Corrado, G., Monga, R., Chen, K., Devin, M., Mao, M., Ranzato, M., Senior, A., Tucker, P., Yang, K., Le, Q. V., and Ng, A. Y · 2012
Earlier work this paper cites.
Optimal distributed online prediction using mini-batches
Dekel, O., Gilad-Bachrach, R., Shamir, O., and Xiao, L · 2012
Earlier work this paper cites.
Lacoste-Julien, S., Schmidt, M. W., and Bach, F. R · 2012
Earlier work this paper cites.
Making gradient descent optimal for strongly convex stochastic optimization
Rakhlin, A., Shamir, O., and Sridharan, K · 2012
Earlier work this paper cites.
Distributed alternating direction method of multipliers
Wei, E. and Ozdaglar, A · 2012
Earlier work this paper cites.
Asynchronous distributed optimization using a randomized alternating direction method of multipliers
Iutzeler, F., Bianchi, P., Ciblat, P., and Hachem, W · 2013
Earlier work this paper cites.
Fast convergence of stochastic gradient descent under a strong growth condition
Schmidt, M. and Roux, N. L · 2013
Earlier work this paper cites.
Fast distributed gradient methods
Jakovetić, D., Xavier, J., and Moura, J. M. F · 2014
Earlier work this paper cites.
Distributed optimization over time-varying directed graphs
Nedić, A. and Olshevsky, A · 2014
Earlier work this paper cites.
Distributed stochastic optimization and learning
Shamir, O. and Srebro, N · 2014
Earlier work this paper cites.
Communication complexity of distributed convex learning and optimization
Arjevani, Y. and Shamir, O · 2015
Earlier work this paper cites.
Asynchronous stochastic convex optimization: the noise is in the noise and SGD don’t care
Chaturapruek, S., Duchi, J. C., and Ré, C · 2015
Earlier work this paper cites.
Iterative parameter mixing for distributed large-margin training of structured predictors for natural language processing
Coppola, G · 2015
Earlier work this paper cites.
Lee, J. D., Lin, Q., Ma, T., and Yang, T · 2015
Earlier work this paper cites.
Asynchronous gossip-based random projection algorithms over networks
Lee, S. and Nedić, A · 2015
Cited alongside, same era.
Decentralized online optimization with global objectives and local communication
Nedić, A., Lee, S., and Raginsky, M · 2015
Cited alongside, same era.
Multi-agent mirror descent for decentralized stochastic optimization
Rabbat, M · 2015
Cited alongside, same era.
EXTRA: An exact first-order algorithm for decentralized consensus optimization
Shi, W., Ling, Q., Wu, G., and Yin, W · 2015
Cited alongside, same era.
Federated optimization: Distributed machine learning for on-device intelligence
Konečnỳ, J., McMahan, H. B., Ramage, D., and Richtárik, P · 2016
Cited alongside, same era.
A new perspective on randomized gossip algorithms
Federated learning with non-iid data
Zhao, Y., Li, M., Lai, L., Suda, N., Civin, D., and Chandra, V · 2018
Later among the works it cites.
On the convergence properties of a k-step averaging stochastic gradient descent algorithm for nonconvex optimization
Zhou, F. and Cong, G · 2018
Later among the works it cites.
Linear convergence of primal-dual gradient methods and their performance in distributed optimization
Alghunaim, S. A. and Sayed, A. H · 2019
Later among the works it cites.
Stochastic gradient push for distributed deep learning
Assran, M., Loizou, N., Ballas, N., and Rabbat, M · 2019
Later among the works it cites.
Qsparse-local-SGD: Distributed SGD with quantization, sparsification, and local computations
Basu, D., Data, D., Karakus, C., and Diggavi, S · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Loizou, N. and Richtárik, P · 2016
Cited alongside, same era.
Federated learning of deep networks using model averaging
McMahan, H. B., Moore, E., Ramage, D., and y Arcas, B. A · 2016
Cited alongside, same era.
Stochastic gradient-push for strongly convex functions on time-varying directed graphs
Nedić, A. and Olshevsky, A · 2016
Cited alongside, same era.
A geometrically convergent method for distributed optimization over time-varying graphs
Nedić, A., Olshevsky, A., and Shi, W · 2016
Cited alongside, same era.
Stochastic gradient descent, weighted sampling, and the randomized Kaczmarz algorithm
Needell, D., Srebro, N., and Ward, R · 2016
Cited alongside, same era.
QSGD: Communication-efficient SGD via gradient quantization and encoding
Alistarh, D., Grubic, D., Li, J., Tomioka, R., and Vojnovic, M · 2017
Cited alongside, same era.
Can decentralized algorithms outperform centralized algorithms? a case study for decentralized parallel stochastic gradient descent
Lian, X., Zhang, C., Zhang, H., Hsieh, C.-J., Zhang, W., and Liu, J · 2017
Cited alongside, same era.
Later among the works it cites.
Robust distributed accelerated stochastic gradient methods for multi-agent networks
Fallah, A., Gürbüzbalaban, M., Ozdaglar, A., Şimşekli, U., and Zhu, L · 2019
Later among the works it cites.
SGD: General analysis and improved rates
Gower, R. M., Loizou, N., Qian, X., Sailanbayev, A., Shulgin, E., and Richtárik, P · 2019
Later among the works it cites.
Advances and open problems in federated learning
Kairouz, P., McMahan, H. B., Avent, B., Bellet, A., Bennis, M., Bhagoji, A. N., Bonawitz, K., Charles, Z., Cormode, G., Cummings, R., D’Oliveira, R. G. L., Rouayheb, S. E., Evans, D., Gardner, J., Garrett, Z., Gascón, A., Ghazi, B., Gibbons, P. B., Gruteser, M., Harchaoui, Z., He, C., He, L., Huo, Z., Hutchinson, B., Hsu, J., Jaggi, M., Javidi, T., Joshi, G., Khodak, M., Konečný, J., Korolova, A., Koushanfar, F., Koyejo, S., Lepoint, T., Liu, Y., Mittal, P., Mohri, M., Nock, R., Özgür, A., Pagh, R., Raykova, M., Qi, H., Ramage, D., Raskar, R., Song, D., Song, W., Stich, S. U., Sun, Z., Suresh, A. T., Tramèr, F., Vepakomma, P., Wang, J., Xiong, L., Xu, Z., Yang, Q., Yu, F. X., Yu, H., and Zhao, S · 2019
Later among the works it cites.
SCAFFOLD: Stochastic controlled averaging for on-device federated learning
Karimireddy, S. P., Kale, S., Mohri, M., Reddi, S. J., Stich, S. U., and Suresh, A. T · 2019
Later among the works it cites.
Decentralized stochastic optimization and gossip algorithms with compressed communication
Koloskova, A., Stich, S., and Jaggi, M · 2019
Later among the works it cites.
Communication efficient decentralized training with multiple local updates
Li, X., Yang, W., Wang, S., and Zhang, Z · 2019
Later among the works it cites.
Clique gossiping
Liu, Y., Li, B., Anderson, B. D., and Shi, G · 2019
Later among the works it cites.
Loizou, N. and Richtárik, P · 2019
Later among the works it cites.
PopSGD: Decentralized Stochastic Gradient Descent in the Population Model
Nadiradze, G., Sabour, A., Sharma, A., Markov, I., Aksenov, V., and Alistarh, D · 2019
Later among the works it cites.
A non-asymptotic analysis of network independence for distributed stochastic gradient descent
Olshevsky, A., Paschalidis, I. C., and Pu, S · 2019
Later among the works it cites.
Communication trade-offs for synchronized distributed SGD with large step size
Patel, K. K. and Dieuleveut, A · 2019
Later among the works it cites.
A sharp estimate on the transient time of distributed stochastic gradient descent
Pu, S., Olshevsky, A., and Paschalidis, I. C · 2019
Later among the works it cites.
Distributed nonconvex constrained optimization over time-varying digraphs
Scutari, G. and Sun, Y · 2019
Later among the works it cites.
Stich, S. U. and Karimireddy, S. P · 2019
Later among the works it cites.
Deepsqueeze: Decentralization meets error-compensated compression
Tang, H., Lian, X., Qiu, S., Yuan, L., Zhang, C., Zhang, T., and Liu, J · 2019
Later among the works it cites.
MATCHA: speeding up decentralized SGD via matching decomposition sampling
Wang, J., Sahu, A. K., Yang, Z., Joshi, G., and Kar, S · 2019
Later among the works it cites.
Parallel restarted SGD with faster convergence and less communication: Demystifying why model averaging works for deep learning
Yu, H., Yang, S., and Zhu, S · 2019
Later among the works it cites.
Tighter theory for local SGD on identical and heterogeneous data
Khaled, A., Mishchenko, K., and Richtárik, P · 2020
Closest in time.
Decentralized deep learning with arbitrary communication compression
Koloskova, A., Lin, T., Stich, S. U., and Jaggi, M · 2020
Closest in time.
Don’t use large mini-batches, use local SGD
Lin, T., Stich, S. U., Patel, K. K., and Jaggi, M · 2020
Closest in time.
Stochastic polyak step-size for SGD: An adaptive learning rate for fast convergence
Loizou, N., Vaswani, S., Laradji, I., and Lacoste-Julien, S · 2020
Closest in time.
Overlap local-SGD: An algorithmic approach to hide communication delays in distributed SGD
Wang, J., Liang, H., and Joshi, G · 2020
Closest in time.
Is local SGD better than minibatch SGD?
Woodworth, B., Patel, K. K., Stich, S. U., Dai, Z., Bullins, B., McMahan, H. B., Shamir, O., and Srebro, N · 2020
Closest in time.