Fetching the paper…
Reading the bibliography…
Gradient tracking (GT) is an algorithm designed for solving decentralized optimization problems over a network (such as training a machine learning model).
J.N. Tsitsiklis, Problems in decentralized decision making and computation. , Tech. Rep., Massachusetts Inst of Tech Cambridge Lab for Information and Decision Systems, 1984
1984
Earlier work this paper cites.
A. Nemirovski, A. Juditsky, G. Lan, and A. Shapiro, Robust stochastic approximation approach to stochastic programming , SIAM Journal on Optimization 19 (2009), pp. 1574–1609
2009
Earlier work this paper cites.
L. Deng, The mnist database of handwritten digit images for machine learning research , IEEE Signal Processing Magazine 29 (2012), pp. 141–142
2012
Earlier work this paper cites.
M. Li, D.G. Andersen, A.J. Smola, and K. Yu, Communication Efficient Distributed Machine Learning with the Parameter Server , in NeurIPS . 2014
2014
Earlier work this paper cites.
S. I. and C. S., Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift , in ICML . 2015
2015
Earlier work this paper cites.
W. Shi, Q. Ling, G. Wu, and W. Yin, EXTRA: An exact first-order algorithm for decentralized consensus optimization , SIAM Journal on Optimization 25 (2015), pp. 944–966
2015
Earlier work this paper cites.
2016
Earlier work this paper cites.
P.D. Lorenzo and G. Scutari, NEXT: In-network nonconvex optimization , IEEE Trans. Signal and Information Processing over Networks 2 (2016), pp. 120–136
2016
Earlier work this paper cites.
2017
Earlier work this paper cites.
X. Lian, C. Zhang, H. Zhang, C.J. Hsieh, W. Zhang, and J. Liu, Can Decentralized Algorithms Outperform Centralized Algorithms? A Case Study for Decentralized Parallel Stochastic Gradient Descent , in NeurIPS . 2017
2017
Earlier work this paper cites.
B. McMahan, E. Moore, D. Ramage, S. Hampson, and B.A. y Arcas, Communication-efficient learning of deep networks from decentralized data , in AISTATS . PMLR, 2017, pp. 1273–1282
2017
Earlier work this paper cites.
A. Nedić, A. Olshevsky, and W. Shi, Achieving geometric convergence for distributed optimization over time-varying graphs , SIAM Journal on Optimization 27 (2017), pp. 2597–2633
2017
Earlier work this paper cites.
X. Lian, W. Zhang, C. Zhang, and J. Liu, Asynchronous Decentralized Parallel Stochastic Gradient Descent , in ICML . 2018
2018
Cited alongside, same era.
S. Pu and A. Nedić, A Distributed Stochastic Gradient Tracking Method , in ICDC . 2018
2018
Cited alongside, same era.
H. Tang, X. Lian, M. Yan, C. Zhang, and J. Liu, D 2 D^{2} : Decentralized Training over Decentralized Data , in ICML . 2018
2018
Cited alongside, same era.
Y. You, Z. Zhang, C.J. Hsieh, J. Demmel, and K. Keutzer, ImageNet Training in Minutes , in ICPP . 2018
2018
Cited alongside, same era.
F. Haddadpour, M.M. Kamani, M. Mahdavi, and V. Cadambe, Local SGD with Periodic Averaging: Tighter Analysis and Adaptive Synchronization , in NeurIPS . 2019
2019
Cited alongside, same era.
A. Koloskova, N. Loizou, S. Boreiri, M. Jaggi, and S. Stich, A Unified Theory of Decentralized SGD with Changing Topology and Local Updates , in ICML . 2020
2020
Later among the works it cites.
B. Li, S. Cen, Y. Chen, and Y. Chi, Communication-Efficient Distributed Optimization in Networks with Gradient Tracking and Variance Reduction , in AISTATS . 2020
2020
Later among the works it cites.
T. Li, A.K. Sahu, M. Zaheer, M. Sanjabi, A. Talwalkar, and V. Smith, Federated Optimization in Heterogeneous Networks , in MLSys . 2020
2020
Later among the works it cites.
T. Lin, S.U. Stich, K.K. Patel, and M. Jaggi, Don’t Use Large Mini-batches, Use Local SGD , in ICLR . 2020
2020
Later among the works it cites.
Y. Liu, Variance reduction on decentralized training over heterogeneous data , Master’s thesis, ETH Zürich, 2021. Available at https://pub.tik.ee.ethz.ch/students/2020-HS/MA-2020-32.pdf
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Koloskova, S. Stich, and M. Jaggi, Decentralized Stochastic Optimization and Gossip Algorithms with Compressed Communication , in ICML . 2019
2019
Cited alongside, same era.
S.U. Stich, Local SGD Converges Fast and Communicates Little , in ICLR . 2019
2019
Cited alongside, same era.
2019
Cited alongside, same era.
S. Alghunaim and A. Sayed, Linear convergence of primal–dual gradient methods and their performance in distributed optimization , Automatica 117 (2020), p. 109003
2020
Cited alongside, same era.
K. Hsieh, A. Phanishayee, O. Mutlu, and P. Gibbons, The Non-IID Data Quagmire of Decentralized Machine Learning , in ICML . 2020
2020
Cited alongside, same era.
S.P. Karimireddy, S. Kale, M. Mohri, S. Reddi, S.U. Stich, and A.T. Suresh, SCAFFOLD: Stochastic Controlled Averaging for Federated Learning , in ICML . 2020
2020
Cited alongside, same era.
A. Koloskova, T. Lin, S.U. Stich, and M. Jaggi, Decentralized deep learning with arbitrary communication compression , ICLR (2020)
2020
Cited alongside, same era.
2020
Later among the works it cites.
R. Xin, U.A. Khan, and S. Kar, Variance-reduced decentralized stochastic optimization with accelerated convergence , IEEE Trans. Signal Process 68 (2020), pp. 6255–6271
2020
Later among the works it cites.
K. Yuan, W. Xu, and Q. Ling, Can primal methods outperform primal-dual methods in decentralized dynamic optimization? , IEEE Trans. Signal Process 68 (2020), pp. 4466–4480
2020
Later among the works it cites.
Chen et al., Accelerating gossip SGD with periodic global averaging , in ICML . 2021
2021
Later among the works it cites.
P. Kairouz, H.B. McMahan, B. Avent, A. Bellet, M. Bennis, A.N. Bhagoji, K. Bonawitz, Z. Charles, G. Cormode, R. Cummings, R.G.L. D’Oliveira, H. Eichner, S.E. Rouayheb, D. Evans, J. Gardner, Z. Garrett, A. Gascón, B. Ghazi, P.B. Gibbons, M. Gruteser, Z. Harchaoui, C. He, L. He, Z. Huo, B. Hutchinson, J. Hsu, M. Jaggi, T. Javidi, G. Joshi, M. Khodak, J. Konečný, A. Korolova, F. Koushanfar, S. Koyejo, T. Lepoint, Y. Liu, P. Mittal, M. Mohri, R. Nock, A. Özgür, R. Pagh, M. Raykova, H. Qi, D. Ramage, R. Raskar, D. Song, W. Song, S.U. Stich, Z. Sun, A.T. Suresh, F. Tramèr, P. Vepakomma, J. Wang, L. Xiong, Z. Xu, Q. Yang, F.X. Yu, H. Yu, and S. Zhao, Advances and open problems in federated learning , Foundations and Trends® in Machine Learning 14 (2021), pp. 1–210
2021
Later among the works it cites.
A. Koloskova, T. Lin, and S. Stich, An improved analysis of gradient tracking for decentralized machine learning , NeurIPS (2021)
2021
Later among the works it cites.