Fetching the paper…
Reading the bibliography…
We study the consensus decentralized optimization problem where the objective function is the average of $n$ agents private non-convex cost functions; moreover, the agents can only communicate to their neighbors on a given network topology.
B. T. Polyak, “Gradient methods for minimizing functionals,”
1963
Earlier work this paper cites.
New York: Optimization Software, 1987
B. Polyak, · 1987
Earlier work this paper cites.
A. Nedic and A. Ozdaglar, “Distributed subgradient methods for multi-agent optimization,”
2009
Earlier work this paper cites.
P. Patarasuk and X. Yuan, “Bandwidth optimal all-reduce algorithms for clusters of workstations,”
2009
Earlier work this paper cites.
S. S. Ram, A. Nedic, and V. V. Veeravalli, “Distributed stochastic subgradient projection algorithms for convex optimization,”
2010
Earlier work this paper cites.
F. S. Cattivelli and A. H. Sayed, “Diffusion LMS strategies for distributed estimation,”
2010
Earlier work this paper cites.
S. Boyd, N. Parikh, E. Chu, B. Peleato, and J. Eckstein, “Distributed optimization and statistical learning via alternating direction method of multipliers,”
2011
Earlier work this paper cites.
A. Antoniadis, I. Gijbels, and M. Nikolova, “Penalized likelihood regression for generalized linear models with non-quadratic penalties,”
2011
Earlier work this paper cites.
Cambridge University Press, 2012
R. A. Horn and C. R. Johnson, · 2012
Earlier work this paper cites.
J. Chen and A. H. Sayed, “Distributed pareto optimization via diffusion strategies,”
2013
Earlier work this paper cites.
P. Bianchi and J. Jakubowicz, “Convergence of a multi-agent projected stochastic gradient algorithm for non-convex optimization,”
2013
Earlier work this paper cites.
Springer, 2013
Y. Nesterov, · 2013
Earlier work this paper cites.
M. Li, D. G. Andersen, J. W. Park, A. J. Smola, A. Ahmed, V. Josifovski, J. Long, E. J. Shekita, and B.-Y. Su, “Scaling distributed machine learning with the parameter server,” in
2014
Earlier work this paper cites.
A. H. Sayed, “Adaptation, learning, and optimization over neworks.,”
2014
Earlier work this paper cites.
W. Shi, Q. Ling, G. Wu, and W. Yin, “EXTRA: An exact first-order algorithm for decentralized consensus optimization,”
2015
Earlier work this paper cites.
J. Xu, S. Zhu, Y. C. Soh, and L. Xie, “Augmented distributed gradient methods for multi-agent optimization under uncoordinated constant stepsizes,” in
2015
Earlier work this paper cites.
J. Chen and A. H. Sayed, “On the learning behavior of adaptive networks - Part I: Transient analysis,”
2015
Earlier work this paper cites.
J. Chen and A. H. Sayed, “On the learning behavior of adaptive networks - Part II: Performance analysis,”
2015
Earlier work this paper cites.
Q. Ling, W. Shi, G. Wu, and A. Ribeiro, “DLM: Decentralized linearized alternating direction method of multipliers,”
2015
Earlier work this paper cites.
K. Yuan, Q. Ling, and W. Yin, “On the convergence of decentralized gradient descent,”
2016
Earlier work this paper cites.
P. Di Lorenzo and G. Scutari, “Next: In-network nonconvex optimization,”
2016
Earlier work this paper cites.
H. Karimi, J. Nutini, and M. Schmidt, “Linear convergence of gradient and proximal-gradient methods under the Polyak-Lojasiewicz condition,” in
2016
Cited alongside, same era.
X. Lian, C. Zhang, H. Zhang, C.-J. Hsieh, W. Zhang, and J. Liu, “Can decentralized algorithms outperform centralized algorithms? A case study for decentralized parallel stochastic gradient descent,” in
2017
Cited alongside, same era.
A. Nedic, A. Olshevsky, and W. Shi, “Achieving geometric convergence for distributed optimization over time-varying graphs,”
2017
Cited alongside, same era.
T. Tatarenko and B. Touri, “Non-convex distributed optimization,”
2017
Cited alongside, same era.
Z. Jiang, A. Balu, C. Hegde, and S. Sarkar, “Collaborative deep learning in fixed topology networks,” in
2017
Cited alongside, same era.
A. Sundararajan, B. Van Scoy, and L. Lessard, “Analysis and design of first-order distributed optimization algorithms over time-varying graphs,”
2020
Later among the works it cites.
B. Swenson, R. Murray, S. Kar, and H. V. Poor, “Distributed stochastic gradient descent: Nonconvexity, nonsmoothness, and convergence to local minima,”
2020
Later among the works it cites.
A. Koloskova, N. Loizou, S. Boreiri, M. Jaggi, and S. Stich, “A unified theory of decentralized SGD with changing topology and local updates,” in
2020
Later among the works it cites.
S. Lu and C. W. Wu, “Decentralized stochastic non-convex optimization over weakly connected time-varying digraphs,” in
2020
Later among the works it cites.
Y. Tang, J. Zhang, and N. Li, “Distributed zero-order algorithms for nonconvex multiagent optimization,”
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2017
Cited alongside, same era.
H. Tang, X. Lian, M. Yan, C. Zhang, and J. Liu, “D
2018
Cited alongside, same era.
G. Qu and N. Li, “Harnessing smoothness to accelerate distributed optimization,”
2018
Cited alongside, same era.
M. Assran, N. Loizou, N. Ballas, and M. Rabbat, “Stochastic gradient push for distributed deep learning,” in
2019
Cited alongside, same era.
K. Yuan, B. Ying, X. Zhao, and A. H. Sayed, “Exact diffusion for distributed optimization and learning-Part I: Algorithm development,”
2019
Cited alongside, same era.
Z. Li, W. Shi, and M. Yan, “A decentralized proximal-gradient method with network independent step-sizes and separated convergence rates,”
2019
Cited alongside, same era.
G. Scutari and Y. Sun, “Distributed nonconvex constrained optimization over time-varying digraphs,”
2019
Cited alongside, same era.
Also available as arXiv preprint:2105.09080
Y. Chen, K. Yuan, Y. Zhang, P. Pan, Y. Xu, and W. Yin, “Accelerating gossip SGD with periodic global averaging,” in · 2021
Closest in time.
K. Yuan and S. A. Alghunaim, “Removing data heterogeneity influence enhances network topology dependence of decentralized SGD,”
2021
Closest in time.
K. Huang and S. Pu, “Improving the transient times for distributed stochastic gradient methods,”
2021
Closest in time.
S. Pu and A. Nedić, “Distributed stochastic gradient tracking methods,”
2021
Closest in time.
Y. Lu and C. De Sa, “Optimal complexity in decentralized training,” in
2021
Closest in time.
R. Xin, U. A. Khan, and S. Kar, “An improved convergence analysis for decentralized online stochastic non-convex optimization,”
2021
Closest in time.
S. A. Alghunaim, E. K. Ryu, K. Yuan, and A. H. Sayed, “Decentralized proximal gradient algorithms with linear convergence rates,”
2021
Closest in time.
J. Xu, Y. Tian, Y. Sun, and G. Scutari, “Distributed algorithms for composite optimization: Unified framework and convergence analysis,”
2021
Closest in time.
S. Pu, A. Olshevsky, and I. C. Paschalidis, “A sharp estimate on the transient time of distributed stochastic gradient descent,”
2021
Closest in time.
S. Vlaski and A. H. Sayed, “Distributed learning in non-convex environments—part II: Polynomial escape from saddle-points,”
2021
Closest in time.
J. Wang and G. Joshi, “Cooperative SGD: A unified framework for the design and analysis of local-update SGD algorithms,”
2021
Closest in time.
R. Xin, S. Das, U. A. Khan, and S. Kar, “A stochastic proximal gradient framework for decentralized non-convex composite optimization: Topology-independent sample complexity and communication efficiency,”
2021
Closest in time.
M. Gurbuzbalaban, U. Simsekli, and L. Zhu, “The heavy-tail phenomenon in SGD,” in
2021
Closest in time.
R. Xin, U. Khan, and S. Kar, “A hybrid variance-reduced method for decentralized stochastic non-convex optimization,” in
2021
Closest in time.
R. Xin, U. A. Khan, and S. Kar, “Fast decentralized nonconvex finite-sum optimization with recursive variance reduction,”
2022
Closest in time.
X. Yi, S. Zhang, T. Yang, T. Chai, and K. H. Johansson, “A primal-dual SGD algorithm for distributed nonconvex optimization,”
2022
Closest in time.