Fetching the paper…
Reading the bibliography…
In this paper we prove that Local (S)GD (or FedAvg) can optimize deep neural networks with Rectified Linear Unit (ReLU) activation function in polynomial time.
Federated optimization: Distributed machine learning for on-device intelligence
Jakub Konečnỳ, H Brendan McMahan, Daniel Ramage, and Peter Richtárik · 2016
Earlier work this paper cites.
Federated learning: Strategies for improving communication efficiency
Jakub Konečnỳ, H Brendan McMahan, Felix X Yu, Peter Richtárik, Ananda Theertha Suresh, and Dave Bacon · 2016
Earlier work this paper cites.
Parallel sgd: When does averaging help?
Jian Zhang, Christopher De Sa, Ioannis Mitliagkas, and Christopher Ré · 2016
Earlier work this paper cites.
Globally optimal gradient descent for a convnet with gaussian inputs
Alon Brutzkus and Amir Globerson · 2017
Earlier work this paper cites.
Convergence analysis of two-layer neural networks with relu activation
Yuanzhi Li and Yang Yuan · 2017
Earlier work this paper cites.
Communication-efficient learning of deep networks from decentralized data
Brendan McMahan, Eider Moore, Daniel Ramage, Seth Hampson, and Blaise Aguera y Arcas · 2017
Earlier work this paper cites.
An analytical formula of population gradient for two-layered relu network and its applications in convergence and critical point analysis
Yuandong Tian · 2017
Earlier work this paper cites.
Recovery guarantees for one-hidden-layer neural networks
Kai Zhong, Zhao Song, Prateek Jain, Peter L Bartlett, and Inderjit S Dhillon · 2017
Earlier work this paper cites.
Learning and generalization in overparameterized neural networks, going beyond two layers
Zeyuan Allen-Zhu, Yuanzhi Li, and Yingyu Liang · 2018
Earlier work this paper cites.
On the power of over-parametrization in neural networks with quadratic activation
Simon Du and Jason Lee · 2018
Earlier work this paper cites.
Gradient descent provably optimizes over-parameterized neural networks
Simon S Du, Xiyu Zhai, Barnabas Poczos, and Aarti Singh · 2018
Earlier work this paper cites.
Federated learning for mobile keyboard prediction
Andrew Hard, Kanishka Rao, Rajiv Mathews, Swaroop Ramaswamy, Françoise Beaufays, Sean Augenstein, Hubert Eichner, Chloé Kiddon, and Daniel Ramage · 2018
Earlier work this paper cites.
Federated optimization in heterogeneous networks
Tian Li, Anit Kumar Sahu, Manzil Zaheer, Maziar Sanjabi, Ameet Talwalkar, and Virginia Smith · 2018
Cited alongside, same era.
Learning overparameterized neural networks via stochastic gradient descent on structured data
Yuanzhi Li and Yingyu Liang · 2018
Cited alongside, same era.
Local sgd converges fast and communicates little
Sebastian U Stich · 2018
Cited alongside, same era.
Federated learning with non-iid data
Yue Zhao, Meng Li, Liangzhen Lai, Naveen Suda, Damon Civin, and Vikas Chandra · 2018
Cited alongside, same era.
A convergence theory for deep learning via over-parameterization
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song · 2019
Tighter theory for local sgd on identical and heterogeneous data
A Khaled, K Mishchenko, and P Richtárik · 2020
Later among the works it cites.
Federated learning: Challenges, methods, and future directions
Tian Li, Anit Kumar Sahu, Ameet Talwalkar, and Virginia Smith · 2020
Later among the works it cites.
Global convergence of deep networks with one wide layer followed by pyramidal topology
Quynh Nguyen and Marco Mondelli · 2020
Later among the works it cites.
Minibatch vs local sgd for heterogeneous distributed learning
Blake Woodworth, Kumar Kshitij Patel, and Nathan Srebro · 2020
Later among the works it cites.
Is local sgd better than minibatch sgd?
Blake Woodworth, Kumar Kshitij Patel, Sebastian Stich, Zhen Dai, Brian Bullins, Brendan Mcmahan, Ohad Shamir, and Nathan Srebro · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Fine-grained analysis of optimization and generalization for overparameterized two-layer neural networks
Sanjeev Arora, Simon Du, Wei Hu, Zhiyuan Li, and Ruosong Wang · 2019
Cited alongside, same era.
Gradient descent finds global minima of deep neural networks
Simon Du, Jason Lee, Haochuan Li, Liwei Wang, and Xiyu Zhai · 2019
Cited alongside, same era.
Local sgd with periodic averaging: Tighter analysis and adaptive synchronization
Farzin Haddadpour, Mohammad Mahdi Kamani, Mehrdad Mahdavi, and Viveck Cadambe · 2019
Cited alongside, same era.
On the convergence of local descent methods in federated learning
Farzin Haddadpour and Mehrdad Mahdavi · 2019
Cited alongside, same era.
Scaffold: Stochastic controlled averaging for on-device federated learning
Sai Praneeth Karimireddy, Satyen Kale, Mehryar Mohri, Sashank J Reddi, Sebastian U Stich, and Ananda Theertha Suresh · 2019
Cited alongside, same era.
On the convergence of fedavg on non-iid data
Xiang Li, Kaixuan Huang, Wenhao Yang, Shusen Wang, and Zhihua Zhang · 2019
Cited alongside, same era.
An improved analysis of training over-parameterized deep neural networks
Difan Zou and Quanquan Gu · 2019
Cited alongside, same era.
Honglin Yuan and Tengyu Ma · 2020
Later among the works it cites.
Gradient descent optimizes over-parameterized deep relu networks
Difan Zou, Yuan Cao, Dongruo Zhou, and Quanquan Gu · 2020
Later among the works it cites.
On the global convergence of training deep linear resnets
Difan Zou, Philip M Long, and Quanquan Gu · 2020
Later among the works it cites.
Fl-ntk: A neural tangent kernel-based framework for federated learning analysis
Baihe Huang, Xiaoxiao Li, Zhao Song, and Xin Yang · 2021
Closest in time.
On the proof of global convergence of gradient descent for deep relu networks with linear widths
Quynh Nguyen · 2021
Closest in time.
On the convergence of deep networks with sample quadratic overparameterization
Asaf Noy, Yi Xu, Yonathan Aflalo, and Rong Jin · 2021
Closest in time.