Fetching the paper…
Reading the bibliography…
Federated Averaging (FedAvg) remains the most popular algorithm for Federated Learning (FL) optimization due to its simple implementation, stateless nature, and privacy guarantees combined with secure aggregation.
Minimization of functions having Lipschitz continuous first partial derivatives
Larry Armijo · 1966
Earlier work this paper cites.
The method of projections for finding the common point of convex sets
Leonid Georgievich Gurin, Boris Teodorovich Polyak, and È V Raik · 1967
Earlier work this paper cites.
Minimization of unsmooth functionals
Boris Teodorovich Polyak · 1969
Earlier work this paper cites.
Optimization of Lipschitz continuous functions
AA Goldstein · 1977
Earlier work this paper cites.
Convergence of the cyclical relaxation method for linear inequalities
Jan Mandel · 1984
Earlier work this paper cites.
Decomposition through formalization in a product space
Guy Pierra · 1984
Earlier work this paper cites.
Two-point step size gradient methods
Jonathan Barzilai and Jonathan M Borwein · 1988
Earlier work this paper cites.
On the barzilai and borwein choice of steplength for the gradient method
Marcos Raydan · 1993
Earlier work this paper cites.
Convex set theoretic image recovery by extrapolated iterations of parallel subgradient projections
Patrick L Combettes · 1997
Earlier work this paper cites.
Alternating projections, 2003
Stephen Boyd and Jon Dattarro · 2003
Earlier work this paper cites.
Mime: Mimicking centralized stochastic algorithms in federated learning
Sai Praneeth Karimireddy, Martin Jaggi, Satyen Kale, Mehryar Mohri, Sashank J Reddi, Sebastian U Stich, and Ananda Theertha Suresh · 2008
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Earlier work this paper cites.
Lecture 6.5-RMSProp: Divide the gradient by a running average of its recent magnitude
Tijmen Tieleman, Geoffrey Hinton, et al · 2012
Earlier work this paper cites.
Adadelta: an adaptive learning rate method
Matthew D Zeiler · 2012
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Practical secure aggregation for federated learning on user-held data
K. A. Bonawitz, Vladimir Ivanov, Ben Kreuter, Antonio Marcedone, H. Brendan McMahan, Sarvar Patel, Daniel Ramage, Aaron Segal, and Karn Seth · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Emnist: Extending mnist to handwritten letters
Gregory Cohen, Saeed Afshar, Jonathan Tapson, and Andre Van Schaik · 2017
Cited alongside, same era.
Communication-efficient learning of deep networks from decentralized data
Brendan McMahan, Eider Moore, Daniel Ramage, Seth Hampson, and Blaise Aguera y Arcas · 2017
Cited alongside, same era.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2017
Cited alongside, same era.
CINIC-10 is not Imagenet or CIFAR-10
Luke N Darlow, Elliot J Crowley, Antreas Antoniou, and Amos J Storkey · 2018
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Cited alongside, same era.
Tighter theory for local sgd on identical and heterogeneous data
Ahmed Khaled, Konstantin Mishchenko, and Peter Richtárik · 2020
Later among the works it cites.
Federated optimization in heterogeneous networks
Tian Li, Anit Kumar Sahu, Manzil Zaheer, Maziar Sanjabi, Ameet Talwalkar, and Virginia Smith · 2020
Later among the works it cites.
Adaptive gradient descent without descent
Yura Malitsky and Konstantin Mishchenko · 2020
Later among the works it cites.
Tackling the objective inconsistency problem in heterogeneous federated optimization
Jianyu Wang, Qinghua Liu, Hao Liang, Gauri Joshi, and H Vincent Poor · 2020
Later among the works it cites.
Why are adaptive methods good for attention models?
Jingzhao Zhang, Sai Praneeth Karimireddy, Andreas Veit, Seungyeon Kim, Sashank Reddi, Sanjiv Kumar, and Suvrit Sra · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dong Yin, Ashwin Pananjady, Max Lam, Dimitris Papailiopoulos, Kannan Ramchandran, and Peter Bartlett · 2018
Cited alongside, same era.
Learning and generalization in overparameterized neural networks, going beyond two layers
Zeyuan Allen-Zhu, Yuanzhi Li, and Yingyu Liang · 2019
Cited alongside, same era.
On exact computation with an infinitely wide neural net
Sanjeev Arora, Simon S Du, Wei Hu, Zhiyuan Li, Russ R Salakhutdinov, and Ruosong Wang · 2019
Cited alongside, same era.
Stabilized Barzilai-Borwein method
Oleg Burdakov, Yu-Hong Dai, and Na Huang · 2019
Cited alongside, same era.
Leaf: A benchmark for federated settings
Sebastian Caldas, Sai Meher Karthik Duddu, Peter Wu, Tian Li, Jakub Konečnỳ, H Brendan McMahan, Virginia Smith, and Ameet Talwalkar · 2019
Cited alongside, same era.
On the convergence of local descent methods in federated learning
Farzin Haddadpour and Mehrdad Mahdavi · 2019
Cited alongside, same era.
Revisiting the Polyak step size
Elad Hazan and Sham Kakade · 2019
Cited alongside, same era.
Durmus Alp Emre Acar, Yue Zhao, Ramon Matas, Matthew Mattina, Paul Whatmough, and Venkatesh Saligrama · 2021
Later among the works it cites.
FL-NTK: A neural tangent kernel-based framework for federated learning analysis
Baihe Huang, Xiaoxiao Li, Zhao Song, and Xin Yang · 2021
Later among the works it cites.
Advances and open problems in federated learning
Peter Kairouz, H Brendan McMahan, Brendan Avent, Aurélien Bellet, Mehdi Bennis, Arjun Nitin Bhagoji, Kallista Bonawitz, Zachary Charles, Graham Cormode, Rachel Cummings, et al · 2021
Later among the works it cites.
Stochastic polyak step-size for sgd: An adaptive learning rate for fast convergence
Nicolas Loizou, Sharan Vaswani, Issam Hadj Laradji, and Simon Lacoste-Julien · 2021
Later among the works it cites.
Linear convergence in federated learning: Tackling client heterogeneity and sparse gradients
Aritra Mitra, Rayana Jaafar, George J Pappas, and Hamed Hassani · 2021
Later among the works it cites.
Adaptive federated optimization
Sashank J. Reddi, Zachary Charles, Manzil Zaheer, Zachary Garrett, Keith Rush, Jakub Konečný, Sanjiv Kumar, and Hugh Brendan McMahan · 2021
Later among the works it cites.
Local SGD optimizes overparameterized neural networks in polynomial time
Yuyang Deng, Mohammad Mahdi Kamani, and Mehrdad Mahdavi · 2022
Later among the works it cites.
Adaptive learning rates for faster stochastic gradient methods
Samuel Horváth, Konstantin Mishchenko, and Peter Richtárik · 2022
Later among the works it cites.
Server-side stepsizes and sampling without replacement provably help in federated optimization
Grigory Malinovsky, Konstantin Mishchenko, and Peter Richtárik · 2022
Later among the works it cites.
ProxSkip: Yes! Local gradient steps provably lead to communication acceleration! Finally!
Konstantin Mishchenko, Grigory Malinovsky, Sebastian Stich, and Peter Richtarik · 2022
Later among the works it cites.
TCT: Convexifying federated learning using bootstrapped neural tangent kernels
Yaodong Yu, Alexander Wei, Sai Praneeth Karimireddy, Yi Ma, and Michael I Jordan · 2022
Later among the works it cites.
Neural tangent kernel empowered federated learning
Kai Yue, Richeng Jin, Ryan Pilgrim, Chau-Wai Wong, Dror Baron, and Huaiyu Dai · 2022
Later among the works it cites.