Fetching the paper…
Reading the bibliography…
We propose Federated Accelerated Stochastic Gradient Descent (FedAc), a principled acceleration of Federated Averaging (FedAvg, also known as Local SGD) for distributed optimization.
Unified Optimal Analysis of the (Stochastic) Gradient Method
Sebastian U. Stich · 1907
Earlier work this paper cites.
A Stochastic Approximation Method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
Problem complexity and method efficiency in optimization
A.S. Nemirovski and D. B. Yudin · 1983
Earlier work this paper cites.
Parallel Gradient Distribution in Unconstrained Optimization
L. O. Mangasarian · 1995
Earlier work this paper cites.
Statistical Learning Theory
Vladimir Naumovich Vapnik · 1998
Earlier work this paper cites.
Stability and generalization
Olivier Bousquet and André Elisseeff · 2002
Earlier work this paper cites.
Stochastic Approximation and Recursive Algorithms and Applications
Harold J Kushner, George Yin, and Harold J Kushner · 2003
Earlier work this paper cites.
Pascal large scale learning challenge
Soeren Sonnenburg, Vojtech Franc, Elad Yom-Tov, and Michele Sebag · 2008
Earlier work this paper cites.
Efficient large-scale distributed training of conditional maximum entropy models
Ryan Mcdonald, Mehryar Mohri, Nathan Silberman, Dan Walker, and Gideon S. Mann · 2009
Earlier work this paper cites.
Parallelized stochastic gradient descent
Martin Zinkevich, Markus Weimer, Lihong Li, and Alex J. Smola · 2010
Earlier work this paper cites.
LIBSVM: A library for support vector machines
Chih-Chung Chang and Chih-Jen Lin · 2011
Earlier work this paper cites.
Better mini-batch algorithms via accelerated gradient methods
Andrew Cotter, Ohad Shamir, Nati Srebro, and Karthik Sridharan · 2011
Earlier work this paper cites.
Optimal distributed online prediction using mini-batches
Ofer Dekel, Ran Gilad-Bachrach, Ohad Shamir, and Lin Xiao · 2012
Earlier work this paper cites.
Optimal Stochastic Approximation Algorithms for Strongly Convex Stochastic Composite Optimization I: A Generic Algorithmic Framework
Saeed Ghadimi and Guanghui Lan · 2012
Earlier work this paper cites.
Simon Lacoste-Julien, Mark Schmidt, and Francis Bach · 2012
Earlier work this paper cites.
Matrix Computations
Gene H. Golub and Charles F. Van Loan · 2013
Earlier work this paper cites.
Accelerating stochastic gradient descent using predictive variance reduction
Rie Johnson and Tong Zhang · 2013
Earlier work this paper cites.
Communication-efficient algorithms for statistical optimization
Yuchen Zhang, John C. Duchi, and Martin J. Wainwright · 2013
Earlier work this paper cites.
Iterative Parameter Mixing for Distributed Large-Margin Training of Structured Predictors for Natural Language Processing
Gregory Francis Coppola · 2014
Earlier work this paper cites.
Distributed stochastic optimization and learning
Ohad Shamir and Nathan Srebro · 2014
Cited alongside, same era.
Divide and conquer kernel ridge regression: A distributed algorithm with minimax optimal rates
Yuchen Zhang, John Duchi, and Martin Wainwright · 2015
Cited alongside, same era.
Train faster, generalize better: Stability of stochastic gradient descent
Moritz Hardt, Ben Recht, and Yoram Singer · 2016
Cited alongside, same era.
Analysis and Design of Optimization Algorithms via Integral Quadratic Constraints
Laurent Lessard, Benjamin Recht, and Andrew Packard · 2016
Cited alongside, same era.
On the optimality of averaging in distributed statistical learning
Jonathan D. Rosenblatt and Boaz Nadler · 2016
Cited alongside, same era.
QSGD: Communication-efficient SGD via gradient quantization and encoding
Dan Alistarh, Demjan Grubic, Jerry Li, Ryota Tomioka, and Milan Vojnovic · 2017
Qsparse-local-sgd: Distributed SGD with quantization, sparsification and local computations
Debraj Basu, Deepesh Data, Can Karakus, and Suhas Diggavi · 2019
Later among the works it cites.
Communication trade-offs for Local-SGD with large step size
Aymeric Dieuleveut and Kumar Kshitij Patel · 2019
Later among the works it cites.
On the Convergence of Local Descent Methods in Federated Learning
Farzin Haddadpour and Mehrdad Mahdavi · 2019
Later among the works it cites.
Advances and Open Problems in Federated Learning
Peter Kairouz, H. Brendan McMahan, Brendan Avent, Aurélien Bellet, Mehdi Bennis, Arjun Nitin Bhagoji, Keith Bonawitz, Zachary Charles, Graham Cormode, Rachel Cummings, Rafael G. L. D’Oliveira, Salim El Rouayheb, David Evans, Josh Gardner, Zachary Garrett, Adrià Gascón, Badih Ghazi, Phillip B. Gibbons, Marco Gruteser, Zaid Harchaoui, Chaoyang He, Lie He, Zhouyuan Huo, Ben Hutchinson, Justin Hsu, Martin Jaggi, Tara Javidi, Gauri Joshi, Mikhail Khodak, Jakub Konečný, Aleksandra Korolova, Farinaz Koushanfar, Sanmi Koyejo, Tancrède Lepoint, Yang Liu, Prateek Mittal, Mehryar Mohri, Richard Nock, Ayfer Özgür, Rasmus Pagh, Mariana Raykova, Hang Qi, Daniel Ramage, Ramesh Raskar, Dawn Song, Weikang Song, Sebastian U. Stich, Ziteng Sun, Ananda Theertha Suresh, Florian Tramèr, Praneeth Vepakomma, Jianyu Wang, Li Xiong, Zheng Xu, Qiang Yang, Felix X. Yu, Han Yu, and Sen Zhao · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
UCI machine learning repository
Dheeru Dua and Casey Graff · 2017
Cited alongside, same era.
Distributed stochastic variance reduced gradient methods by sampling extra data with replacement
Jason D. Lee, Qihang Lin, Tengyu Ma, and Tianbao Yang · 2017
Cited alongside, same era.
Perturbed Iterate Analysis for Asynchronous Stochastic Optimization
Horia Mania, Xinghao Pan, Dimitris Papailiopoulos, Benjamin Recht, Kannan Ramchandran, and Michael I. Jordan · 2017
Cited alongside, same era.
Communication-efficient learning of deep networks from decentralized data
Brendan McMahan, Eider Moore, Daniel Ramage, Seth Hampson, and Blaise Aguera y Arcas · 2017
Cited alongside, same era.
TernGrad: Ternary gradients to reduce communication in distributed deep learning
Wei Wen, Cong Xu, Feng Yan, Chunpeng Wu, Yandan Wang, Yiran Chen, and Hai Li · 2017
Cited alongside, same era.
Accelerated Methods for NonConvex Optimization
Yair Carmon, John C. Duchi, Oliver Hinder, and Aaron Sidford · 2018
Cited alongside, same era.
Distributed Learning with Compressed Gradient Differences
Konstantin Mishchenko, Eduard Gorbunov, Martin Takáč, and Peter Richtárik · 2019
Later among the works it cites.
Sebastian U. Stich and Sai Praneeth Karimireddy · 2019
Later among the works it cites.
Cooperative SGD: A unified Framework for the Design and Analysis of Communication-Efficient SGD Algorithms
Jianyu Wang and Gauri Joshi · 2019
Later among the works it cites.
On the computation and communication complexity of parallel SGD with dynamic batch sizes for stochastic non-convex optimization
Hao Yu and Rong Jin · 2019
Later among the works it cites.
On the rates of convergence of parallelized averaged stochastic gradient algorithms
Antoine Godichon-Baggioni and Sofiane Saadane · 2020
Closest in time.
SCAFFOLD: Stochastic Controlled Averaging for Federated Learning
Sai Praneeth Karimireddy, Satyen Kale, Mehryar Mohri, Sashank J. Reddi, Sebastian U. Stich, and Ananda Theertha Suresh · 2020
Closest in time.
Tighter Theory for Local SGD on Identical and Heterogeneous Data
Ahmed Khaled, Konstantin Mishchenko, and Peter Richtárik · 2020
Closest in time.
A Unified Theory of Decentralized SGD with Changing Topology and Local Updates
Anastasia Koloskova, Nicolas Loizou, Sadra Boreiri, Martin Jaggi, and Sebastian U. Stich · 2020
Closest in time.
Federated optimization in heterogeneous networks
Tian Li, Anit Kumar Sahu, Manzil Zaheer, Maziar Sanjabi, Ameet Talwalkar, and Virginia Smith · 2020
Closest in time.
FedSplit: An algorithmic framework for fast federated optimization
Reese Pathak and Martin J. Wainwright · 2020
Closest in time.
FedPAQ: A Communication-Efficient Federated Learning Method with Periodic Averaging and Quantization
Amirhossein Reisizadeh, Aryan Mokhtari, Hamed Hassani, Ali Jadbabaie, and Ramtin Pedarsani · 2020
Closest in time.
SlowMo: Improving communication-efficient distributed SGD with slow momentum
Jianyu Wang, Vinayak Tantia, Nicolas Ballas, and Michael Rabbat · 2020
Closest in time.
Is Local SGD Better than Minibatch SGD?
Blake Woodworth, Kumar Kshitij Patel, Sebastian U. Stich, Zhen Dai, Brian Bullins, H. Brendan McMahan, Ohad Shamir, and Nathan Srebro · 2020
Closest in time.
Veridical data science
Bin Yu and Karl Kumbier · 2020
Closest in time.