Fetching the paper…
Reading the bibliography…
Federated Averaging (FedAvg) has emerged as the algorithm of choice for federated learning due to its simplicity and low communication cost.
Federated learning of out-of-vocabulary words
Chen, M., Mathews, R., Ouyang, T., and Beaufays, F · 1903
Earlier work this paper cites.
Fair resource allocation in federated learning
Li, T., Sanjabi, M., and Smith, V · 1905
Earlier work this paper cites.
On the convergence of FedAvg on non-iid data
Li, X., Huang, K., Yang, W., Wang, S., and Zhang, Z · 1907
Earlier work this paper cites.
Fast incremental method for smooth nonconvex optimization
Reddi, S. J., Sra, S., Póczos, B., and Smola, A · 1977
Earlier work this paper cites.
Parallelized stochastic gradient descent
Zinkevich, M., Weimer, M., Li, L., and Smola, A. J · 2010
Earlier work this paper cites.
Differentially private empirical risk minimization
Chaudhuri, K., Monteleoni, C., and Sarwate, A. D · 2011
Earlier work this paper cites.
Large scale distributed deep networks
Dean, J., Corrado, G., Monga, R., Chen, K., Devin, M., Mao, M., Ranzato, M., Senior, A., Tucker, P., Yang, K., Le, Q. V., and Ng, A. Y · 2012
Earlier work this paper cites.
Private convex empirical risk minimization and high-dimensional regression
Kifer, D., Smith, A., and Thakurta, A · 2012
Earlier work this paper cites.
Monte Carlo methods in financial engineering , volume 53
Glasserman, P · 2013
Earlier work this paper cites.
Accelerating stochastic gradient descent using predictive variance reduction
Johnson, R. and Zhang, T · 2013
Earlier work this paper cites.
Private empirical risk minimization: Efficient algorithms and tight error bounds
Bassily, R., Smith, A., and Thakurta, A · 2014
Earlier work this paper cites.
SAGA: A fast incremental gradient method with support for non-strongly convex composite objectives
Defazio, A., Bach, F., and Lacoste-Julien, S · 2014
Earlier work this paper cites.
Communication-efficient distributed optimization using an approximate newton-type method
Shamir, O., Srebro, N., and Zhang, T · 2014
Earlier work this paper cites.
Communication complexity of distributed convex learning and optimization
Arjevani, Y. and Shamir, O · 2015
Earlier work this paper cites.
Lee, J. D., Lin, Q., Ma, T., and Yang, T · 2015
Earlier work this paper cites.
EXTRA: An exact first-order algorithm for decentralized consensus optimization
Shi, W., Ling, Q., Wu, G., and Yin, W · 2015
Earlier work this paper cites.
Deep learning with elastic averaging sgd
Zhang, S., Choromanska, A. E., and LeCun, Y · 2015
Earlier work this paper cites.
Firecaffe: near-linear acceleration of deep neural network training on compute clusters
Iandola, F. N., Moskewicz, M. W., Ashraf, K., and Keutzer, K · 2016
Earlier work this paper cites.
Linear convergence of gradient and proximal-gradient methods under the polyak-łojasiewicz condition
Karimi, H., Nutini, J., and Schmidt, M · 2016
Earlier work this paper cites.
A geometrically convergent method for distributed optimization over time-varying graphs
Nedich, A., Olshevsky, A., and Shi, W · 2016
Earlier work this paper cites.
Practical secure aggregation for privacy-preserving machine learning
Bonawitz, K., Ivanov, V., Kreuter, B., Marcedone, A., McMahan, H. B., Patel, S., Ramage, D., Segal, A., and Seth, K · 2017
Earlier work this paper cites.
Emnist: Extending mnist to handwritten letters
Cohen, G., Afshar, S., Tapson, J., and Van Schaik, A · 2017
Earlier work this paper cites.
Accurate, large minibatch SGD: Training imagenet in 1 hour
Goyal, P., Dollár, P., Girshick, R., Noordhuis, P., Wesolowski, L., Kyrola, A., Tulloch, A., Jia, Y., and He, K · 2017
Earlier work this paper cites.
Less than a single pass: Stochastically controlled stochastic gradient
Lei, L. and Jordan, M · 2017
Cited alongside, same era.
Communication-efficient learning of deep networks from decentralized data
McMahan, B., Moore, E., Ramage, D., Hampson, S., and y Arcas, B. A · 2017
Cited alongside, same era.
Sarah: A novel method for machine learning problems using stochastic recursive gradient
Nguyen, L. M., Liu, J., Scheinberg, K., and Takáč, M · 2017
Cited alongside, same era.
Minimizing finite sums with the stochastic average gradient
Schmidt, M., Le Roux, N., and Bach, F · 2017
Cited alongside, same era.
Distributed mean estimation with limited communication
Suresh, A. T., Yu, F. X., Kumar, S., and McMahan, H. B · 2017
Cited alongside, same era.
cpSGD: Communication-efficient and differentially-private distributed SGD
Agarwal, N., Suresh, A. T., Yu, F. X., Kumar, S., and McMahan, B · 2018
On the convergence of local descent methods in federated learning
Haddadpour, F. and Mahdavi, M · 2019
Closest in time.
One method to rule them all: Variance reduction for data, parameters and many new methods
Hanzely, F. and Richtárik, P · 2019
Closest in time.
Measuring the effects of non-identical data distribution for federated visual classification
Hsu, T.-M. H., Qi, H., and Brown, M · 2019
Closest in time.
Advances and open problems in federated learning
Kairouz, P., McMahan, H. B., Avent, B., Bellet, A., Bennis, M., Bhagoji, A. N., Bonawitz, K., Charles, Z., Cormode, G., Cummings, R., et al · 2019
Closest in time.
Error feedback fixes SignSGD and other gradient compression schemes
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Federated learning of predictive models from federated electronic health records
Brisimi, T. S., Chen, R., Mela, T., Olshevsky, A., Paschalidis, I. C., and Shi, W · 2018
Cited alongside, same era.
SPIDER: Near-optimal non-convex optimization via stochastic path-integrated differential estimator
Fang, C., Li, C. J., Lin, Z., and Zhang, T · 2018
Cited alongside, same era.
Federated learning for mobile keyboard prediction
Hard, A., Rao, K., Mathews, R., Beaufays, F., Augenstein, S., Eichner, H., Kiddon, C., and Ramage, D · 2018
Cited alongside, same era.
On the convergence of federated optimization in heterogeneous networks
Li, T., Sahu, A. K., Sanjabi, M., Zaheer, M., Talwalkar, A., and Smith, V · 2018
Cited alongside, same era.
Lectures on convex optimization , volume 137
Nesterov, Y · 2018
Cited alongside, same era.
Inexact SARAH algorithm for stochastic optimization
Nguyen, L. M., Scheinberg, K., and Takáč, M · 2018
Cited alongside, same era.
Karimireddy, S. P., Rebjock, Q., Stich, S. U., and Jaggi, M · 2019
Closest in time.
Kulunchakov, A. and Mairal, J · 2019
Closest in time.
Variance reduced local sgd with lower communication complexity
Liang, X., Shen, S., Liu, J., Pan, Z., Chen, E., and Cheng, Y · 2019
Closest in time.
Distributed learning with compressed gradient differences
Mishchenko, K., Gorbunov, E., Takáč, M., and Richtárik, P · 2019
Closest in time.
Mohri, M., Sivek, G., and Suresh, A. T · 2019
Closest in time.
Communication trade-offs for synchronized distributed SGD with large step size
Patel, K. K. and Dieuleveut, A · 2019
Closest in time.
Federated learning for emoji prediction in a mobile keyboard
Ramaswamy, S., Mathews, R., Rao, K., and Beaufays, F · 2019
Closest in time.
How good is sgd with random shuffling?
Safran, I. and Shamir, O · 2019
Closest in time.
Unified optimal analysis of the (stochastic) gradient method
Stich, S. U · 2019
Closest in time.
Stich, S. U. and Karimireddy, S. P · 2019
Closest in time.
Hybrid stochastic gradient descent algorithms for stochastic nonconvex optimization
Tran-Dinh, Q., Pham, N. H., Phan, D. T., and Nguyen, L. M · 2019
Closest in time.
Fast and faster convergence of sgd for over-parameterized models and an accelerated perceptron
Vaswani, S., Bach, F., and Schmidt, M · 2019
Closest in time.
Adaptive federated learning in resource constrained edge computing systems
Wang, S., Tuor, T., Salonidis, T., Leung, K. K., Makaya, C., He, T., and Chan, K · 2019
Closest in time.
Parallel restarted SGD with faster convergence and less communication: Demystifying why model averaging works for deep learning
Yu, H., Yang, S., and Zhu, S · 2019
Closest in time.
On convergence of distributed approximate newton methods: Globalization, sharper bounds and beyond
Yuan, X.-T. and Li, P · 2019
Closest in time.
Tighter theory for local SGD on indentical and heterogeneous data
Khaled, A., Mishchenko, K., and Richtárik, P · 2020
Closest in time.
Feddane: A federated newton-type method
Li, T., Sahu, A. K., Zaheer, M., Sanjabi, M., Talwalkar, A., and Smith, V · 2020
Closest in time.
Federated learning based on dynamic regularization
Acar, D. A. E., Zhao, Y., Navarro, R. M., Mattina, M., Whatmough, P. N., and Saligrama, V · 2021
Closest in time.