Fetching the paper…
Reading the bibliography…
A major obstacle to achieving global convergence in distributed and federated learning is the misalignment of gradients across clients, or mini-batches due to heterogeneity and stochasticity of the distributed data.
Stiffness: A New Perspective on Generalization in Neural Networks
Fort, S.; Nowak, P. K.; Jastrzebski, S.; and Narayanan, S. 2020 · 1901
Earlier work this paper cites.
Measuring the Effects of Non-Identical Data Distribution for Federated Visual Classification
Hsu, T.-M. H.; Qi, H.; and Brown, M. 2019a · 1909
Earlier work this paper cites.
Measuring the effects of non-identical data distribution for federated visual classification
Hsu, T.-M. H.; Qi, H.; and Brown, M. 2019b · 1909
Earlier work this paper cites.
Advances and open problems in federated learning
Kairouz, P.; McMahan, H. B.; Avent, B.; Bellet, A.; Bennis, M.; Bhagoji, A. N.; Bonawitz, K.; Charles, Z.; Cormode, G.; Cummings, R.; et al. 2019 · 1912
Earlier work this paper cites.
A Stochastic Approximation Method
Robbins, H.; and Monro, S. 1951 · 1951
Earlier work this paper cites.
Gradient Diversity: a Key Ingredient for Scalable Distributed Learning
Yin, D.; Pananjady, A.; Lam, M.; Papailiopoulos, D.; Ramchandran, K.; and Bartlett, P. 2018 · 2007
Earlier work this paper cites.
Making Coherence Out of Nothing At All: Measuring the Evolution of Gradient Alignment
Chatterjee, S.; and Zielinski, P. 2020 · 2008
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A.; Hinton, G.; et al. 2009 · 2009
Earlier work this paper cites.
Parallelized Stochastic Gradient Descent
Zinkevich, M. A.; Weimer, M.; Smola, A.; and Li, L. 2010 · 2010
Earlier work this paper cites.
Large Scale Distributed Deep Networks
Dean, J.; Corrado, G.; Monga, R.; Chen, K.; Devin, M.; Mao, M.; Ranzato, M. a.; Senior, A.; Tucker, P.; Yang, K.; Le, Q.; and Ng, A. 2012 · 2012
Earlier work this paper cites.
Accelerating Stochastic Gradient Descent using Predictive Variance Reduction
Johnson, R.; and Zhang, T. 2013 · 2013
Earlier work this paper cites.
Fast Convergence of Stochastic Gradient Descent under a Strong Growth Condition
Schmidt, M.; and Roux, N. L. 2013 · 2013
Earlier work this paper cites.
Federated Optimization: Distributed Machine Learning for On-Device Intelligence
Konečný, J.; McMahan, H. B.; Ramage, D.; and Richtárik, P. 2016 · 2016
Earlier work this paper cites.
EMNIST: Extending MNIST to handwritten letters
Cohen, G.; Afshar, S.; Tapson, J.; and Van Schaik, A. 2017 · 2017
Cited alongside, same era.
Sharp Minima Can Generalize For Deep Nets
Dinh, L.; Pascanu, R.; Bengio, S.; and Bengio, Y. 2017 · 2017
Cited alongside, same era.
On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima
Keskar, N. S.; Mudigere, D.; Nocedal, J.; Smelyanskiy, M.; and Tang, P. T. P. 2017 · 2017
Cited alongside, same era.
Stochastic Gradient Descent as Approximate Bayesian Inference
Mandt, S.; Hoffman, M. D.; and Blei, D. M. 2017 · 2017
Cited alongside, same era.
Leaf: A benchmark for federated settings
Caldas, S.; Duddu, S. M. K.; Wu, P.; Li, T.; Konečnỳ, J.; McMahan, H. B.; Smith, V.; and Talwalkar, A. 2018 · 2018
Cited alongside, same era.
Measuring the Effects of Data Parallelism on Neural Network Training
Shallue, C. J.; Lee, J.; Antognini, J.; Sohl-Dickstein, J.; Frostig, R.; and Dahl, G. E. 2019 · 2019
Later among the works it cites.
Coherent Gradients: An Approach to Understanding Generalization in Gradient Descent-based Optimization
Chatterjee, S. 2020 · 2020
Later among the works it cites.
SCAFFOLD: Stochastic Controlled Averaging for Federated Learning
Karimireddy, S. P.; Kale, S.; Mohri, M.; Reddi, S.; Stich, S.; and Suresh, A. T. 2020 · 2020
Later among the works it cites.
Tighter theory for local SGD on identical and heterogeneous data
Khaled, A.; Mishchenko, K.; and Richtárik, P. 2020 · 2020
Later among the works it cites.
Federated Optimization in Heterogeneous Networks
Li, T.; Sahu, A. K.; Zaheer, M.; Sanjabi, M.; Talwalkar, A.; and Smith, V. 2020 · 2020
Later among the works it cites.
Accelerating federated learning via momentum gradient descent
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chaudhari, P.; and Soatto, S. 2018 · 2018
Cited alongside, same era.
Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour
Goyal, P.; Dollár, P.; Girshick, R.; Noordhuis, P.; Wesolowski, L.; Kyrola, A.; Tulloch, A.; Jia, Y.; and He, K. 2018 · 2018
Cited alongside, same era.
Neural Tangent Kernel: Convergence and Generalization in Neural Networks
Jacot, A.; Gabriel, F.; and Hongler, C. 2018 · 2018
Cited alongside, same era.
Three Factors Influencing Minima in SGD
Jastrzębski, S.; Kenton, Z.; Arpit, D.; Ballas, N.; Fischer, A.; Bengio, Y.; and Storkey, A. 2018 · 2018
Cited alongside, same era.
The Power of Interpolation: Understanding the Effectiveness of SGD in Modern Over-parametrized Learning
Ma, S.; Bassily, R.; and Belkin, M. 2018 · 2018
Cited alongside, same era.
On First-Order Meta-Learning Algorithms
Nichol, A.; Achiam, J.; and Schulman, J. 2018 · 2018
Cited alongside, same era.
A Bayesian Perspective on Generalization and Stochastic Gradient Descent
Smith, S. L.; and Le, Q. V. 2018 · 2018
Cited alongside, same era.
Liu, W.; Chen, L.; Chen, Y.; and Zhang, W. 2020 · 2020
Later among the works it cites.
Distributed Gradient Methods for Convex Machine Learning Problems in Networks: Distributed Optimization
Nedic, A. 2020 · 2020
Later among the works it cites.
Is Local SGD Better than Minibatch SGD?
Woodworth, B.; Patel, K. K.; Stich, S.; Dai, Z.; Bullins, B.; Mcmahan, B.; Shamir, O.; and Srebro, N. 2020 · 2020
Later among the works it cites.
Federated Learning Based on Dynamic Regularization
Acar, D. A. E.; Zhao, Y.; Matas, R.; Mattina, M.; Whatmough, P.; and Saligrama, V. 2021 · 2021
Closest in time.
Implicit Gradient Regularization
Barrett, D.; and Dherin, B. 2021 · 2021
Closest in time.
Advances and Open Problems in Federated Learning
Kairouz, P.; and McMahan, H. B. 2021 · 2021
Closest in time.
Adaptive Federated Optimization
Reddi, S. J.; Charles, Z.; Zaheer, M.; Garrett, Z.; Rush, K.; Konečný, J.; Kumar, S.; and McMahan, H. B. 2021 · 2021
Closest in time.
On the Origin of Implicit Regularization in Stochastic Gradient Descent
Smith, S. L.; Dherin, B.; Barrett, D.; and De, S. 2021 · 2021
Closest in time.