Fetching the paper…
Reading the bibliography…
Federated Learning is a distributed learning paradigm with two key challenges that differentiate it from traditional distributed optimization: (1) significant variability in terms of the systems characteristics on each device in the network (systems heterogeneity), and (2) non-identically distributed data across the network (statistical heterogeneity).
Approximate solution of systems of linear equations
Kaczmarz, S · 1993
Earlier work this paper cites.
Gradient-based learning applied to document recognition
LeCun, Y., Bottou, L., Bengio, Y., and Haffner, P · 1998
Earlier work this paper cites.
On early stopping in gradient descent learning
Yao, Y., Rosasco, L., and Caponnetto, A · 2007
Earlier work this paper cites.
Twitter sentiment classification using distant supervision
Go, A., Bhayani, R., and Huang, L · 2009
Earlier work this paper cites.
A randomized kaczmarz algorithm with exponential convergence
Strohmer, T. and Vershynin, R · 2009
Earlier work this paper cites.
Distributed optimization and statistical learning via the alternating direction method of multipliers
Boyd, S., Parikh, N., Chu, E., Peleato, B., and Eckstein, J · 2010
Earlier work this paper cites.
Large scale distributed deep networks
Dean, J., Corrado, G., Monga, R., Chen, K., Devin, M., Le, Q. V., Mao, M., Ranzato, M., Senior, A., Tucker, P., Yang, K., and Ng, A · 2012
Earlier work this paper cites.
Optimal Distributed Online Prediction Using Mini-Batches
Dekel, O., Gilad-Bachrach, R., Shamir, O., and Xiao, L · 2012
Earlier work this paper cites.
Stochastic first-and zeroth-order methods for nonconvex stochastic programming
Ghadimi, S. and Lan, G · 2013
Earlier work this paper cites.
Fast convergence of stochastic gradient descent under a strong growth condition
Schmidt, M. and Roux, N. L · 2013
Earlier work this paper cites.
Communication-efficient algorithms for statistical optimization
Zhang, Y., Duchi, J. C., and Wainwright, M. J · 2013
Earlier work this paper cites.
Glove: Global vectors for word representation
Pennington, J., Socher, R., and Manning, C · 2014
Earlier work this paper cites.
Communication-efficient distributed optimization using an approximate newton-type method
Shamir, O., Srebro, N., and Zhang, T · 2014
Earlier work this paper cites.
Deep learning with elastic averaging sgd
Zhang, S., Choromanska, A. E., and LeCun, Y · 2015
Earlier work this paper cites.
Tensorflow: A system for large-scale machine learning
Abadi, M., Barham, P., Chen, J., Chen, Z., Davis, A., Dean, J., Devin, M., Ghemawat, S., Irving, G., Isard, M., Kudlur, M. K., Levenberg, J., Monga, R., Moore, S., Murray, D. G., Steiner, B., Tucker, P., Vasudevan, V., Warden, P., Wicke, M., Yu, Y., and Zheng, X · 2016
Earlier work this paper cites.
Aide: Fast and communication efficient distributed optimization
Reddi, S. J., Konečnỳ, J., Richtárik, P., Póczós, B., and Smola, A · 2016
Cited alongside, same era.
Distributed coordinate descent method for learning with big data
Richtárik, P. and Takáč, M · 2016
Cited alongside, same era.
Emnist: an extension of mnist to handwritten letters
Cohen, G., Afshar, S., Tapson, J., and van Schaik, A · 2017
Cited alongside, same era.
Communication-efficient learning of deep networks from decentralized data
McMahan, H. B., Moore, E., Ramage, D., Hampson, S., and Arcas, B. A. y · 2017
Cited alongside, same era.
Federated multi-task learning
Smith, V., Chiang, C.-K., Sanjabi, M., and Talwalkar, A. S · 2017
Cited alongside, same era.
Parallel restarted sgd for non-convex optimization with faster convergence and less communication
Yu, H., Yang, S., and Zhu, S · 2018
Closest in time.
Federated learning with non-iid data
Zhao, Y., Li, M., Lai, L., Suda, N., Civin, D., and Chandra, V · 2018
Closest in time.
On the convergence properties of a k k -step averaging stochastic gradient descent algorithm for nonconvex optimization
Zhou, F. and Cong, G · 2018
Closest in time.
Towards federated learning at scale: system design
Bonawitz, K., Eichner, H., Grieskamp, W., Huba, D., Ingerman, A., Ivanov, V., Kiddon, C., Konecny, J., Mazzocchi, S., McMahan, H. B., Overveldt, T. V., Petrou, D., Ramage, D., and Roselander, J · 2019
Closest in time.
On the linear speedup analysis of communication efficient momentum sgd for distributed non-convex optimization
Hao, Y., Rong, J., and Sen, Y · 2019
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
How to make the gradients small stochastically: Even faster convex and nonconvex sgd
Allen-Zhu, Z · 2018
Cited alongside, same era.
Leaf: A benchmark for federated settings
Caldas, S., Wu, P., Li, T., Konečnỳ, J., McMahan, H. B., Smith, V., and Talwalkar, A · 2018
Cited alongside, same era.
Loadaboost: Loss-based adaboost federated machine learning on medical data
Huang, L., Yin, Y., Fu, Z., Zhang, S., Deng, H., and Liu, D · 2018
Cited alongside, same era.
Jeong, E., Oh, S., Kim, H., Park, J., Bennis, M., and Kim, S.-L · 2018
Cited alongside, same era.
A linear speedup analysis of distributed deep learning with sparse and quantized communication
Jiang, P. and Agrawal, G · 2018
Cited alongside, same era.
Cocoa: A general framework for communication-efficient distributed optimization
Smith, V., Forte, S., Ma, C., Takac, M., Jordan, M. I., and Jaggi, M · 2018
Cited alongside, same era.
Wang, J. and Joshi, G · 2018
Cited alongside, same era.
Adaptive gradient-based meta-learning methods
Khodak, M., Balcan, M.-F. F., and Talwalkar, A. S · 2019
Closest in time.
Federated learning: Challenges, methods, and future directions
Li, T., Sahu, A., Talwalkar, A., and Smith, V · 2019
Closest in time.
Local sgd converges fast and communicates little
Stich, S. U · 2019
Closest in time.
Fast and faster convergence of sgd for over-parameterized models (and an accelerated perceptron)
Vaswani, S., Bach, F., and Schmidt, M · 2019
Closest in time.
Adaptive federated learning in resource constrained edge computing systems
Wang, S., Tuor, T., Salonidis, T., Leung, K. K., Makaya, C., He, T., and Chan, K · 2019
Closest in time.
Efficient meta learning via minibatch proximal update
Zhou, P., Yuan, X., Xu, H., Yan, S., and Feng, J · 2019
Closest in time.
Unraveling meta-learning: Understanding feature representations for few-shot tasks
Goldblum, M., Reich, S., Fowl, L., Ni, R., Cherepanova, V., and Goldstein, T · 2020
Closest in time.
Feddane: A federated newton-type method
Li, T., Sahu, A. K., Zaheer, M., Sanjabi, M., Talwalkar, A., and Smith, V · 2020
Closest in time.
Don’t use large mini-batches, use local sgd
Lin, T., Stich, S. U., and Jaggi, M · 2020
Closest in time.