Fetching the paper…
Reading the bibliography…
We propose a new optimization formulation for training federated learning models.
L-SVRG and L-Katyusha with arbitrary sampling
Qian, X., Qu, Z., and Richtárik, P · 1906
Earlier work this paper cites.
Is local sgd better than minibatch sgd?
Woodworth, B., Patel, K. K., Stich, S. U., Dai, Z., Bullins, B., McMahan, H. B., Shamir, O., and Srebro, N · 2002
Earlier work this paper cites.
Introductory lectures on convex optimization: a basic course (Applied Optimization)
Nesterov, Y · 2004
Earlier work this paper cites.
Minibatch vs local sgd for heterogeneous distributed learning
Woodworth, B., Patel, K. K., and Srebro, N · 2006
Earlier work this paper cites.
A convex formulation for learning task relationships in multi-task learning
Zhang, Y. and Yeung, D.-Y · 2010
Earlier work this paper cites.
LibSVM: A library for support vector machines
Chang, C.-C. and Lin, C.-J · 2011
Earlier work this paper cites.
Large scale distributed deep networks
Dean, J., Corrado, G., Monga, R., Chen, K., Devin, M., Mao, M., Senior, A., Tucker, P., Yang, K., Le, Q. V., and et al · 2012
Earlier work this paper cites.
Accelerating stochastic gradient descent using predictive variance reduction
Johnson, R. and Zhang, T · 2013
Earlier work this paper cites.
SAGA: a fast incremental gradient method with support for non-strongly convex composite objectives
Defazio, A., Bach, F., and Lacoste-Julien, S · 2014
Earlier work this paper cites.
A proximal stochastic gradient method with progressive variance reduction
Xiao, L. and Zhang, T · 2014
Earlier work this paper cites.
Variance reduced stochastic gradient descent with neighbors
Hofmann, T., Lucchi, A., Lacoste-Julien, S., and McWilliams, B · 2015
Earlier work this paper cites.
Deep learning with elastic averaging sgd
Zhang, S., Choromanska, A. E., and LeCun, Y · 2015
Earlier work this paper cites.
Stochastic optimization with importance sampling for regularized loss minimization
Zhao, P. and Zhang, T · 2015
Earlier work this paper cites.
Federated learning of deep networks using model averaging
McMahan, B., Moore, E., Ramage, D., and Agüera y Arcas, B · 2016
Earlier work this paper cites.
Coordinate descent with arbitrary sampling II: Expected separable overapproximation
Qu, Z. and Richtárik, P · 2016
Earlier work this paper cites.
AIDE: fast and communication efficient distributed optimization
Reddi, S. J., Konečný, J., Richtárik, P., Póczos, B., and Smola, A · 2016
Earlier work this paper cites.
Parallel coordinate descent methods for big data optimization
Richtárik, P. and Takáč, M · 2016
Cited alongside, same era.
Katyusha: The first direct acceleration of stochastic gradient methods
Allen-Zhu, Z · 2017
Cited alongside, same era.
Practical secure aggregation for privacy-preserving machine learning
Bonawitz, K., Ivanov, V., Kreuter, B., Marcedone, A., McMahan, H. B., Patel, S., Ramage, D., Segal, A., and Seth, K · 2017
Cited alongside, same era.
Model-agnostic meta-learning for fast adaptation of deep networks
Finn, C., Abbeel, P., and Levine, S · 2017
Cited alongside, same era.
Distributed multi-task relationship learning
Liu, S., Pan, S. J., and Ho, Q · 2017
Cited alongside, same era.
Communication-efficient learning of deep networks from decentralized data
McMahan, H. B., Moore, E., Ramage, D., Hampson, S., and Agüera y Arcas, B · 2017
One method to rule them all: Variance reduction for data, parameters and many new methods
Hanzely, F. and Richtárik, P · 2019
Later among the works it cites.
Advances and open problems in federated learning
Kairouz, P., McMahan, H. B., and et al · 2019
Later among the works it cites.
SCAFFOLD: stochastic controlled averaging for on-device federated learning
Karimireddy, S. P., Kale, S., Mohri, M., Reddi, S. J., Stich, S. U., and Suresh, A. T · 2019
Later among the works it cites.
First analysis of local GD on heterogeneous data
Khaled, A., Mishchenko, K., and Richtárik, P · 2019
Later among the works it cites.
Adaptive gradient-based meta-learning methods
Khodak, M., Balcan, M.-F., and Talwalkar, A · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Federated multi-task learning
Smith, V., Chiang, C.-K., Sanjabi, M., and Talwalkar, A. S · 2017
Cited alongside, same era.
Stochastic quasi-gradient methods: variance reduction via Jacobian sketching
Gower, R. M., Richtárik, P., and Bach, F · 2018
Cited alongside, same era.
Communication-efficient algorithms for decentralized and stochastic optimization
Lan, G., Lee, S., and Zhou, Y · 2018
Cited alongside, same era.
Distributed stochastic multi-task learning with graph regularization
Wang, W., Wang, J., Kolar, M., and Srebro, N · 2018
Cited alongside, same era.
Federated learning with non-iid data
Zhao, Y., Li, M., Lai, L., Suda, N., Civin, D., and Chandra, V · 2018
Cited alongside, same era.
Variational federated multi-task learning
Corinzia, L. and Buhmann, J. M · 2019
Cited alongside, same era.
Federated learning: challenges, methods, and future directions
Li, T., Sahu, A. K., Talwalkar, A., and Smith, V · 2019
Later among the works it cites.
Variance reduced local SGD with lower communication complexity
Liang, X., Shen, S., Liu, J., Pan, Z., Chen, E., and Cheng, Y · 2019
Later among the works it cites.
A stochastic decoupling method for minimizing the sum of smooth and non-smooth functions
Mishchenko, K. and Richtárik, P · 2019
Later among the works it cites.
Private federated learning with domain adaptation
Peterson, D., Kanani, P., and Marathe, V. J · 2019
Later among the works it cites.
Federated variance-reduced stochastic gradient descent with robustness to byzantine attacks
Wu, Z., Ling, Q., Chen, T., and Giannakis, G. B · 2019
Later among the works it cites.
A unified theory of sgd: Variance reduction, sampling, quantization and coordinate descent
Gorbunov, E., Hanzely, F., and Richtárik, P · 2020
Closest in time.
Tighter theory for local SGD on identical and heterogeneous data
Khaled, A., Mishchenko, K., and Richtárik, P · 2020
Closest in time.
Don’t jump through hoops and remove those loops: SVRG and Katyusha are better without the outer loop
Kovalev, D., Horváth, S., and Richtárik, P · 2020
Closest in time.
Local SGD converges fast and communicates little
Stich, S. U · 2020
Closest in time.
Overlap local-sgd: An algorithmic approach to hide communication delays in distributed sgd
Wang, J., Liang, H., and Joshi, G · 2020
Closest in time.