Fetching the paper…
Reading the bibliography…
We propose \texttt{FedGLOMO}, a novel federated learning (FL) algorithm with an iteration complexity of $\mathcal{O}(\epsilon^{-1.5})$ to converge to an $\epsilon$-stationary point (i.e., $\mathbb{E}[\|\nabla f(\bm{x})\|^2] \leq \epsilon$) for smooth non-convex functions -- under arbitrary client heterogeneity and compressed communication -- compared to the $\mathcal{O}(\epsilon^{-2})$ complexity of most prior works.
Parallelized stochastic gradient descent
Zinkevich, M., Weimer, M., Li, L., and Smola, A · 2010
Earlier work this paper cites.
Stochastic gradient descent tricks
Bottou, L · 2012
Earlier work this paper cites.
Qsgd: Communication-efficient sgd via gradient quantization and encoding
Alistarh, D., Grubic, D., Li, J., Tomioka, R., and Vojnovic, M · 2017
Earlier work this paper cites.
Deep gradient compression: Reducing the communication bandwidth for distributed training
Lin, Y., Han, S., Mao, H., Wang, Y., and Dally, W. J · 2017
Earlier work this paper cites.
Communication-efficient learning of deep networks from decentralized data
McMahan, B., Moore, E., Ramage, D., Hampson, S., and y Arcas, B. A · 2017
Earlier work this paper cites.
Distributed mean estimation with limited communication
Suresh, A. T., Felix, X. Y., Kumar, S., and McMahan, H. B · 2017
Earlier work this paper cites.
Fashion-mnist: a novel image dataset for benchmarking machine learning algorithms
Xiao, H., Rasul, K., and Vollgraf, R · 2017
Earlier work this paper cites.
The convergence of sparsified gradient methods
Alistarh, D., Hoefler, T., Johansson, M., Konstantinov, N., Khirirat, S., and Renggli, C · 2018
Earlier work this paper cites.
signsgd: Compressed optimisation for non-convex problems
Bernstein, J., Wang, Y.-X., Azizzadenesheli, K., and Anandkumar, A · 2018
Earlier work this paper cites.
Spider: Near-optimal non-convex optimization via stochastic path-integrated differential estimator
Fang, C., Li, C. J., Lin, Z., and Zhang, T · 2018
Earlier work this paper cites.
Federated optimization in heterogeneous networks
Li, T., Sahu, A. K., Zaheer, M., Sanjabi, M., Talwalkar, A., and Smith, V · 2018
Earlier work this paper cites.
Local sgd converges fast and communicates little
Stich, S. U · 2018
Earlier work this paper cites.
Sparsified sgd with memory
Stich, S. U., Cordonnier, J.-B., and Jaggi, M · 2018
Earlier work this paper cites.
Communication compression for decentralized training
Tang, H., Gan, S., Zhang, C., Zhang, T., and Liu, J · 2018
Earlier work this paper cites.
Wang, J., and Joshi, G · 2018
Cited alongside, same era.
Error compensated quantized sgd and its applications to large-scale distributed optimization
Wu, J., Huang, W., Huang, J., and Zhang, T · 2018
Cited alongside, same era.
Parallel restarted sgd for non-convex optimization with faster convergence and less communication
Yu, H., Yang, S., and Zhu, S · 2018
Cited alongside, same era.
Stochastic nested variance reduced gradient descent for nonconvex optimization
Zhou, D., Xu, P., and Gu, Q · 2018
Cited alongside, same era.
Tighter theory for local sgd on identical and heterogeneous data
Bayoumi, A. K. R., Mishchenko, K., and Richtárik, P · 2020
Closest in time.
Communication-efficient algorithms for decentralized optimization over directed graphs
Chen, Y., Hashemi, A., and Vikalo, H · 2020
Closest in time.
Federated learning with compression: Unified analysis and sharp guarantees
Haddadpour, F., Kamani, M. M., Mokhtari, A., and Mahdavi, M · 2020
Closest in time.
On the benefits of multiple gossip steps in communication-constrained decentralized optimization
Hashemi, A., Acharya, A., Das, R., Vikalo, H., Sanghavi, S., and Dhillon, I · 2020
Closest in time.
Faster on-device training using new federated momentum algorithm
Huo, Z., Yang, Q., Gu, B., Huang, L. C., et al · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Arjevani, Y., Carmon, Y., Duchi, J. C., Foster, D. J., Srebro, N., and Woodworth, B · 2019
Cited alongside, same era.
Qsparse-local-sgd: Distributed sgd with quantization, sparsification and local computations
Basu, D., Data, D., Karakus, C., and Diggavi, S · 2019
Cited alongside, same era.
Momentum-based variance reduction in non-convex sgd
Cutkosky, A., and Orabona, F · 2019
Cited alongside, same era.
Stochastic distributed learning with gradient quantization and variance reduction
Horváth, S., Kovalev, D., Mishchenko, K., Stich, S., and Richtárik, P · 2019
Cited alongside, same era.
Scaffold: Stochastic controlled averaging for federated learning
Karimireddy, S. P., Kale, S., Mohri, M., Reddi, S. J., Stich, S. U., and Suresh, A. T · 2019
Cited alongside, same era.
On the convergence of fedavg on non-iid data
Li, X., Huang, K., Yang, W., Wang, S., and Zhang, Z · 2019
Cited alongside, same era.
Variance reduced local sgd with lower communication complexity
Liang, X., Shen, S., Liu, J., Pan, Z., Chen, E., and Cheng, Y · 2019
Cited alongside, same era.
Communication trade-offs for synchronized distributed sgd with large step size
Patel, K. K., and Dieuleveut, A · 2019
Cited alongside, same era.
Closest in time.
Mime: Mimicking centralized stochastic algorithms in federated learning
Karimireddy, S. P., Jaggi, M., Kale, S., Mohri, M., Reddi, S. J., Stich, S. U., and Suresh, A. T · 2020
Closest in time.
A unified theory of decentralized sgd with changing topology and local updates
Koloskova, A., Loizou, N., Boreiri, S., Jaggi, M., and Stich, S · 2020
Closest in time.
An optimal hybrid variance-reduced algorithm for stochastic composite nonconvex optimization
Liu, D., Nguyen, L. M., and Tran-Dinh, Q · 2020
Closest in time.
Distributed stochastic gradient tracking methods
Pu, S., and Nedić, A · 2020
Closest in time.
Federated learning’s blessing: Fedavg has linear speedup
Qu, Z., Lin, K., Kalagnanam, J., Li, Z., Zhou, J., and Zhou, Z · 2020
Closest in time.
Adaptive federated optimization
Reddi, S., Charles, Z., Zaheer, M., Garrett, Z., Rush, K., Konečnỳ, J., Kumar, S., and McMahan, H. B · 2020
Closest in time.
Is local sgd better than minibatch sgd?
Woodworth, B., Patel, K. K., Stich, S. U., Dai, Z., Bullins, B., McMahan, H. B., Shamir, O., and Srebro, N · 2020
Closest in time.
Marina: Faster non-convex distributed learning with compression
Gorbunov, E., Burlachenko, K., Li, Z., and Richtárik, P · 2021
Closest in time.
Fedpaq: A communication-efficient federated learning method with periodic averaging and quantization
Reisizadeh, A., Mokhtari, A., Hassani, H., Jadbabaie, A., and Pedarsani, R · 2031
Closest in time.