Fetching the paper…
Reading the bibliography…
In federated learning, communication cost is often a critical bottleneck to scale up distributed optimization algorithms to collaboratively learn a model from millions of devices with potentially unreliable or limited communication and heterogeneous data distributions.
Learning with kernels
Bernhard Schölkopf and Alexander J Smola · 2002
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
Stochastic gradient descent tricks
Léon Bottou · 2012
Earlier work this paper cites.
Stochastic gradient descent, weighted sampling, and the randomized kaczmarz algorithm
Deanna Needell, Rachel Ward, and Nati Srebro · 2014
Earlier work this paper cites.
1-bit stochastic gradient descent and its application to data-parallel distributed training of speech dnns
Frank Seide, Hao Fu, Jasha Droppo, Gang Li, and Dong Yu · 2014
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Linear convergence of gradient and proximal-gradient methods under the polyak-łojasiewicz condition
Hamed Karimi, Julie Nutini, and Mark Schmidt · 2016
Earlier work this paper cites.
Communication-efficient learning of deep networks from decentralized data
H Brendan McMahan, Eider Moore, Daniel Ramage, Seth Hampson, et al · 2016
Earlier work this paper cites.
Parallel sgd: When does averaging help?
Jian Zhang, Christopher De Sa, Ioannis Mitliagkas, and Christopher Ré · 2016
Earlier work this paper cites.
Sparse communication for distributed gradient descent
Alham Fikri Aji and Kenneth Heafield · 2017
Earlier work this paper cites.
Qsgd: Communication-efficient sgd via gradient quantization and encoding
Dan Alistarh, Demjan Grubic, Jerry Li, Ryota Tomioka, and Milan Vojnovic · 2017
Earlier work this paper cites.
Deep gradient compression: Reducing the communication bandwidth for distributed training
Yujun Lin, Song Han, Huizi Mao, Yu Wang, and William J Dally · 2017
Earlier work this paper cites.
Federated multi-task learning
Virginia Smith, Chao-Kai Chiang, Maziar Sanjabi, and Ameet S Talwalkar · 2017
Earlier work this paper cites.
Distributed mean estimation with limited communication
Ananda Theertha Suresh, Felix X Yu, Sanjiv Kumar, and H Brendan McMahan · 2017
Earlier work this paper cites.
Terngrad: Ternary gradients to reduce communication in distributed deep learning
Wei Wen, Cong Xu, Feng Yan, Chunpeng Wu, Yandan Wang, Yiran Chen, and Hai Li · 2017
Earlier work this paper cites.
Fashion-mnist: a novel image dataset for benchmarking machine learning algorithms, 2017
Han Xiao, Kashif Rasul, and Roland Vollgraf · 2017
Earlier work this paper cites.
The convergence of sparsified gradient methods
Dan Alistarh, Torsten Hoefler, Mikael Johansson, Nikola Konstantinov, Sarit Khirirat, and Cédric Renggli · 2018
Earlier work this paper cites.
signsgd: Compressed optimisation for non-convex problems
Jeremy Bernstein, Yu-Xiang Wang, Kamyar Azizzadenesheli, and Animashree Anandkumar · 2018
Earlier work this paper cites.
Optimization methods for large-scale machine learning
Léon Bottou, Frank E Curtis, and Jorge Nocedal · 2018
Earlier work this paper cites.
Leaf: A benchmark for federated settings
Sebastian Caldas, Peter Wu, Tian Li, Jakub Konečnỳ, H Brendan McMahan, Virginia Smith, and Ameet Talwalkar · 2018
Earlier work this paper cites.
Don’t use large mini-batches, use local sgd
Tao Lin, Sebastian U Stich, Kumar Kshitij Patel, and Martin Jaggi · 2018
Earlier work this paper cites.
Deep gradient compression: Reducing the communication bandwidth for distributed training
Yujun Lin, Song Han, Huizi Mao, Yu Wang, and Bill Dally · 2018
Cited alongside, same era.
Deep gradient compression: Reducing the communication bandwidth for distributed training
Yujun Lin, Song Han, Huizi Mao, Yu Wang, and Bill Dally · 2018
Cited alongside, same era.
Quantized decentralized consensus optimization
Amirhossein Reisizadeh, Aryan Mokhtari, Hamed Hassani, and Ramtin Pedarsani · 2018
Cited alongside, same era.
Local sgd converges fast and communicates little
Sebastian U Stich · 2018
Cited alongside, same era.
Sparsified sgd with memory
Sebastian U Stich, Jean-Baptiste Cordonnier, and Martin Jaggi · 2018
Cited alongside, same era.
Communication compression for decentralized training
On the convergence of fedavg on non-iid data
Xiang Li, Kaixuan Huang, Wenhao Yang, Shusen Wang, and Zhihua Zhang · 2019
Later among the works it cites.
Variance reduced local sgd with lower communication complexity
Xianfeng Liang, Shuheng Shen, Jingchang Liu, Zhen Pan, Enhong Chen, and Yifei Cheng · 2019
Later among the works it cites.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al · 2019
Later among the works it cites.
Robust and communication-efficient collaborative learning
Amirhossein Reisizadeh, Hossein Taheri, Aryan Mokhtari, Hamed Hassani, and Ramtin Pedarsani · 2019
Later among the works it cites.
Sebastian U Stich and Sai Praneeth Karimireddy · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hanlin Tang, Shaoduo Gan, Ce Zhang, Tong Zhang, and Ji Liu · 2018
Cited alongside, same era.
Jianyu Wang and Gauri Joshi · 2018
Cited alongside, same era.
Gradient sparsification for communication-efficient distributed optimization
Jianqiao Wangni, Jialei Wang, Ji Liu, and Tong Zhang · 2018
Cited alongside, same era.
Error compensated quantized sgd and its applications to large-scale distributed optimization
Jiaxiang Wu, Weidong Huang, Junzhou Huang, and Tong Zhang · 2018
Cited alongside, same era.
Parallel restarted sgd for non-convex optimization with faster convergence and less communication
Hao Yu, Sen Yang, and Shenghuo Zhu · 2018
Cited alongside, same era.
Fan Zhou and Guojing Cong · 2018
Cited alongside, same era.
Qsparse-local-sgd: Distributed sgd with quantization, sparsification and local computations
Debraj Basu, Deepesh Data, Can Karakus, and Suhas Diggavi · 2019
Cited alongside, same era.
Later among the works it cites.
Matcha: Speeding up decentralized sgd via matching decomposition sampling
Jianyu Wang, Anit Kumar Sahu, Zhouyi Yang, Gauri Joshi, and Soummya Kar · 2019
Later among the works it cites.
Parallel restarted sgd with faster convergence and less communication: Demystifying why model averaging works for deep learning
Hao Yu, Sen Yang, and Shenghuo Zhu · 2019
Later among the works it cites.
Adaptive personalized federated learning
Yuyang Deng, Mohammad Mahdi Kamani, and Mehrdad Mahdavi · 2020
Closest in time.
Personalized federated learning: A meta-learning approach
Alireza Fallah, Aryan Mokhtari, and Asuman Ozdaglar · 2020
Closest in time.
Stochastic-sign sgd for federated learning with theoretical guarantees
Richeng Jin, Yufan Huang, Xiaofan He, Huaiyu Dai, and Tianfu Wu · 2020
Closest in time.
Tighter theory for local sgd on identical and heterogeneous data
A Khaled, K Mishchenko, and P Richtárik · 2020
Closest in time.
A unified theory of decentralized sgd with changing topology and local updates
Anastasia Koloskova, Nicolas Loizou, Sadra Boreiri, Martin Jaggi, and Sebastian U Stich · 2020
Closest in time.
Federated learning: Challenges, methods, and future directions
Tian Li, Anit Kumar Sahu, Ameet Talwalkar, and Virginia Smith · 2020
Closest in time.
Acceleration for compressed gradient descent in distributed and federated optimization
Zhize Li, Dmitry Kovalev, Xun Qian, and Peter Richtárik · 2020
Closest in time.
A collaborative learning framework via federated meta-learning
Sen Lin, Guang Yang, and Junshan Zhang · 2020
Closest in time.
Three approaches for personalization with applications to federated learning
Yishay Mansour, Mehryar Mohri, Jae Ro, and Ananda Theertha Suresh · 2020
Closest in time.
Adaptive federated optimization
Sashank Reddi, Zachary Charles, Manzil Zaheer, Zachary Garrett, Keith Rush, Jakub Konečnỳ, Sanjiv Kumar, and H Brendan McMahan · 2020
Closest in time.
Fedpaq: A communication-efficient federated learning method with periodic averaging and quantization
Amirhossein Reisizadeh, Aryan Mokhtari, Hamed Hassani, Ali Jadbabaie, and Ramtin Pedarsani · 2020
Closest in time.
Squarm-sgd: Communication-efficient momentum sgd for decentralized optimization
Navjot Singh, Deepesh Data, Jemin George, and Suhas Diggavi · 2020
Closest in time.
Is local sgd better than minibatch sgd?
Blake Woodworth, Kumar Kshitij Patel, Sebastian U Stich, Zhen Dai, Brian Bullins, H Brendan McMahan, Ohad Shamir, and Nathan Srebro · 2020
Closest in time.
Fedpd: A federated learning framework with optimal rates and adaptivity to non-iid data
Xinwei Zhang, Mingyi Hong, Sairaj Dhople, Wotao Yin, and Yang Liu · 2020
Closest in time.