Fetching the paper…
Reading the bibliography…
Federated learning (FL) is an emerging technique for training machine learning models using geographically dispersed data collected by local entities.
P. Auer, N. Cesa-Bianchi, Y. Freund, and R. E. Schapire, “The nonstochastic multiarmed bandit problem,” SIAM journal on computing , vol. 32, no. 1, pp. 48–77, 2002
2002
Earlier work this paper cites.
A. D. Flaxman, A. T. Kalai, A. T. Kalai, and H. B. McMahan, “Online convex optimization in the bandit setting: Gradient descent without a gradient,” in ACM-SIAM Symposium on Discrete Algorithms , 2005
2005
Earlier work this paper cites.
2009
Earlier work this paper cites.
S. Mannor and O. Shamir, “From bandits to experts: On the value of side-observations,” in NeurIPS , 2011, pp. 684–692
2011
Earlier work this paper cites.
S. Caron, B. Kveton, M. Lelarge, and S. Bhagat, “Leveraging side observations in stochastic bandits,” in UAI , 2012
2012
Earlier work this paper cites.
S. Gupta, A. Agrawal, K. Gopalakrishnan, and P. Narayanan, “Deep learning with limited numerical precision,” in ICML , 2015
2015
Earlier work this paper cites.
S. Bubeck, “Convex optimization: Algorithms and complexity,” Foundations and trends in Machine Learning , vol. 8, no. 3-4, 2015
2015
Earlier work this paper cites.
J. Konečný, H. B. McMahan, F. X. Yu et al. , “Federated learning: Strategies for improving communication efficiency,” in NeurIPS Workshop on Private Multi-Party Machine Learning , 2016
2016
Earlier work this paper cites.
E. Hazan et al. , “Introduction to online convex optimization,” Foundations and Trends® in Optimization , vol. 2, no. 3-4, pp. 157–325, 2016
2016
Earlier work this paper cites.
I. Goodfellow, Y. Bengio, and A. Courville, Deep Learning . MIT Press, 2016, http://www.deeplearningbook.org
2016
Earlier work this paper cites.
H. B. McMahan, E. Moore, D. Ramage, S. Hampson, and B. A. y Arcas, “Communication-efficient learning of deep networks from decentralized data,” in AISTATS , 2017
2017
Earlier work this paper cites.
H. Zhang, Z. Zheng, S. Xu et al. , “Poseidon: An efficient communication architecture for distributed deep learning on GPU clusters,” in USENIX ATC , 2017
2017
Earlier work this paper cites.
Y. You, A. Buluç, and J. Demmel, “Scaling deep learning on GPU and knights landing clusters,” in International Conference for High Performance Computing, Networking, Storage and Analysis , 2017
2017
Earlier work this paper cites.
K. Hsieh, A. Harlap, N. Vijaykumar et al. , “Gaia: Geo-distributed machine learning approaching LAN speeds,” in USENIX NSDI , 2017
2017
Earlier work this paper cites.
A. F. Aji and K. Heafield, “Sparse communication for distributed gradient descent,” in Proceedings of Empirical Methods in Natural Language Processing , 2017, pp. 440–445
2017
Earlier work this paper cites.
C. Hardy, E. Le Merrer, and B. Sericola, “Distributed deep learning on edge-devices: feasibility via adaptive compression,” in IEEE NCA , 2017
2017
Cited alongside, same era.
M. McHugh, “GPUs are the new star of moore’s law, nvidia channel boss claims,” 2018. [Online]. Available: https://www.channelweb.co.uk/crn-uk/news/3032004/gpus-are-the-new-star-of-moores-law-nvidia-channel-boss-claims
2018
Cited alongside, same era.
A. Wong, “The mobile GPU comparison guide rev. 18.2,” 2018. [Online]. Available: https://www.techarp.com/computer/mobile-gpu-comparison-guide/
2018
Cited alongside, same era.
P. Jiang and G. Agrawal, “A linear speedup analysis of distributed deep learning with sparse and quantized communication,” in NeurIPS , 2018
2018
Cited alongside, same era.
T. D. Nguyen, S. Marchal, M. Miettinen et al. , “Guardiot: A federated self-learning anomaly detection system for IoT,” in IEEE ICDCS , 2019
2019
Later among the works it cites.
C. Chen, W. Wang, and B. Li, “Round-robin synchronization: Mitigating communication bottlenecks in parameter servers,” in IEEE INFOCOM , 2019
2019
Later among the works it cites.
S. Shi, X. Chu, and B. Li, “MG-WFBP: Efficient data communication for distributed synchronous SGD algorithms,” in IEEE INFOCOM , 2019
2019
Later among the works it cites.
S. Wang, T. Tuor, T. Salonidis et al. , “Adaptive federated learning in resource constrained edge computing systems,” IEEE Journal on Selected Areas in Communications , vol. 37, no. 6, pp. 1205–1221, 2019
2019
Later among the works it cites.
J. Wang and G. Joshi, “Adaptive communication strategies to achieve the best error-runtime trade-off in local-update SGD,” in SysML , 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2018
Cited alongside, same era.
J. Wangni, J. Wang, J. Liu, and T. Zhang, “Gradient sparsification for communication-efficient distributed optimization,” in NeurIPS , 2018
2018
Cited alongside, same era.
Y. Lin, S. Han, H. Mao, Y. Wang, and W. J. Dally, “Deep gradient compression: Reducing the communication bandwidth for distributed training,” in ICLR , 2018
2018
Cited alongside, same era.
D. Alistarh, T. Hoefler, M. Johansson et al. , “The convergence of sparsified gradient methods,” in NeurIPS , 2018, pp. 5977–5987
2018
Cited alongside, same era.
C.-Y. Chen, J. Choi, D. Brand et al. , “Adacomp: Adaptive residual gradient compression for data-parallel distributed training,” in AAAI , 2018
2018
Cited alongside, same era.
2018
Cited alongside, same era.
2018
Cited alongside, same era.
2019
Cited alongside, same era.
2019
Later among the works it cites.
N. H. Tran, W. Bao, A. Zomaya, N. Minh N.H., and C. S. Hong, “Federated learning over wireless networks: Optimization model design and analysis,” in IEEE INFOCOM , 2019
2019
Later among the works it cites.
L. Wang, W. Wang, and B. Li, “CMFL: Mitigating communication overhead for federated learning,” in IEEE ICDCS , 2019
2019
Later among the works it cites.
2019
Later among the works it cites.
2019
Later among the works it cites.
S. Shi, Q. Wang, K. Zhao et al. , “A distributed synchronous SGD algorithm with global top-k sparsification for low bandwidth networks,” in IEEE ICDCS , 2019
2019
Later among the works it cites.
S. Shi, K. Zhao, Q. Wang, Z. Tang, and X. Chu, “A convergence analysis of distributed SGD with communication-efficient gradient sparsification,” in IJCAI , 2019
2019
Later among the works it cites.
F. Sattler, S. Wiedemann, K. Müller, and W. Samek, “Robust and communication-efficient federated learning from non-i.i.d. data,” IEEE Transactions on Neural Networks and Learning Systems , Nov. 2019
2019
Later among the works it cites.
2019
Later among the works it cites.
2019
Later among the works it cites.
J. Wang and G. Joshi, “Cooperative SGD: A unified framework for the design and analysis of communication-efficient SGD algorithms,” in ICML , 2019
2019
Later among the works it cites.