Fetching the paper…
Reading the bibliography…
In this work, we study the problem of federated learning (FL), where distributed users aim to jointly train a machine learning model with the help of a parameter server (PS).
1901
Earlier work this paper cites.
1907
Earlier work this paper cites.
1907
Earlier work this paper cites.
1907
Earlier work this paper cites.
1908
Earlier work this paper cites.
1909
Earlier work this paper cites.
1909
Earlier work this paper cites.
1911
Earlier work this paper cites.
A. Schrijver, Theory of linear and integer programming . John Wiley & Sons, 1998
1998
Earlier work this paper cites.
T. M. Cover and J. A. Thomas, Elements of Information Theory . Wiley-Interscience, 2006
2006
Cited alongside, same era.
2012
Cited alongside, same era.
F. Seide, H. Fu, J. Droppo, G. Li, and D. Yu, “1-Bit Stochastic Gradient Descent and Application to Data-Parallel Distributed Training of Speech DNNs,” in Interspeech 2014 , September 2014
2014
Cited alongside, same era.
N. Dryden, S. A. Jacobs, T. Moon, and B. Van Essen, “Communication quantization for data-parallel training of deep neural networks,” in Proceedings of the Workshop on Machine Learning in High Performance Computing Environments , ser. MLHPC ’16. Piscataway, NJ, USA: IEEE Press, 2016, pp. 1–8. [Online]. Available: https://doi.org/10.1109/MLHPC.2016.4
2016
Cited alongside, same era.
J. Bernstein, Y.-X. Wang, K. Azizzadenesheli, and A. Anandkumar, “signSGD: Compressed optimisation for non-convex problems,” in Proceedings of the 35th International Conference on Machine Learning , ser. Proceedings of Machine Learning Research, vol. 80, 10–15 Jul 2018, pp. 560–569
2018
Later among the works it cites.
J. Wangni, J. Wang, J. Liu, and T. Zhang, “Gradient sparsification for communication-efficient distributed optimization,” in Advances in Neural Information Processing Systems 31 , 2018, pp. 1299–1309
2018
Later among the works it cites.
2018
Later among the works it cites.
2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
W. Wen, C. Xu, F. Yan, C. Wu, Y. Wang, Y. Chen, and H. Li, “TernGrad: Ternary Gradients to Reduce Communication in Distributed Deep Learning,” in Advances in Neural Information Processing Systems 30 , 2017, pp. 1509–1519
2017
Cited alongside, same era.
2017
Cited alongside, same era.
2017
Cited alongside, same era.
D. Alistarh, D. Grubic, J. Li, R. Tomioka, and M. Vojnovic, “QSGD: Communication-Efficient SGD via Gradient Quantization and Encoding,” in Advances in Neural Information Processing Systems 30 , 2017, pp. 1709–1720
2017
Cited alongside, same era.
A. T. Suresh, F. X. Yu, S. Kumar, and H. B. McMahan, “Distributed mean estimation with limited communication,” in Proceedings of the 34th International Conference on Machine Learning , 2017, p. 3329–3337
2017
Cited alongside, same era.
Later among the works it cites.
2018
Later among the works it cites.
2018
Later among the works it cites.
M. M. Amiri and D. Gündüz, “Over-the-air machine learning at the wireless edge,” in 2019 IEEE 20th International Workshop on Signal Processing Advances in Wireless Communications (SPAWC) , July 2019, pp. 1–5
2019
Later among the works it cites.
T. Sery and K. Cohen, “A sequential gradient-based multiple access for distributed learning over fading channels,” in 2019 57th Annual Allerton Conference on Communication, Control, and Computing (Allerton) , Sep. 2019, pp. 303–307
2019
Later among the works it cites.
G. Zhu, Y. Wang, and K. Huang, “Broadband analog aggregation for low-latency federated edge learning,” IEEE Transactions on Wireless Communications , vol. 19, no. 1, pp. 491–506, Jan. 2020
2020
Closest in time.