Fetching the paper…
Reading the bibliography…
Machine Learning (ML) solutions are nowadays distributed, according to the so-called server/worker architecture.
Dropout: A simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov. 2014 · 1958
Earlier work this paper cites.
The Byzantine generals problem
Leslie Lamport, Robert Shostak, and Marshall Pease. 1982 · 1982
Earlier work this paper cites.
Impossibility of distributed consensus with one faulty process
Michael J Fischer, Nancy A Lynch, and Michael S Paterson. 1985 · 1985
Earlier work this paper cites.
Multivariate estimation with high breakdown point
Peter J Rousseeuw. 1985 · 1985
Earlier work this paper cites.
Reaching approximate agreement in the presence of faults
Danny Dolev, Nancy A Lynch, Shlomit S Pinter, Eugene W Stark, and William E Weihl. 1986 · 1986
Earlier work this paper cites.
Asymptotically optimal algorithms for approximate agreement. In Proceedings of the fifth annual ACM symposium on Principles of distributed computing . 73–87
AD Fekete. 1986 · 1986
Earlier work this paper cites.
Learning representations by back-propagating errors
David E Rumelhart, Geoffrey E Hinton, and Ronald J Williams. 1986 · 1986
Earlier work this paper cites.
Asynchronous approximate agreement. In Proceedings of the sixth annual ACM Symposium on Principles of distributed computing . 64–76
Alan David Fekete. 1987 · 1987
Earlier work this paper cites.
Implementing fault-tolerant services using the state machine approach: A tutorial
Fred B Schneider. 1990 · 1990
Earlier work this paper cites.
Theory of the backpropagation neural network
Robert Hecht-Nielsen. 1992 · 1992
Earlier work this paper cites.
Online learning and stochastic approximations
Léon Bottou. 1998 · 1998
Earlier work this paper cites.
MNIST dataset
Yann Lecunn. 1998 · 1998
Earlier work this paper cites.
Practical Byzantine fault tolerance. In OSDI , Vol. 99. 173–186
Miguel Castro, Barbara Liskov, et al · 1999
Earlier work this paper cites.
Optimal resilience asynchronous approximate agreement. In International Conference On Principles Of Distributed Systems . Springer, 229–239
Ittai Abraham, Yonatan Amit, and Danny Dolev. 2004 · 2004
Earlier work this paper cites.
Are loss functions all the same?
Lorenzo Rosasco, Ernesto De Vito, Andrea Caponnetto, Michele Piana, and Alessandro Verri. 2004 · 2004
Earlier work this paper cites.
The tradeoffs of large scale learning. In Neural Information Processing Systems . 161–168
Olivier Bousquet and Léon Bottou. 2008 · 2008
Earlier work this paper cites.
Introduction to reliable and secure distributed programming
Christian Cachin, Rachid Guerraoui, and Luis Rodrigues. 2011 · 2011
Earlier work this paper cites.
Poisoning attacks against support vector machines
Battista Biggio, Blaine Nelson, and Pavel Laskov. 2012 · 2012
Earlier work this paper cites.
How many ads does Google serve in a day?
Larry Kim. 2012 · 2012
Cited alongside, same era.
Parameter server for distributed machine learning. In Big Learning NIPS Workshop , Vol. 6. 2
Mu Li, Li Zhou, Zichao Yang, et al · 2013
Cited alongside, same era.
Project Adam: Building an Efficient and Scalable Deep Learning Training System.. In OSDI , Vol. 14. 571–582
Trishul M Chilimbi, Yutaka Suzue, Johnson Apacible, and Karthik Kalyanaraman. 2014 · 2014
Cited alongside, same era.
Scaling distributed machine learning with the parameter server. In 11th { \{ USENIX } \} Symposium on Operating Systems Design and Implementation ( { \{ OSDI } \} 14) . 583–598
Mu Li, David G Andersen, Jun Woo Park, et al · 2014
Cited alongside, same era.
Federated optimization: Distributed optimization beyond the datacenter
Jakub Konečnỳ, Brendan McMahan, and Daniel Ramage. 2015 · 2015
Cited alongside, same era.
Defending distributed systems against adversarial attacks: consensus, consensus-based learning, and statistical learning
Lili Su. 2017 · 2017
Later among the works it cites.
Poseidon: An Efficient Communication Architecture for Distributed Deep Learning on GPU Clusters. In USENIX ATC . 181–193
Hao Zhang, Zeyu Zheng, Shizhen Xu, Wei Dai, Qirong Ho, Xiaodan Liang, Zhiting Hu, Jinliang Wei, Pengtao Xie, and Eric P. Xing. 2017 · 2017
Later among the works it cites.
Byzantine stochastic gradient descent. In Advances in Neural Information Processing Systems . 4613–4623
Dan Alistarh, Zeyuan Allen-Zhu, and Jerry Li. 2018 · 2018
Later among the works it cites.
signSGD with majority vote is communication efficient and fault tolerant
Jeremy Bernstein, Jiawei Zhao, Kamyar Azizzadenesheli, and Anima Anandkumar. 2018 · 2018
Later among the works it cites.
DRACO: Byzantine-resilient Distributed Training via Redundant Gradients. In International Conference on Machine Learning . 902–911
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Deep learning
Yann LeCun, Yoshua Bengio, and Geoffrey Hinton. 2015 · 2015
Cited alongside, same era.
Multidimensional agreement in Byzantine systems
Hammurabi Mendes, Maurice Herlihy, Nitin H. Vaidya, and Vijay K. Garg. 2015 · 2015
Cited alongside, same era.
Is feature selection secure against training data poisoning?. In ICML . 1689–1698
Huang Xiao, Battista Biggio, Gavin Brown, Giorgio Fumera, Claudia Eckert, and Fabio Roli. 2015 · 2015
Cited alongside, same era.
Deep learning with elastic averaging SGD. In NIPS . 685–693
Sixin Zhang, Anna E Choromanska, and Yann LeCun. 2015 · 2015
Cited alongside, same era.
Tensorflow: A system for large-scale machine learning. In 12th { \{ USENIX } \} Symposium on Operating Systems Design and Implementation ( { \{ OSDI } \} 16) . 265–283
Martín Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, et al · 2016
Cited alongside, same era.
QSGD: Randomized Quantization for Communication-Optimal Stochastic Gradient Descent
Dan Alistarh, Jerry Li, Ryota Tomioka, and Milan Vojnovic. 2016 · 2016
Cited alongside, same era.
Mllib: Machine learning in apache spark
Xiangrui Meng, Joseph Bradley, Burak Yavuz, et al · 2016
Cited alongside, same era.
Lingjiao Chen, Hongyi Wang, Zachary Charles, and Dimitris Papailiopoulos. 2018 · 2018
Later among the works it cites.
Asynchronous Byzantine Machine Learning (the case of SGD). In ICML . 1153–1162
Georgios Damaskinos, El Mahdi El Mhamdi, Rachid Guerraoui, Rhicheek Patra, and Mahsa Taziki. 2018 · 2018
Later among the works it cites.
Robustly learning a gaussian: Getting optimal error, efficiently. In Proceedings of the Twenty-Ninth Annual ACM-SIAM Symposium on Discrete Algorithms . 2683–2702
Ilias Diakonikolas, Gautam Kamath, Daniel M Kane, Jerry Li, Ankur Moitra, and Alistair Stewart. 2018 · 2018
Later among the works it cites.
The Hidden Vulnerability of Distributed Learning in Byzantium. In International Conference on Machine Learning . 3521–3530
El Mahdi El Mhamdi, Rachid Guerraoui, and Sébastien Rouault. 2018 · 2018
Later among the works it cites.
Justin Gilmer, Luke Metz, Fartash Faghri, et al · 2018
Later among the works it cites.
Finite-time Guarantees for Byzantine-Resilient Distributed State Estimation with Noisy Measurements
Lili Su and Shahin Shahrampour. 2018 · 2018
Later among the works it cites.
Lipschitz regularity of deep neural networks: analysis and efficient estimation. In Advances in Neural Information Processing Systems . 3835–3844
Aladin Virmaux and Kevin Scaman. 2018 · 2018
Later among the works it cites.
Generalized Byzantine-tolerant SGD
Cong Xie, Oluwasanmi Koyejo, and Indranil Gupta. 2018a · 2018
Later among the works it cites.
Phocas: dimensional Byzantine-resilient stochastic gradient descent
Cong Xie, Oluwasanmi Koyejo, and Indranil Gupta. 2018b · 2018
Later among the works it cites.
Zeno: Byzantine-suspicious stochastic gradient descent
Cong Xie, Oluwasanmi Koyejo, and Indranil Gupta. 2018c · 2018
Later among the works it cites.
A Little Is Enough: Circumventing Defenses For Distributed Learning
Moran Baruch, Gilad Baruch, and Yoav Goldberg. 2019 · 2019
Closest in time.
AggregaThor: Byzantine Machine Learning via Robust Gradient Aggregation. In SysML
Georgios Damaskinos, El Mahdi El Mhamdi, Rachid Guerraoui, Arsany Guirguis, and Sébastien Rouault. 2019 · 2019
Closest in time.
DETOX: A Redundancy-based Framework for Faster and More Robust Gradient Aggregation
Shashank Rajput, Hongyi Wang, Zachary Charles, and Dimitris Papailiopoulos. 2019 · 2019
Closest in time.
Distributed Learning with Adversarial Agents Under Relaxed Network Condition
Pooja Vyavahare, Lili Su, and Nitin H Vaidya. 2019 · 2019
Closest in time.