Fetching the paper…
Reading the bibliography…
In this work, we consider the distributed stochastic optimization problem of minimizing a non-convex function $f(x) = \mathbb{E}_{\xi \sim \mathcal{D}} f(x; \xi)$ in an adversarial setting, where the individual functions $f(x; \xi)$ can also be potentially non-convex.
Parallelized stochastic gradient descent
Martin Zinkevich, Markus Weimer, Lihong Li, and Alex J Smola · 2010
Earlier work this paper cites.
Hogwild: A lock-free approach to parallelizing stochastic gradient descent
Benjamin Recht, Christopher Re, Stephen Wright, and Feng Niu · 2011
Earlier work this paper cites.
Optimal distributed online prediction using mini-batches
Ofer Dekel, Ran Gilad-Bachrach, Ohad Shamir, and Lin Xiao · 2012
Earlier work this paper cites.
More effective distributed ml via a stale synchronous parallel parameter server
Qirong Ho, James Cipar, Henggang Cui, Seunghak Lee, Jin Kyu Kim, Phillip B Gibbons, Garth A Gibson, Greg Ganger, and Eric P Xing · 2013
Earlier work this paper cites.
Stochastic variance reduction for nonconvex optimization
Sashank J. Reddi, Ahmed Hefny, Suvrit Sra, Barnabas Poczos, and Alex Smola · 2016
Earlier work this paper cites.
Machine learning with adversaries: Byzantine tolerant gradient descent
Peva Blanchard, El Mahdi El Mhamdi, Rachid Guerraoui, and Julien Stainer · 2017
Cited alongside, same era.
Less than a Single Pass: Stochastically Controlled Stochastic Gradient
Lihua Lei and Michael Jordan · 2017
Cited alongside, same era.
Non-convex finite-sum optimization via scsg methods
Lihua Lei, Cheng Ju, Jianbo Chen, and Michael I Jordan · 2017
Cited alongside, same era.
Byzantine stochastic gradient descent
Dan Alistarh, Zeyuan Allen-Zhu, and Jerry Li · 2018
Cited alongside, same era.
A linear speedup analysis of distributed deep learning with sparse and quantized communication
Peng Jiang and Gagan Agrawal · 2018
Cited alongside, same era.
Phocas: dimensional byzantine-resilient stochastic gradient descent
Securing distributed machine learning in high dimensions
Lili Su and Jiaming Xu · 2018
Later among the works it cites.
Byzantine-robust distributed learning: Towards optimal statistical rates
Dong Yin, Yudong Chen, Ramchandran Kannan, and Peter Bartlett · 2018
Later among the works it cites.
Rsa: Byzantine-robust stochastic aggregation methods for distributed learning from heterogeneous datasets
Liping Li, Wei Xu, Tianyi Chen, Georgios B Giannakis, and Qing Ling · 2019
Closest in time.
Defending against saddle point attack in Byzantine-robust distributed learning
Dong Yin, Yudong Chen, Ramchandran Kannan, and Peter Bartlett · 2019
Closest in time.
On the linear speedup analysis of communication efficient momentum sgd for distributed non-convex optimization
Hao Yu, Rong Jin, and Sen Yang · 2019
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cong Xie, Oluwasanmi Koyejo, and Indranil Gupta
Cited in the paper.
Zeno: Distributed stochastic gradient descent with suspicion-based fault-tolerance
Cong Xie, Oluwasanmi Koyejo, and Indranil Gupta
Cited in the paper.