Fetching the paper…
Reading the bibliography…
Distributed model training is vulnerable to byzantine system failures and adversarial compute nodes, i.e., nodes that use malicious updates to corrupt the global model stored at a parameter server (PS).
Probabilistic logics and the synthesis of reliable organisms from unreliable components
John Von Neumann · 1956
Earlier work this paper cites.
The byzantine generals problem
Leslie Lamport, Robert Shostak, and Marshall Pease · 1982
Earlier work this paper cites.
Reliable computation by formulas in the presence of noise
Nicholas Pippenger · 1988
Earlier work this paper cites.
Mjrty—a fast majority vote algorithm
Robert S Boyer and J Strother Moore · 1991
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner · 1998
Earlier work this paper cites.
Practical byzantine fault tolerance
Miguel Castro, Barbara Liskov, et al · 1999
Earlier work this paper cites.
Seeing stars: Exploiting class relationships for sentiment categorization with respect to rating scales
Bo Pang and Lillian Lee · 2005
Earlier work this paper cites.
Zyzzyva: speculative byzantine fault tolerance
Ramakrishna Kotla, Lorenzo Alvisi, Mike Dahlin, Allen Clement, and Edmund Wong · 2007
Earlier work this paper cites.
Improving mapreduce performance in heterogeneous environments
Matei Zaharia, Andy Konwinski, Anthony D Joseph, Randy H Katz, and Ion Stoica · 2008
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky and Geoffrey Hinton · 2009
Earlier work this paper cites.
Distributed dual averaging in networks
Alekh Agarwal, Martin J Wainwright, and John C Duchi · 2010
Earlier work this paper cites.
Better mini-batch algorithms via accelerated gradient methods
Andrew Cotter, Ohad Shamir, Nati Srebro, and Karthik Sridharan · 2011
Earlier work this paper cites.
Parallel distributed computing using python
Lisandro D Dalcin, Rodrigo R Paz, Pablo A Kler, and Alejandro Cosimo · 2011
Earlier work this paper cites.
Hogwild: A lock-free approach to parallelizing stochastic gradient descent
Benjamin Recht, Christopher Re, Stephen Wright, and Feng Niu · 2011
Earlier work this paper cites.
Large scale distributed deep networks
Jeffrey Dean, Greg Corrado, Rajat Monga, Kai Chen, Matthieu Devin, Mark Mao, Andrew Senior, Paul Tucker, Ke Yang, Quoc V Le, et al · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Accelerating stochastic gradient descent using predictive variance reduction
Rie Johnson and Tong Zhang · 2013
Earlier work this paper cites.
Project adam: Building an efficient and scalable deep learning training system
Trishul M Chilimbi, Yutaka Suzue, Johnson Apacible, and Karthik Kalyanaraman · 2014
Cited alongside, same era.
Communication-efficient distributed dual coordinate ascent
Martin Jaggi, Virginia Smith, Martin Takác, Jonathan Terhorst, Sanjay Krishnan, Thomas Hofmann, and Michael I Jordan · 2014
Cited alongside, same era.
Convolutional neural networks for sentence classification
Yoon Kim · 2014
Cited alongside, same era.
Scaling distributed machine learning with the parameter server
Mu Li, David G Andersen, Jun Woo Park, Alexander J Smola, Amr Ahmed, Vanja Josifovski, James Long, Eugene J Shekita, and Bor-Yiing Su · 2014
Cited alongside, same era.
An asynchronous parallel stochastic coordinate descent algorithm
Ji Liu, Steve Wright, Christopher Re, Victor Bittorf, and Srikrishna Sridhar · 2014
Cited alongside, same era.
Federated learning: Strategies for improving communication efficiency
Jakub Konečnỳ, H Brendan McMahan, Felix X Yu, Peter Richtárik, Ananda Theertha Suresh, and Dave Bacon · 2016
Later among the works it cites.
When do redundant requests reduce latency?
Nihar B Shah, Kangwook Lee, and Kannan Ramchandran · 2016
Later among the works it cites.
Qsgd: Communication-efficient sgd via gradient quantization and encoding
Dan Alistarh, Demjan Grubic, Jerry Li, Ryota Tomioka, and Milan Vojnovic · 2017
Later among the works it cites.
Machine learning with adversaries: Byzantine tolerant gradient descent
Peva Blanchard, Rachid Guerraoui, Julien Stainer, et al · 2017
Later among the works it cites.
Approximate gradient coding via sparse random graphs
Zachary Charles, Dimitris Papailiopoulos, and Jordan Ellenberg · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Karen Simonyan and Andrew Zisserman · 2014
Cited alongside, same era.
Mxnet: A flexible and efficient machine learning library for heterogeneous distributed systems
Tianqi Chen, Mu Li, Yutian Li, Min Lin, Naiyan Wang, Minjie Wang, Tianjun Xiao, Bing Xu, Chiyuan Zhang, and Zheng Zhang · 2015
Cited alongside, same era.
Federated optimization: Distributed optimization beyond the datacenter
Jakub Konečnỳ, Brendan McMahan, and Daniel Ramage · 2015
Cited alongside, same era.
Coded mapreduce
Songze Li, Mohammad Ali Maddah-Ali, and A Salman Avestimehr · 2015
Cited alongside, same era.
Perturbed iterate analysis for asynchronous stochastic optimization
Horia Mania, Xinghao Pan, Dimitris Papailiopoulos, Benjamin Recht, Kannan Ramchandran, and Michael I Jordan · 2015
Cited alongside, same era.
Deep learning with elastic averaging sgd
Sixin Zhang, Anna E Choromanska, and Yann LeCun · 2015
Cited alongside, same era.
Tensorflow: A system for large-scale machine learning
Martín Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, et al · 2016
Cited alongside, same era.
Distributed statistical machine learning in adversarial settings: Byzantine gradient descent
Yudong Chen, Lili Su, and Jiaming Xu · 2017
Later among the works it cites.
Coded convolution for parallel and distributed computing within a deadline
Sanghamitra Dutta, Viveck Cadambe, and Pulkit Grover · 2017
Later among the works it cites.
Speeding up distributed machine learning using codes
Kangwook Lee, Maximilian Lam, Ramtin Pedarsani, Dimitris Papailiopoulos, and Kannan Ramchandran · 2017
Later among the works it cites.
Automatic differentiation in pytorch
Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer · 2017
Later among the works it cites.
Pytorch, 2017
Adam Paszke, Sam Gross, Soumith Chintala, and Gregory Chanan · 2017
Later among the works it cites.
Coded computation over heterogeneous clusters
Amirhossein Reisizadeh, Saurav Prakash, Ramtin Pedarsani, and Salman Avestimehr · 2017
Later among the works it cites.
Gradient coding from cyclic mds codes and expander graphs
Netanel Raviv, Itzhak Tamo, Rashish Tandon, and Alexandros G Dimakis · 2017
Later among the works it cites.
Gradient coding: Avoiding stragglers in distributed learning
Rashish Tandon, Qi Lei, Alexandros G Dimakis, and Nikos Karampatziakis · 2017
Later among the works it cites.
Coded distributed computing for inverse problems
Yaoqing Yang, Pulkit Grover, and Soummya Kar · 2017
Later among the works it cites.
DRACO: robust distributed training via redundant gradients
Lingjiao Chen, Hongyi Wang, Zachary B. Charles, and Dimitris S. Papailiopoulos · 2018
Closest in time.
Draco: Robust distributed training against adversaries
Lingjiao Chen, Hongyi Wang, and Dimitris Papailiopoulos · 2018
Closest in time.