Fetching the paper…
Reading the bibliography…
Machine learning (ML) training algorithms often possess an inherent self-correcting behavior due to their iterative-convergent nature.
Innovations in internetworking
R. Sandberg, D. Golgberg, S. Kleiman, D. Walsh, and B. Lyon · 1988
Earlier work this paper cites.
The collapsed gibbs sampler in bayesian computations with applications to a gene regulation problem
Jun S. Liu · 1994
Earlier work this paper cites.
Newsweeder: Learning to filter netnews
Ken Lang · 1995
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Y. Lecun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
Eigentaste: A constant time collaborative filtering algorithm
Ken Goldberg, Theresa Roeder, Dhruv Gupta, and Chris Perkins · 2001
Earlier work this paper cites.
Accuracy and Stability of Numerical Algorithms
Nicholas J. Higham · 2002
Earlier work this paper cites.
A higher order estimate of the optimum checkpoint interval for restart dumps
J. T. Daly · 2004
Earlier work this paper cites.
Rcv1: A new benchmark collection for text categorization research
David D. Lewis, Yiming Yang, Tony G. Rose, and Fan Li · 2004
Earlier work this paper cites.
Ceph: A scalable, high-performance distributed file system
Sage A. Weil, Scott A. Brandt, Ethan L. Miller, Darrell D. E. Long, and Carlos Maltzahn · 2006
Earlier work this paper cites.
Robust stochastic approximation approach to stochastic programming
Arkadi Nemirovski, Anatoli Juditsky, Guanghui Lan, and Alexander Shapiro · 2009
Earlier work this paper cites.
Proximal alternating minimization and projection methods for nonconvex problems: An approach based on the kurdyka-łojasiewicz inequality
Hédy Attouch, Jérôme Bolte, Patrick Redont, and Antoine Soubeyran · 2010
Earlier work this paper cites.
Zookeeper: Wait-free coordination for internet-scale systems
Patrick Hunt, Mahadev Konar, Flavio P. Junqueira, and Benjamin Reed · 2010
Earlier work this paper cites.
Cassandra: A decentralized structured storage system
Avinash Lakshman and Prashant Malik · 2010
Earlier work this paper cites.
Rectified linear units improve restricted boltzmann machines
Vinod Nair and Geoffrey E. Hinton · 2010
Earlier work this paper cites.
Mesos: A platform for fine-grained resource sharing in the data center
Benjamin Hindman, Andy Konwinski, Matei Zaharia, Ali Ghodsi, Anthony D. Joseph, Randy Katz, Scott Shenker, and Ion Stoica · 2011
Earlier work this paper cites.
Hogwild!: A lock-free approach to parallelizing stochastic gradient descent
Feng Niu, Benjamin Recht, Christopher Re, and Stephen J. Wright · 2011
Earlier work this paper cites.
Large scale distributed deep networks
Jeffrey Dean, Greg S. Corrado, Rajat Monga, Kai Chen, Matthieu Devin, Quoc V. Le, Mark Z. Mao, Marc’Aurelio Ranzato, Andrew Senior, Paul Tucker, Ke Yang, and Andrew Y. Ng · 2012
Earlier work this paper cites.
Distributed graphlab: A framework for machine learning and data mining in the cloud
Yucheng Low, Danny Bickson, Joseph Gonzalez, Carlos Guestrin, Aapo Kyrola, and Joseph M. Hellerstein · 2012
Cited alongside, same era.
Making gradient descent optimal for strongly convex stochastic optimization
Alexander Rakhlin, Ohad Shamir, Karthik Sridharan, et al · 2012
Cited alongside, same era.
Solving the straggler problem with bounded staleness
James Cipar, Qirong Ho, Jin Kyu Kim, Seunghak Lee, Gregory R. Ganger, Garth Gibson, Kimberly Keeton, and Eric Xing · 2013
Cited alongside, same era.
Facc1: Freebase annotation of clueweb corpora, version 1 (release date 2013-06-26, format version 1, correction level 0)
Evgeniy Gabrilovich, Michael Ringgaard, and Amarnag Subramanya · 2013
Cited alongside, same era.
More effective distributed ml via a stale synchronous parallel parameter server
Qirong Ho, James Cipar, Henggang Cui, Jin Kyu Kim, Seunghak Lee, Phillip B. Gibbons, Garth A. Gibson, Gregory R. Ganger, and Eric P. Xing · 2013
Cited alongside, same era.
Optimization methods for large-scale machine learning
Léon Bottou, Frank E Curtis, and Jorge Nocedal · 2016
Later among the works it cites.
Addressing the straggler problem for iterative convergent parallel ml
Aaron Harlap, Henggang Cui, Wei Dai, Jinliang Wei, Gregory R. Ganger, Phillip B. Gibbons, Garth A. Gibson, and Eric P. Xing · 2016
Later among the works it cites.
Machine learning with adversaries: Byzantine tolerant gradient descent
Peva Blanchard, Rachid Guerraoui, Julien Stainer, et al · 2017
Later among the works it cites.
Distributed statistical machine learning in adversarial settings: Byzantine gradient descent
Yudong Chen, Lili Su, and Jiaming Xu · 2017
Later among the works it cites.
UCI machine learning repository, 2017
Dua Dheeru and Efi Karra Taniskidou · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Introductory lectures on convex optimization: A basic course , volume 87
Yurii Nesterov · 2013
Cited alongside, same era.
Apache hadoop yarn: Yet another resource negotiator
Vinod Kumar Vavilapalli, Arun C. Murthy, Chris Douglas, Sharad Agarwal, Mahadev Konar, Robert Evans, Thomas Graves, Jason Lowe, Hitesh Shah, Siddharth Seth, Bikas Saha, Carlo Curino, Owen O’Malley, Sanjay Radia, Benjamin Reed, and Eric Baldeschwieler · 2013
Cited alongside, same era.
Low precision arithmetic for deep learning
Matthieu Courbariaux, Yoshua Bengio, and Jean-Pierre David · 2014
Cited alongside, same era.
Exploiting bounded staleness to speed up big data analytics
Henggang Cui, James Cipar, Qirong Ho, Jin Kyu Kim, Seunghak Lee, Abhimanu Kumar, Jinliang Wei, Wei Dai, Gregory R. Ganger, Phillip B. Gibbons, Garth A. Gibson, and Eric P. Xing · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2014
Cited alongside, same era.
High-performance distributed ml at scale through parameter server consistency models
Wei Dai, Abhimanu Kumar, Jinliang Wei, Qirong Ho, Garth Gibson, and Eric P. Xing · 2015
Cited alongside, same era.
Escaping from saddle points — online stochastic gradient for tensor decomposition
Rong Ge, Furong Huang, Chi Jin, and Yang Yuan · 2015
Cited alongside, same era.
Simon S Du, Chi Jin, Jason D Lee, Michael I Jordan, Barnabas Poczos, and Aarti Singh · 2017
Later among the works it cites.
On randomized distributed coordinate descent with quantized updates
Mostafa El Gamal and Lifeng Lai · 2017
Later among the works it cites.
Proteus: Agile ml elasticity through tiered reliability in dynamic resource markets
Aaron Harlap, Alexey Tumanov, Andrew Chung, Gregory R. Ganger, and Phillip B. Gibbons · 2017
Later among the works it cites.
Quantized neural networks: Training neural networks with low precision weights and activations
Itay Hubara, Matthieu Courbariaux, Daniel Soudry, Ran El-Yaniv, and Yoshua Bengio · 2017
Later among the works it cites.
How to escape saddle points efficiently
Chi Jin, Rong Ge, Praneeth Netrapalli, Sham M Kakade, and Michael I Jordan · 2017
Later among the works it cites.
A globally convergent algorithm for nonconvex optimization based on block coordinate update
Yangyang Xu and Wotao Yin · 2017
Later among the works it cites.
Poseidon: An efficient communication architecture for distributed deep learning on gpu clusters
Hao Zhang, Zeyu Zheng, Shizhen Xu, Wei Dai, Qirong Ho, Xiaodan Liang, Zhiting Hu, Jinliang Wei, Pengtao Xie, and Eric P. Xing · 2017
Later among the works it cites.
Asynchronous Byzantine machine learning (the case of SGD)
Georgios Damaskinos, El Mahdi El Mhamdi, Rachid Guerraoui, Rhicheek Patra, and Mahsa Taziki · 2018
Closest in time.
The hidden vulnerability of distributed learning in byzantium
Rachid Guerraoui, Sébastien Rouault, et al · 2018
Closest in time.
Highly Scalable Deep Learning Training System with Mixed-Precision: Training ImageNet in Four Minutes
X. Jia, S. Song, W. He, Y. Wang, H. Rong, F. Zhou, L. Xie, Z. Guo, Y. Yang, L. Yu, T. Chen, G. Hu, S. Shi, and X. Chu · 2018
Closest in time.
Litz: Elastic framework for high-performance distributed machine learning
Aurick Qiao, Abutalib Aghayev, Weiren Yu, Haoyang Chen, Qirong Ho, Garth A. Gibson, and Eric P. Xing · 2018
Closest in time.
Horovod: fast and easy distributed deep learning in tensorflow
Alexander Sergeev and Mike Del Balso · 2018
Closest in time.