Fetching the paper…
Reading the bibliography…
The growth of data, the need for scalability and the complexity of models used in modern machine learning calls for distributed implementations.
Some studies in machine learning using the game of checkers
A. L. Samuel · 1959
Earlier work this paper cites.
The byzantine generals problem
L. Lamport, R. Shostak, and M. Pease · 1982
Earlier work this paper cites.
Semi-Martingales
M. Métivier · 1983
Earlier work this paper cites.
Distributed asynchronous deterministic and stochastic gradient optimization algorithms
J. Tsitsiklis, D. Bertsekas, and M. Athans · 1986
Earlier work this paper cites.
Implementing fault-tolerant services using the state machine approach: A tutorial
F. B. Schneider · 1990
Earlier work this paper cites.
Acceleration of stochastic approximation by averaging
B. T. Polyak and A. B. Juditsky · 1992
Earlier work this paper cites.
Distributed algorithms
N. A. Lynch · 1996
Earlier work this paper cites.
Online learning and stochastic approximations
L. Bottou · 1998
Earlier work this paper cites.
Solving large scale linear prediction problems using stochastic gradient descent algorithms
T. Zhang · 2004
Earlier work this paper cites.
Neural networks and learning machines
S. S. Haykin · 2009
Cited alongside, same era.
Large-scale machine learning with stochastic gradient descent
L. Bottou · 2010
Cited alongside, same era.
Large-scale matrix factorization with distributed stochastic gradient descent
R. Gemulla, E. Nijkamp, P. J. Haas, and Y. Sismanis · 2011
Cited alongside, same era.
Large scale distributed deep networks
J. Dean, G. Corrado, R. Monga, K. Chen, M. Devin, M. Mao, A. Senior, P. Tucker, K. Yang, Q. V. Le, et al · 2012
Cited alongside, same era.
How many computers to identify a cat? 16,000
J. Markoff · 2012
Cited alongside, same era.
Multidimensional approximate agreement in byzantine asynchronous systems
H. Mendes and M. Herlihy · 2013
Cited alongside, same era.
Asynchronous parallel stochastic gradient for nonconvex optimization
X. Lian, Y. Huang, Y. Li, and J. Liu · 2015
Later among the works it cites.
Training very deep networks
R. K. Srivastava, K. Greff, and J. Schmidhuber · 2015
Later among the works it cites.
Modeling order in neural word embeddings at scale
A. Trask, D. Gilmore, and M. Russell · 2015
Later among the works it cites.
Deep learning with elastic averaging sgd
S. Zhang, A. E. Choromanska, and Y. LeCun · 2015
Later among the works it cites.
Tensorflow: A system for large-scale machine learning
M. Abadi, P. Barham, J. Chen, Z. Chen, A. Davis, J. Dean, M. Devin, S. Ghemawat, G. Irving, M. Isard, et al · 2016
Later among the works it cites.
Fault-tolerant multi-agent optimization: optimal iterative distributed algorithms
L. Su and N. H. Vaidya · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Computing in the presence of concurrent solo executions
M. Herlihy, S. Rajsbaum, M. Raynal, and J. Stainer · 2014
Cited alongside, same era.
Machine learning: Trends, perspectives, and prospects
M. Jordan and T. Mitchell · 2015
Cited alongside, same era.
Non-bayesian learning in the presence of byzantine agents
L. Su and N. H. Vaidya · 2016
Later among the works it cites.
Forbes: The World’s Largest Tech Companies Are Making Massive AI Investments
D. Newman · 2017
Closest in time.