Fetching the paper…
Reading the bibliography…
This paper presents a new class of gradient methods for distributed machine learning that adaptively skip the gradient calculations to learn with reduced communication and computation.
A stochastic approximation method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
Hedonic housing prices and the demand for clean air
David Harrison Jr. and Daniel L. Rubinfeld · 1978
Earlier work this paper cites.
Classification of radar returns from the ionosphere using neural networks
Vincent G. Sigillito, Simon P. Wing, Larrie V. Hutton, and Kile B. Baker · 1989
Earlier work this paper cites.
Scaling up the accuracy of Naive-Bayes classifiers: a decision-tree hybrid
Ron Kohavi · 1996
Earlier work this paper cites.
Learning differential diagnosis of erythemato-squamous diseases using voting feature intervals
H. Altay Güvenir, Gülşen Demiröz, and Nilsel Ilter · 1998
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner · 1998
Earlier work this paper cites.
A convergent incremental gradient method with a constant step size
Doron Blatt, Alfred O. Hero, and Hillel Gauchman · 2007
Earlier work this paper cites.
Consensus in ad hoc WSNs with noisy links – Part I: Distributed estimation of deterministic signals
Ioannis D. Schizas, Alejandro Ribeiro, and Georgios B. Giannakis · 2008
Earlier work this paper cites.
Distributed subgradient methods for multi-agent optimization
Angelia Nedic and Asuman Ozdaglar · 2009
Earlier work this paper cites.
Large-Scale Machine Learning with Stochastic Gradient Descent
Léon Bottou · 2010
Earlier work this paper cites.
Hogwild: A lock-free approach to parallelizing stochastic gradient descent
Benjamin Recht, Christopher Re, Stephen Wright, and Feng Niu · 2011
Earlier work this paper cites.
Large scale distributed deep networks
Jeffrey Dean, Greg Corrado, Rajat Monga, Kai Chen, Matthieu Devin, Mark Mao, Andrew Senior, Paul Tucker, Ke Yang, Quoc V Le, et al · 2012
Earlier work this paper cites.
UCI machine learning repository, 2013
Moshe Lichman · 2013
Earlier work this paper cites.
Introductory Lectures on Convex Optimization: A basic course , volume 87
Yurii Nesterov · 2013
Cited alongside, same era.
Communication-efficient algorithms for statistical optimization
Yuchen Zhang, John C. Duchi, and Martin J. Wainwright · 2013
Cited alongside, same era.
Communication-efficient distributed dual coordinate ascent
Martin Jaggi, Virginia Smith, Martin Takác, Jonathan Terhorst, Sanjay Krishnan, Thomas Hofmann, and Michael I Jordan · 2014
Cited alongside, same era.
Communication efficient distributed machine learning with the parameter server
Mu Li, David G Andersen, Alexander J Smola, and Kai Yu · 2014
Cited alongside, same era.
Communication-efficient distributed optimization using an approximate newton-type method
Ohad Shamir, Nati Srebro, and Tong Zhang · 2014
Cited alongside, same era.
Arock: an algorithmic framework for asynchronous parallel coordinate updates
Zhimin Peng, Yangyang Xu, Ming Yan, and Wotao Yin · 2016
Later among the works it cites.
On the convergence rate of incremental aggregated gradient algorithms
Mert Gurbuzbalaban, Asuman Ozdaglar, and Pablo A. Parrilo · 2017
Later among the works it cites.
Communication-efficient algorithms for decentralized and stochastic optimization
Guanghui Lan, Soomin Lee, and Yi Zhou · 2017
Later among the works it cites.
Asynchronous periodic event-triggered coordination of multi-agent systems
Yaohua Liu, Cameron Nowzari, Zhi Tian, and Qing Ling · 2017
Later among the works it cites.
Distributed optimization with arbitrary local solvers
Chenxin Ma, Jakub Konečnỳ, Martin Jaggi, Virginia Smith, Michael I Jordan, Peter Richtárik, and Martin Takáč · 2017
Later among the works it cites.
Federated learning: Collaborative machine learning without centralized training data
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
An asynchronous parallel stochastic coordinate descent algorithm
Ji Liu, Stephen Wright, Christopher Ré, Victor Bittorf, and Srikrishna Sridhar · 2015
Cited alongside, same era.
DiSCO: Distributed optimization for self-concordant empirical loss
Yuchen Zhang and Xiao Lin · 2015
Cited alongside, same era.
Optimization methods for large-scale machine learning
Léon Bottou, Frank E. Curtis, and Jorge Nocedal · 2016
Cited alongside, same era.
Asynchronous parallel algorithms for nonconvex big-data optimization: Model and convergence
Loris Cannelli, Francisco Facchinei, Vyacheslav Kungurtsev, and Gesualdo Scutari · 2016
Cited alongside, same era.
Convergence Rate Analysis of Several Splitting Schemes
Damek Davis and Wotao Yin · 2016
Cited alongside, same era.
Decentralized Learning for Wireless Communications and Networking
Georgios B. Giannakis, Qing Ling, Gonzalo Mateos, Ioannis D. Schizas, and Hao Zhu · 2016
Cited alongside, same era.
Linear convergence of gradient and proximal-gradient methods under the polyak-łojasiewicz condition
Hamed Karimi, Julie Nutini, and Mark Schmidt · 2016
Cited alongside, same era.
Brendan McMahan and Daniel Ramage · 2017
Later among the works it cites.
Communication-efficient learning of deep networks from decentralized data
Brendan McMahan, Eider Moore, Daniel Ramage, Seth Hampson, and Blaise Aguera y Arcas · 2017
Later among the works it cites.
Federated multi-task learning
Virginia Smith, Chao-Kai Chiang, Maziar Sanjabi, and Ameet S Talwalkar · 2017
Later among the works it cites.
A Berkeley view of systems challenges for AI
Ion Stoica, Dawn Song, Raluca Ada Popa, David Patterson, Michael W. Mahoney, Randy Katz, Anthony D. Joseph, Michael Jordan, Joseph M Hellerstein, Joseph E Gonzalez, et al · 2017
Later among the works it cites.
Asynchronous coordinate descent under more realistic assumptions
Tao Sun, Robert Hannah, and Wotao Yin · 2017
Later among the works it cites.
Distributed mean estimation with limited communication
Ananda Theertha Suresh, X. Yu Felix, Sanjiv Kumar, and H Brendan McMahan · 2017
Later among the works it cites.
Communication-efficient distributed statistical inference
Michael I. Jordan, Jason D. Lee, and Yun Yang · 2018
Closest in time.