Fetching the paper…
Reading the bibliography…
The DANE algorithm is an approximate Newton method popularly used for communication-efficient distributed machine learning.
Efficient stochastic gradient hard thresholding
Pan Zhou, Xiaotong Yuan, and Jiashi Feng · 1908
Earlier work this paper cites.
Memory and communication efficient distributed stochastic optimization with minibatch prox
Jialei Wang, Weiran Wang, and Nathan Srebro · 1919
Earlier work this paper cites.
Some methods of speeding up the convergence of iteration methods
B. Polyak · 1964
Earlier work this paper cites.
On the momentum term in gradient descent learning algorithms
N. Qian · 1999
Earlier work this paper cites.
Rcv1: A new benchmark collection for text categorization research
David D Lewis, Yiming Yang, Tony G Rose, and Fan Li · 2004
Earlier work this paper cites.
Result analysis of the nips 2003 feature selection challenge
Isabelle Guyon, Steve Gunn, Asa Ben-Hur, and Gideon Dror · 2005
Earlier work this paper cites.
Mapreduce: Simplified data processing on large clusters
Jeffrey Dean and Sanjay Ghemawat · 2008
Earlier work this paper cites.
Stochastic convex optimization
Shai Shalev-Shwartz, Ohad Shamir, Nathan Srebro, and Karthik Sridharan · 2009
Earlier work this paper cites.
Distributed optimization and statistical learning via the alternating direction method of multipliers
Stephen Boyd, Neal Parikh, Eric Chu, Borja Peleato, and Jonathan Eckstein · 2011
Earlier work this paper cites.
User-friendly tail bounds for sums of random matrices
Joel A Tropp · 2012
Earlier work this paper cites.
Accelerating stochastic gradient descent using predictive variance reduction
Rie Johnson and Tong Zhang · 2013
Earlier work this paper cites.
Communication-efficient distributed dual coordinate ascent
Martin Jaggi, Virginia Smith, Martin Takác, Jonathan Terhorst, Sanjay Krishnan, Thomas Hofmann, and Michael I Jordan · 2014
Earlier work this paper cites.
Communication efficient distributed machine learning with the parameter server
Mu Li, David G Andersen, Alex J Smola, and Kai Yu · 2014
Earlier work this paper cites.
Communication-efficient distributed optimization using an approximate newton-type method
Ohad Shamir, Nati Srebro, and Tong Zhang · 2014
Cited alongside, same era.
Communication complexity of distributed convex learning and optimization
Yossi Arjevani and Ohad Shamir · 2015
Cited alongside, same era.
Global convergence of the heavy-ball method for convex optimization
Euhanna Ghadimi, Hamid Reza Feyzmahdavian, and Mikael Johansson · 2015
Cited alongside, same era.
A universal catalyst for first-order optimization
Hongzhou Lin, Julien Mairal, and Zaid Harchaoui · 2015
Cited alongside, same era.
Adding vs. averaging in distributed primal-dual optimization
Chenxin Ma, Virginia Smith, Martin Jaggi, Michael Jordan, Peter Richtarik, and Martin Takac · 2015
Cited alongside, same era.
Petuum: A new platform for distributed machine learning on big data
Eric P Xing, Qirong Ho, Wei Dai, Jin Kyu Kim, Jinliang Wei, Seunghak Lee, Xun Zheng, Pengtao Xie, Abhimanu Kumar, and Yaoliang Yu · 2015
Apache spark: A unified engine for big data processing
Matei Zaharia, Reynold S. Xin, Patrick Wendell, Tathagata Das, Michael Armbrust, Ankur Dave, Xiangrui Meng, Josh Rosen, Shivaram Venkataraman, Michael J. Franklin, Ali Ghodsi, Joseph Gonzalez, Scott Shenker, and Ion Stoica · 2016
Later among the works it cites.
Sgdlibrary: A matlab library for stochastic optimization algorithms
Hiroyuki Kasai · 2017
Later among the works it cites.
Distributed stochastic variance reduced gradient methods by sampling extra data with replacement
Jason D Lee, Qihang Lin, Tengyu Ma, and Tianbao Yang · 2017
Later among the works it cites.
Linearly convergent stochastic heavy ball method for minimizing generalization error
Nicolas Loizou and Peter Richtárik · 2017
Later among the works it cites.
Communication-efficient learning of deep networks from decentralized data
Brendan McMahan, Eider Moore, Daniel Ramage, Seth Hampson, and Blaise Aguera y Arcas · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
DiSCO: Distributed optimization for self-concordant empirical loss
Yuchen Zhang and Lin Xiao · 2015
Cited alongside, same era.
Federated optimization: Distributed machine learning for on-device intelligence
Jakub Konečnỳ, H Brendan McMahan, Daniel Ramage, and Peter Richtárik · 2016
Cited alongside, same era.
AIDE: Fast and communication efficient distributed optimization
Sashank J Reddi, Jakub Konečnỳ, Peter Richtárik, Barnabás Póczós, and Alex Smola · 2016
Cited alongside, same era.
Distributed coordinate descent method for learning with big data
Peter Richtárik and Martin Takáč · 2016
Cited alongside, same era.
Without-replacement sampling for stochastic gradient methods
Ohad Shamir · 2016
Cited alongside, same era.
A lyapunov analysis of momentum methods in optimization
Ashia C Wilson, Benjamin Recht, and Michael I Jordan · 2016
Cited alongside, same era.
Stochastic primal-dual coordinate method for regularized empirical risk minimization
Yuchen Zhang and Lin Xiao · 2017
Later among the works it cites.
Communication-efficient distributed statistical inference
Michael I Jordan, Jason D Lee, and Yun Yang · 2018
Later among the works it cites.
The landscape of empirical risk for nonconvex losses
Song Mei, Yu Bai, Andrea Montanari, et al · 2018
Later among the works it cites.
Cocoa: A general framework for communication-efficient distributed optimization
Virginia Smith, Simone Forte, Ma Chenxin, Martin Takáč, Michael I Jordan, and Martin Jaggi · 2018
Later among the works it cites.
Giant: Globally improved approximate newton method for distributed optimization
Shusen Wang, Farbod Roosta-Khorasani, Peng Xu, and Michael W Mahoney · 2018
Later among the works it cites.
Distributed inexact newton-type pursuit for non-convex sparse learning
Bo Liu, Xiao-Tong Yuan, Lezi Wang, Qingshan Liu, Junzhou Huang, and Dimitris Metaxas · 2019
Closest in time.
Dscovr: Randomized primal-dual block coordinate algorithms for asynchronous distributed optimization
Lin Xiao, Adams Wei Yu, Qihang Lin, and Weizhu Chen · 2019
Closest in time.