Fetching the paper…
Reading the bibliography…
For distributed computing environment, we consider the empirical risk minimization problem and propose a distributed and communication-efficient Newton-type optimization method.
Problem Complexity and Method Efficiency in Optimization
A.S. Nemirovskii and D.B. Yudin · 1983
Earlier work this paper cites.
A method of solving a convex programming problem with convergence rate 𝒪 ( 1 / k 2 ) \mathcal{O}(1/k^{2})
Yurii Nesterov · 1983
Earlier work this paper cites.
Extensions of Lipschitz mappings into a Hilbert space
William B. Johnson and Joram Lindenstrauss · 1984
Earlier work this paper cites.
On the limited memory BFGS method for large scale optimization
Dong C. Liu and Jorge Nocedal · 1989
Earlier work this paper cites.
Numerical optimization
Stephen Wright and Jorge Nocedal · 1999
Earlier work this paper cites.
Database-friendly random projections: Johnson-Lindenstrauss with binary coins
Dimitris Achlioptas · 2003
Earlier work this paper cites.
Sampling algorithms for ℓ 2 \ell_{2} regression and applications
Petros Drineas, Michael W. Mahoney, and S. Muthukrishnan · 2006
Earlier work this paper cites.
Numerical optimization
Jorge Nocedal and Stephen Wright · 2006
Earlier work this paper cites.
Random features for large-scale kernel machines
Ali Rahimi and Benjamin Recht · 2007
Earlier work this paper cites.
MapReduce: simplified data processing on large clusters
Jeffrey Dean and Sanjay Ghemawat · 2008
Earlier work this paper cites.
Spark: Cluster computing with working sets
Matei Zaharia, Mosharaf Chowdhury, Michael J Franklin, Scott Shenker, and Ion Stoica · 2010
Earlier work this paper cites.
Parallelized stochastic gradient descent
Martin Zinkevich, Markus Weimer, Lihong Li, and Alex J Smola · 2010
Earlier work this paper cites.
Faster least squares approximation
Petros Drineas, Michael W. Mahoney, S. Muthukrishnan, and Tamás Sarlós · 2011
Earlier work this paper cites.
Randomized algorithms for matrices and data
Michael W. Mahoney · 2011
Earlier work this paper cites.
Hogwild: a lock-free approach to parallelizing stochastic gradient descent
Benjamin Recht, Christopher Re, Stephen Wright, and Feng Niu · 2011
Earlier work this paper cites.
Improved analysis of the subsampled randomized hadamard transform
Joel A Tropp · 2011
Earlier work this paper cites.
Fast approximation of matrix coherence and statistical leverage
Petros Drineas, Malik Magdon-Ismail, Michael W. Mahoney, and David P. Woodruff · 2012
Earlier work this paper cites.
Matrix computations
Gene H Golub and Charles F Van Loan · 2012
Earlier work this paper cites.
Distributed GraphLab: A framework for machine learning and data mining in the cloud
Yucheng Low, Danny Bickson, Joseph Gonzalez, Carlos Guestrin, Aapo Kyrola, and Joseph M. Hellerstein · 2012
Earlier work this paper cites.
Low rank approximation and regression in input sparsity time
Kenneth L. Clarkson and David P. Woodruff · 2013
Earlier work this paper cites.
Accelerating stochastic gradient descent using predictive variance reduction
Rie Johnson and Tong Zhang · 2013
Cited alongside, same era.
Faster ridge regression via the subsampled randomized Hadamard transform
Yichao Lu, Paramveer Dhillon, Dean P Foster, and Lyle Ungar · 2013
Cited alongside, same era.
An efficient distributed learning algorithm based on effective local functional approximations
Dhruv Mahajan, Nikunj Agrawal, S Sathiya Keerthi, S Sundararajan, and Léon Bottou · 2013
Cited alongside, same era.
A parallel SGD method with strong convergence
Dhruv Mahajan, S Sathiya Keerthi, S Sundararajan, and Léon Bottou · 2013
Cited alongside, same era.
Low-distortion subspace embeddings in input-sparsity time and applications to robust linear regression
Xiangrui Meng and Michael W Mahoney · 2013
Cited alongside, same era.
RandNLA: randomized numerical linear algebra
Petros Drineas and Michael W Mahoney · 2016
Later among the works it cites.
Optimization in high dimensions via accelerated, parallel, and proximal coordinate descent
Olivier Fercoq and Peter Richtárik · 2016
Later among the works it cites.
Federated optimization: distributed machine learning for on-device intelligence
Jakub Konecnỳ, H Brendan McMahan, Daniel Ramage, and Peter Richtárik · 2016
Later among the works it cites.
Federated learning: strategies for improving communication efficiency
Jakub Konecnỳ, H Brendan McMahan, Felix X Yu, Peter Richtárik, Ananda Theertha Suresh, and Dave Bacon · 2016
Later among the works it cites.
MLlib: machine learning in Apache Spark
Xiangrui Meng, Joseph Bradley, Burak Yavuz, Evan Sparks, Shivaram Venkataraman, Davies Liu, Jeremy Freeman, DB Tsai, Manish Amde, Sean Owen, et al · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
OSNAP: Faster numerical linear algebra algorithms via sparser subspace embeddings
John Nelson and Huy L Nguyên · 2013
Cited alongside, same era.
Introductory lectures on convex optimization: A basic course
Yurii Nesterov · 2013
Cited alongside, same era.
Trading computation for communication: distributed stochastic dual coordinate ascent
Tianbao Yang · 2013
Cited alongside, same era.
Communication-efficient algorithms for statistical optimization
Yuchen Zhang, John C. Duchi, and Martin J. Wainwright · 2013
Cited alongside, same era.
Communication-efficient distributed dual coordinate ascent
Martin Jaggi, Virginia Smith, Martin Takác, Jonathan Terhorst, Sanjay Krishnan, Thomas Hofmann, and Michael I Jordan · 2014
Cited alongside, same era.
Scaling distributed machine learning with the parameter server
Mu Li, David G Andersen, Jun Woo Park, Alexander J Smola, Amr Ahmed, Vanja Josifovski, James Long, Eugene J Shekita, and Bor-Yiing Su · 2014
Cited alongside, same era.
Understanding machine learning: from theory to algorithms
Shai Shalev-Shwartz and Shai Ben-David · 2014
Cited alongside, same era.
Parallel random coordinate descent method for composite minimization: Convergence analysis and error bounds
Ion Necoara and Dragos Clipici · 2016
Later among the works it cites.
AIDE: fast and communication efficient distributed optimization
Sashank J Reddi, Jakub Konecnỳ, Peter Richtárik, Barnabás Póczós, and Alex Smola · 2016
Later among the works it cites.
Distributed coordinate descent method for learning with big data
Peter Richtárik and Martin Takác · 2016
Later among the works it cites.
Parallel coordinate descent methods for big data optimization
Peter Richtárik and Martin Taká v · 2016
Later among the works it cites.
Sub-sampled Newton methods I: globally convergent algorithms
Farbod Roosta-Khorasani and Michael W Mahoney · 2016
Later among the works it cites.
Sub-sampled Newton methods II: Local convergence rates
Farbod Roosta-Khorasani and Michael W Mahoney · 2016
Later among the works it cites.
CoCoA: A general framework for communication-efficient distributed optimization
Virginia Smith, Simone Forte, Chenxin Ma, Martin Takac, Michael I Jordan, and Martin Jaggi · 2016
Later among the works it cites.
SPSD matrix approximation vis column selection: Theories, algorithms, and extensions
Shusen Wang, Luo Luo, and Zhihua Zhang · 2016
Later among the works it cites.
Sub-sampled Newton methods with non-uniform sampling
Peng Xu, Jiyan Yang, Farbod Roosta-Khorasani, Christopher Ré, and Michael W Mahoney · 2016
Later among the works it cites.
A general distributed dual coordinate optimization framework for regularized loss minimization
Shun Zheng, Fen Xia, Wei Xu, and Tong Zhang · 2016
Later among the works it cites.
Practical secure aggregation for privacy preserving machine learning
Keith Bonawitz, Vladimir Ivanov, Ben Kreuter, Antonio Marcedone, H Brendan McMahan, Sarvar Patel, Daniel Ramage, Aaron Segal, and Karn Seth · 2017
Closest in time.
Communication-efficient learning of deep networks from decentralized data
Brendan McMahan, Eider Moore, Daniel Ramage, Seth Hampson, and Blaise Aguera y Arcas · 2017
Closest in time.
Newton sketch: A near linear-time optimization algorithm with linear-quadratic convergence
Mert Pilanci and Martin J Wainwright · 2017
Closest in time.
Virginia Smith, Chao-Kai Chiang, Maziar Sanjabi, and Ameet Talwalkar · 2017
Closest in time.
Sketched ridge regression: Optimization perspective, statistical perspective, and model averaging
Shusen Wang, Alex Gittens, and Michael W. Mahoney · 2017
Closest in time.