Fetching the paper…
Reading the bibliography…
We demonstrate that distributed block coordinate descent can quickly solve kernel regression and classification problems with millions of data points.
Parallel and distributed computation: numerical methods
Dimitri P. Bertsekas and John N. Tsitsiklis · 1989
Earlier work this paper cites.
Summa: Scalable universal matrix multiplication algorithm
Robert A. Van De Geijn and Jerrell Watts · 1997
Earlier work this paper cites.
Sequential minimal optimization: A fast algorithm for training support vector machines
John Platt · 1998
Earlier work this paper cites.
Making large-scale SVM learning practical
Thorsten Joachims · 1999
Earlier work this paper cites.
Learning with kernels
Bernhard Schölkopf and Alexander J. Smola · 2001
Earlier work this paper cites.
Using the nyström method to speed up kernel machines
Christopher Williams and Matthias Seeger · 2001
Earlier work this paper cites.
Everything old is new again: a fresh look at historical approaches in machine learning
Ryan Rifkin · 2002
Earlier work this paper cites.
In defense of one-vs-all classification
Ryan Rifkin and Aldebaro Klautau · 2004
Earlier work this paper cites.
Predictive low-rank decomposition for kernel methods
Francis Bach and Michael I. Jordan · 2005
Earlier work this paper cites.
On the nyström method for approximating a gram matrix for improved kernel-based learning
Petros Drineas and Michael W. Mahoney · 2005
Earlier work this paper cites.
Working set selection using second order information for training support vector machines
Rong-En Fan, Pai-Hsuen Chen, and Chih-Jen Lin · 2005
Earlier work this paper cites.
Accurate error bounds for the eigenvalues of the kernel matrix
Mikio L. Braun · 2006
Earlier work this paper cites.
Random features for large-scale kernel machines
Ali Rahimi and Benjamin Recht · 2007
Earlier work this paper cites.
Large scale learning with string kernels
Sören Sonnenburg, Gunnar Rätsch, and Konrad Rieck · 2007
Earlier work this paper cites.
A dual coordinate descent method for large-scale linear svm
Cho-Jui Hsieh, Kai-Wei Chang, Chih-Jen Lin, S. Sathiya Keerthi, and S. Sundararajan · 2008
Cited alongside, same era.
Feature hashing for large scale multitask learning
Kilian Weinberger, Anirban Dasgupta, John Langford, Alex Smola, and Josh Attenberg · 2009
Cited alongside, same era.
Large linear classification when data cannot fit in memory
Hsiang-Fu Yu, Cho-Jui Hsieh, Kai-Wei Chang, and Chih-Jen Lin · 2010
Cited alongside, same era.
Distributed delayed stochastic optimization
Alekh Agarwal and John C. Duchi · 2011
Cited alongside, same era.
Distributed optimization and statistical learning via the alternating direction method of multipliers
Stephen Boyd, Neal Parikh, Eric Chu, Borja Peleato, and Jonathan Eckstein · 2011
Cited alongside, same era.
Hogwild!: A lock-free approach to parallelizing stochastic gradient descent
Distributed coordinate descent method for learning with big data
Peter Richtárik and Martin Takác̆ · 2013
Later among the works it cites.
Stochastic dual coordinate ascent methods for regularized loss minimization
Shai Shalev-Shwartz and Tong Zhang · 2013
Later among the works it cites.
Trading computation for communication: Distributed stochastic dual coordinate ascent
Tianbao Yang · 2013
Later among the works it cites.
Kernel methods match deep neural networks on timit
Po-Sen Huang, Haim Avron, Tara N Sainath, Vikas Sindhwani, and Bhuvana Ramabhadran · 2014
Later among the works it cites.
Communication-efficient distributed dual coordinate ascent
Martin Jaggi, Virginia Smith, Martin Takác̆, Jonathan Terhorst, Sanjay Krishnan, Thomas Hofmann, and Michael I. Jordan · 2014
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Feng Niu, Benjamin Recht, Christopher Ré, and Stephen J. Wright · 2011
Cited alongside, same era.
Improved analysis of the subsampled randomized hadamard transform
Joel A. Tropp · 2011
Cited alongside, same era.
Parallelized stochastic gradient descent
Martin A. Zinkevich, Markus Weimer, Alex Smola, and Lihong Li · 2011
Cited alongside, same era.
Learning feature representations with k-means
Adam Coates and Andrew Y. Ng · 2012
Cited alongside, same era.
Nyström method vs random fourier features: A theoretical and empirical comparison
Tianbao Yang, Yu-Feng Li, Mehrdad Mahdavi, Rong Jin, and Zhi-Hua Zhou · 2012
Cited alongside, same era.
Resilient distributed datasets: A fault-tolerant abstraction for in-memory cluster computing
Matei Zaharia, Mosharaf Chowdhury, Tathagata Das, Ankur Dave, Justin Ma, Murphy McCauley, Michael J. Franklin, Scott Shenker, and Ion Stoica · 2012
Cited alongside, same era.
Least squares revisited: Scalable approaches for multi-class prediction
Alekh Agarwal, Sham M. Kakade, Nikos Karampatziakis, Le Song, and Gregory Valiant · 2013
Cited alongside, same era.
Ce Zhang, Arun Kumar, and Christopher Ré · 2014
Later among the works it cites.
Kernel methods and regularization techniques for nonparametric regression: Minimax optimality and adaptation
Lee H. Dicker, Dean P. Foster, and Daniel Hsu · 2015
Later among the works it cites.
Fast randomized kernel methods with statistical guarantees
Ahmed El Alaoui and Michael W. Mahoney · 2015
Later among the works it cites.
Rows vs. columns: Randomized kaczmarz or gauss-seidel for ridge regression
Ahmed Hefny, Deanna Needell, and Aaditya Ramdas · 2015
Later among the works it cites.
An asynchronous parallel stochastic coordinate descent algorithm
Ji Liu, Stephen J Wright, Christopher Ré, Victor Bittorf, and Srikrishna Sridhar · 2015
Later among the works it cites.
Adding vs. averaging in distributed primal-dual optimization
Chenxin Ma, Virginia Smith, Martin Jaggi, Michael I. Jordan, Peter Richtárik, and Martin Takác̆ · 2015
Later among the works it cites.
Less is more: Nyström computational regularization
Alessandro Rudi, Raffaello Camoriano, and Lorenzo Rosasco · 2015
Later among the works it cites.
An introduction to matrix concentration inequalities
Joel A. Tropp · 2015
Later among the works it cites.
Coordinate descent algorithms
Stephen J. Wright · 2015
Later among the works it cites.