Fetching the paper…
Reading the bibliography…
The general perception is that kernel methods are not scalable, and neural nets are the methods of choice for nonlinear learning problems.
Integral representation of pd functions
A. Devinatz · 1953
Earlier work this paper cites.
Predicting time series with support vector machines
K.-R. Müller, A. J. Smola, G. Rätsch, B. Schölkopf, J. Kohlmorgen, and V. Vapnik · 1997
Earlier work this paper cites.
Sequential minimal optimization: A fast algorithm for training support vector machines
John C. Platt · 1998
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
Making large-scale SVM learning practical
T. Joachims · 1999
Earlier work this paper cites.
Sparse greedy matrix approximation for machine learning
A. J. Smola and B. Schölkopf · 2000
Earlier work this paper cites.
Using the Nystrom method to speed up kernel machines
C. K. I. Williams and M. Seeger · 2000
Earlier work this paper cites.
Efficient SVM training using low-rank kernel representations
S. Fine and K. Scheinberg · 2001
Earlier work this paper cites.
Estimating the support of a high-dimensional distribution
B. Schölkopf, J. Platt, J. Shawe-Taylor, A. J. Smola, and R. C. Williamson · 2001
Earlier work this paper cites.
Learning with Kernels
Bernhard Schölkopf and A. J. Smola · 2002
Earlier work this paper cites.
Online learning with kernels
J. Kivinen, A. J. Smola, and R. C. Williamson · 2004
Earlier work this paper cites.
Kernels, associated structures, and generalizations
M. Hein and O. Bousquet · 2004
Earlier work this paper cites.
On the nystr om method for approximating a gram matrix for improved kernel-based learning
P. Drineas and M. Mahoney · 2005
Earlier work this paper cites.
A modified finite Newton method for fast solution of large scale linear SVMs
S. S. Keerthi and D. DeCoste · 2005
Earlier work this paper cites.
Scattered Data Approximation
H. Wendland · 2005
Earlier work this paper cites.
Gaussian Processes for Machine Learning
C. E. Rasmussen and C. K. I. Williams · 2006
Earlier work this paper cites.
Kernel conjugate gradient for fast kernel machines
N. Ratliff and J. Bagnell · 2007
Cited alongside, same era.
Pegasos: Primal estimated sub-gradient solver for SVM
Shai Shalev-Shwartz, Yoram Singer, and Nathan Srebro · 2007
Cited alongside, same era.
Training invariant support vector machines with selective sampling
G. Loosli, S. Canu, and L. Bottou · 2007
Cited alongside, same era.
Random features for large-scale kernel machines
A. Rahimi and B. Recht · 2008
Cited alongside, same era.
Estimating divergence functionals and the likelihood ratio by penalized convex risk minimization
X.L. Nguyen, M. Wainwright, and M. Jordan · 2008
Cited alongside, same era.
Weighted sums of random kitchen sinks: Replacing minimization with randomization in learning
Ali Rahimi and Benjamin Recht · 2009
Cited alongside, same era.
Efficient additive kernels via explicit feature maps
Andrea Vedaldi and Andrew Zisserman · 2012
Later among the works it cites.
Efficiency of coordinate descent methods on huge-scale optimization problems
Yurii Nesterov · 2012
Later among the works it cites.
Fastfood — computing hilbert space expansions in loglinear time
Q.V. Le, T. Sarlos, and A. J. Smola · 2013
Later among the works it cites.
Stochastic dual coordinate ascent methods for regularized loss
Shai Shalev-Shwartz and Tong Zhang · 2013
Later among the works it cites.
Fast and scalable polynomial kernels via explicit feature maps
N. Pham and R. Pagh · 2013
Later among the works it cites.
Accelerating stochastic gradient descent using predictive variance reduction
Rie Johnson and Tong Zhang · 2013
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Robust stochastic approximation approach to stochastic programming
A. Nemirovski, A. Juditsky, G. Lan, and A. Shapiro · 2009
Cited alongside, same era.
Relative novelty detection
Alex J Smola, Le Song, and Choon H Teo · 2009
Cited alongside, same era.
Learning multiple layers of features from tiny images
A. Krizhevsky · 2009
Cited alongside, same era.
Kernel methods for deep learning
Youngmin Cho and Lawrence K. Saul · 2009
Cited alongside, same era.
On the impact of kernel approximation on learning accuracy
Corinna Cortes, Mehryar Mohri, and Ameet Talwalkar · 2010
Cited alongside, same era.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. Hinton · 2012
Cited alongside, same era.
Stochastic block mirror descent methods for nonsmooth and stochastic optimization
Cong D. Dang and Guanghui Lan · 2013
Later among the works it cites.
Learning optimally sparse support vector machines
Andrew Cotter, Shai Shalev-Shwartz, and Nati Srebro · 2013
Later among the works it cites.
Randomized nonlinear component analysis
David Lopez-Paz, Suvrit Sra, A. J. Smola, Zoubin Ghahramani, and Bernhard Schölkopf · 2014
Closest in time.
Random laplace feature maps for semigroup kernels on histograms
Jiyan Yang, Vikas Sindhwani, Quanfu Fan, Haim Avron, and Michael W. Mahoney · 2014
Closest in time.
Least squares revisited: Scalable approaches for multi-class prediction
A. Agarwal, S. Kakade, N. Karampatziakis, L. Song, and G. Valiant · 2014
Closest in time.
On data preconditioning for regularized loss minimization
Tianbao Yang, Rong Jin, and Shenghuo Zhu · 2014
Closest in time.
Quasi-monte carlo feature maps for shift-invariant kernels
Jiyan Yang, Vikas Sindhwani, Haim Avron, and Michael W. Mahoney · 2014
Closest in time.
Learning by stretching deep networks
Gaurav Pandey and Ambedkar Dukkipati · 2014
Closest in time.
On the equivalence between quadrature rules and random features
Francis R. Bach · 2015
Closest in time.