Fetching the paper…
Reading the bibliography…
Despite their successes, what makes kernel methods difficult to use in many large scale problems is the fact that storing and computing the decision function is typically expensive, especially at prediction time.
Functions of positive and negative type and their connection with the theory of integral equations
J. Mercer · 1909
Earlier work this paper cites.
La théorie générale des noyaux réproduisants et ses applications
N. Aronszajn · 1944
Earlier work this paper cites.
Special functions of mathematical physics
H. Hochstadt · 1961
Earlier work this paper cites.
Theoretical foundations of the potential function method in pattern recognition learning
M. A. Aizerman, A. M. Braverman, and L. I. Rozonoér · 1964
Earlier work this paper cites.
A correspondence between Bayesian estimation on stochastic processes and smoothing by splines
G. S. Kimeldorf and G. Wahba · 1970
Earlier work this paper cites.
Harmonic Analysis on Semigroups
C. Berg, J. P. R. Christensen, and P. Ressel · 1984
Earlier work this paper cites.
Interpolation of scattered data: distance matrices and conditionally positive definite functions
C. A. Micchelli · 1986
Earlier work this paper cites.
Introductory Functional Analysis with Applications
E. Kreyszig · 1989
Earlier work this paper cites.
Spline Models for Observational Data , volume 59 of CBMS-NSF Regional Conference Series in Applied Mathematics
G. Wahba · 1990
Earlier work this paper cites.
A training algorithm for optimal margin classifiers
B. Boser, I. Guyon, and V. Vapnik · 1992
Earlier work this paper cites.
Rates of convergence for radial basis functions and neural networks
F. Girosi and G. Anzellotti · 1993
Earlier work this paper cites.
Priors for infinite networks
R. Neal · 1994
Earlier work this paper cites.
Support vector networks
C. Cortes and V. Vapnik · 1995
Earlier work this paper cites.
Regularization theory and neural networks architectures
F. Girosi, M. Jones, and T. Poggio · 1995
Earlier work this paper cites.
Simplified support vector decision rules
C. J. C. Burges · 1996
Earlier work this paper cites.
Isoperimetry and gaussian analysis
M. Ledoux · 1996
Earlier work this paper cites.
Support vector method for function approximation, regression estimation, and signal processing
V. Vapnik, S. Golowich, and A. J. Smola · 1997
Earlier work this paper cites.
An equivalence between sparse approximation and support vector machines
F. Girosi · 1998
Earlier work this paper cites.
Nonlinear component analysis as a kernel eigenvalue problem
B. Schölkopf, A. J. Smola, and K.-R. Müller · 1998
Cited alongside, same era.
Learning with Kernels
A. J. Smola · 1998
Cited alongside, same era.
Prediction with Gaussian processes: From linear regression to linear prediction and beyond
C. K. I. Williams · 1998
Cited alongside, same era.
Convolution kernels on discrete structures
David Haussler · 1999
Cited alongside, same era.
Sparse greedy matrix approximation for machine learning
A. J. Smola and B. Schölkopf · 2000
Cited alongside, same era.
Efficient SVM training using low-rank kernel representations
S. Fine and K. Scheinberg · 2001
Cited alongside, same era.
Group theoretical methods in machine learning
R. Kondor · 2008
Later among the works it cites.
Random features for large-scale kernel machines
A. Rahimi and B. Recht · 2008
Later among the works it cites.
Support Vector Machines
Ingo Steinwart and Andreas Christmann · 2008
Later among the works it cites.
The fast johnson-lindenstrauss transform and approximate nearest neighbors
N. Ailon and B. Chazelle · 2009
Later among the works it cites.
Learning multiple layers of features from tiny images
A. Krizhevsky · 2009
Later among the works it cites.
Fast high-dimensional kernel summations using the monte carlo multipole method
Dongryeol Lee and Alexander G. Gray · 2009
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. J. Smola, Z. L. Óvári, and R. C. Williamson · 2001
Cited alongside, same era.
Using the Nystrom method to speed up kernel machines
Christoper K. I. Williams and Matthias Seeger · 2001
Cited alongside, same era.
Generalization bounds for regularization networks and support vector machines via entropy numbers of compact operators
R. C. Williamson, A. J. Smola, and B. Schölkopf · 2001
Cited alongside, same era.
Learning with Kernels
Bernhard Schölkopf and A. J. Smola · 2002
Cited alongside, same era.
Marginalized kernels for biological sequences
K. Tsuda, T. Kin, and K. Asai · 2002
Cited alongside, same era.
Rapid evaluation of multiple density models
Alexander G. Gray and Andrew W. Moore · 2003
Cited alongside, same era.
Weighted sums of random kitchen sinks: Replacing minimization with randomization in learning
Ali Rahimi and Benjamin Recht · 2009
Later among the works it cites.
Distributed optimization and statistical learning via the alternating direction method of multipliers
S. Boyd, N. Parikh, E. Chu, B. Peleato, and J. Eckstein · 2010
Later among the works it cites.
UCI machine learning repository, 2010
A. Frank and A. Asuncion · 2010
Later among the works it cites.
Bundle methods for regularized risk minimization
Choon Hui Teo, S. V. N. Vishwanthan, A. J. Smola, and Quoc V. Le · 2010
Later among the works it cites.
Improved analysis of the subsampled randomized hadamard transform
J. A. Tropp · 2010
Later among the works it cites.
Submodular meets spectral: Greedy algorithms for subset selection, sparse approximation and dictionary selection
A. Das and D. Kempe · 2011
Later among the works it cites.
Fast locality-sensitive hashing
A. Dasgupta, R. Kumar, and T. Sarlós · 2011
Later among the works it cites.
Improved bound for the nystrom’s method and its application to kernel classification, 2011
R. Jin, T. Yang, M. Mahdavi, Y.F. Li, and Z.H. Zhou · 2011
Later among the works it cites.
O ( 1 ) O(1) computation of legendre polynomials and gauss–legendre nodes and weights for parallel computing
I. Bogaert, B. Michiels, and J. Fostier · 2012
Later among the works it cites.
Linear support vector machines via dual cached loops
S. Matsushima, S.V.N. Vishwanathan, and A.J. Smola · 2012
Later among the works it cites.
The random forest kernel and other kernels for big data from random partitions
A. Davies and Z. Ghahramani · 2014
Closest in time.