Fetching the paper…
Reading the bibliography…
In this paper we first identify a basic limitation in gradient descent-based optimization methods when used in conjunctions with smooth kernels.
The approximate arithmetical solution by finite differences of physical problems involving differential equations, with an application to the stresses in a masonry dam
L. F. Richardson · 1911
Earlier work this paper cites.
Theory of reproducing kernels
N. Aronszajn · 1950
Earlier work this paper cites.
Eigenvalues of integral operators with smooth positive definite kernels
T. Kühn · 1987
Earlier work this paper cites.
On the limited memory bfgs method for large scale optimization
D. C. Liu and J. Nocedal · 1989
Earlier work this paper cites.
Darpa timit acoustic-phonetic continous speech corpus cd-rom
J. S. Garofolo, L. F. Lamel, W. M. Fisher, J. G. Fiscus, and D. S. Pallett · 1993
Earlier work this paper cites.
An introduction to the conjugate gradient method without the agonizing pain
J. R. Shewchuk · 1994
Earlier work this paper cites.
Numerical methods for unconstrained optimization and nonlinear equations
J. E. Dennis Jr and R. B. Schnabel · 1996
Earlier work this paper cites.
The Laplacian on a Riemannian manifold: an introduction to analysis on manifolds
S. Rosenberg · 1997
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
A statistical study of on-line learning
N. Murata · 1998
Earlier work this paper cites.
Using the Nyström method to speed up kernel machines
C. Williams and M. Seeger · 2001
Earlier work this paper cites.
Stability and generalization
O. Bousquet and A. Elisseeff · 2002
Earlier work this paper cites.
Convex optimization
S. Boyd and L. Vandenberghe · 2004
Earlier work this paper cites.
Online learning with kernels
J. Kivinen, A. J. Smola, and R. C. Williamson · 2004
Earlier work this paper cites.
Kernel methods for pattern analysis
J. Shawe-Taylor and N. Cristianini · 2004
Earlier work this paper cites.
Optimal aggregation of classifiers in statistical learning
A. B. Tsybakov · 2004
Earlier work this paper cites.
Spectral properties of the kernel matrix and their relation to kernel methods in machine learning
M. L. Braun et al · 2005
Earlier work this paper cites.
Beyond the point cloud: from transductive to semi-supervised learning
V. Sindhwani, P. Niyogi, and M. Belkin · 2005
Earlier work this paper cites.
Pattern recognition and machine learning
C. Bishop · 2007
Earlier work this paper cites.
Random features for large-scale kernel machines
A. Rahimi and B. Recht · 2007
Earlier work this paper cites.
A stochastic quasi-newton method for online convex optimization
N. N. Schraudolph, J. Yu, and S. Günter · 2007
Earlier work this paper cites.
On early stopping in gradient descent learning
Y. Yao, L. Rosasco, and A. Caponnetto · 2007
Earlier work this paper cites.
A dual coordinate descent method for large-scale linear svm
C.-J. Hsieh, K.-W. Chang, C.-J. Lin, S. S. Keerthi, and S. Sundararajan · 2008
Cited alongside, same era.
Support vector machines
I. Steinwart and A. Christmann · 2008
Cited alongside, same era.
SGD-QN: Careful quasi-newton stochastic gradient descent
A. Bordes, L. Bottou, and P. Gallinari · 2009
Cited alongside, same era.
Learning multiple layers of features from tiny images
A. Krizhevsky and G. Hinton · 2009
Cited alongside, same era.
On learning with integral operators
L. Rosasco, M. Belkin, and E. D. Vito · 2010
Cited alongside, same era.
Arccosine kernels: Acoustic modeling with infinite neural networks
C.-C. Cheng and B. Kingsbury · 2011
Cited alongside, same era.
Adam: A method for stochastic optimization
D. Kingma and J. Ba · 2014
Later among the works it cites.
How to scale up kernel methods to be as good as deep neural nets
Z. Lu, A. May, K. Liu, A. B. Garakani, D. Guo, A. Bellet, L. Fan, M. Collins, B. Kingsbury, M. Picheny, and F. Sha · 2014
Later among the works it cites.
Early stopping and non-parametric regression: an optimal data-dependent stopping rule
G. Raskutti, M. Wainwright, and B. Yu · 2014
Later among the works it cites.
Convex optimization: Algorithms and complexity
S. Bubeck et al · 2015
Later among the works it cites.
Convergence rates of sub-sampled newton methods
M. A. Erdogdu and A. Montanari · 2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Adaptive subgradient methods for online learning and stochastic optimization
J. Duchi, E. Hazan, and Y. Singer · 2011
Cited alongside, same era.
Finding structure with randomness: Probabilistic algorithms for constructing approximate matrix decompositions
N. Halko, P.-G. Martinsson, and J. A. Tropp · 2011
Cited alongside, same era.
Reading digits in natural images with unsupervised feature learning
Y. Netzer, T. Wang, A. Coates, A. Bissacco, B. Wu, and A. Ng · 2011
Cited alongside, same era.
The kaldi speech recognition toolkit
D. Povey, A. Ghoshal, G. Boulianne, L. Burget, O. Glembek, N. Goel, M. Hannemann, P. Motlicek, Y. Qian, P. Schwarz, et al · 2011
Cited alongside, same era.
Pegasos: Primal estimated sub-gradient solver for SVM
S. Shalev-Shwartz, Y. Singer, N. Srebro, and A. Cotter · 2011
Cited alongside, same era.
Stable evaluation of gaussian radial basis function interpolants
G. Fasshauer and M. McCourt · 2012
Cited alongside, same era.
J. A. Tropp · 2015
Later among the works it cites.
Second order stochastic optimization in linear time
N. Agarwal, B. Bullins, and E. Hazan · 2016
Later among the works it cites.
Faster kernel ridge regression using sketching and preconditioning
H. Avron, K. Clarkson, and D. Woodruff · 2016
Later among the works it cites.
A stochastic quasi-newton method for large-scale optimization
R. H. Byrd, S. Hansen, J. Nocedal, and Y. Singer · 2016
Later among the works it cites.
NYTRO: When subsampling meets early stopping
R. Camoriano, T. Angles, A. Rudi, and L. Rosasco · 2016
Later among the works it cites.
Hierarchically compositional kernels for scalable nonparametric learning
J. Chen, H. Avron, and V. Sindhwani · 2016
Later among the works it cites.
Preconditioning kernel matrices
K. Cutajar, M. Osborne, J. Cunningham, and M. Filippone · 2016
Later among the works it cites.
Solving ridge regression using sketched preconditioned svrg
A. Gonen, F. Orabona, and S. Shalev-Shwartz · 2016
Later among the works it cites.
Universal adversarial perturbations
S.-M. Moosavi-Dezfooli, A. Fawzi, O. Fawzi, and P. Frossard · 2016
Later among the works it cites.
A linearly-convergent stochastic l-bfgs algorithm
P. Moritz, R. Nishihara, and M. Jordan · 2016
Later among the works it cites.
Back to the future: Radial basis function networks revisited
Q. Que and M. Belkin · 2016
Later among the works it cites.
Approximation of eigenfunctions in kernel-based spaces
G. Santin and R. Schaback · 2016
Later among the works it cites.
Large scale kernel learning using block coordinate descent
S. Tu, R. Roelofs, S. Venkataraman, and B. Recht · 2016
Later among the works it cites.
Understanding deep learning requires rethinking generalization
C. Zhang, S. Bengio, M. Hardt, B. Recht, and O. Vinyals · 2016
Later among the works it cites.
Kernel approximation methods for speech recognition
A. May, A. B. Garakani, Z. Lu, D. Guo, K. Liu, A. Bellet, L. Fan, M. Collins, D. Hsu, B. Kingsbury, et al · 2017
Closest in time.
On some extensions of bernstein’s inequality for self-adjoint operators
S. Minsker · 2017
Closest in time.