Fetching the paper…
Reading the bibliography…
Stochastic gradient methods enable learning probabilistic models from large amounts of data.
Generalized linear models
P. McCullagh · 1984
Earlier work this paper cites.
Acceleration of stochastic approximation by averaging
B. T. Polyak and A. B. Juditsky · 1992
Earlier work this paper cites.
Markov chains and stochastic stability
S. P. Meyn and R. L. Tweedie · 1993
Earlier work this paper cites.
Markov chain Monte Carlo in practice
W. R. Gilks, S. Richardson, and D. Spiegelhalter · 1995
Earlier work this paper cites.
Conditional random fields: Probabilistic models for segmenting and labeling sequence data
J. Lafferty, A. McCallum, and F. Pereira · 2001
Earlier work this paper cites.
Learning with Kernels: Support Vector Machines, Regularization, Optimization, and beyond
B. Scholkopf and A. J. Smola · 2001
Earlier work this paper cites.
Using the nyström method to speed up kernel machines
C. K.I. Williams and M. Seeger · 2001
Earlier work this paper cites.
Kernel Methods for Pattern Analysis
J. Shawe-Taylor and N. Cristianini · 2004
Earlier work this paper cites.
Fast kernel classifiers with online and active learning
A. Bordes, S. Ertekin, J. Weston, and L. Bottou · 2005
Earlier work this paper cites.
Gaussian Processes for Machine Learning
C. E. Rasmussen and C. K. I. Williams · 2006
Cited alongside, same era.
Optimal rates for the regularized least-squares algorithm
A. Caponnetto and E. De Vito · 2007
Cited alongside, same era.
Injective hilbert space embeddings of probability measures
B. K. Sriperumbudur, A. Gretton, K. Fukumizu, G. Lanckriet, and B. Schölkopf · 2008
Cited alongside, same era.
Probabilistic Graphical Models: Principles and Techniques - Adaptive Computation and Machine Learning
D. Koller and N. Friedman · 2009
Cited alongside, same era.
Non-asymptotic analysis of stochastic approximation algorithms for machine learning
F. Bach and E. Moulines · 2011
Cited alongside, same era.
Negative binomial regression
J. M. Hilbe · 2011
Cited alongside, same era.
Non-strongly-convex smooth stochastic approximation with convergence rate O ( 1 / n ) {O}(1/n)
F. Bach and E. Moulines · 2013
Later among the works it cites.
UCI machine learning repository, 2013
M. Lichman · 2013
Later among the works it cites.
Sketching as a tool for numerical linear algebra
D. P. Woodruff et al · 2014
Later among the works it cites.
Optimization methods for large-scale machine learning
L. Bottou, F. E. Curtis, and J. Nocedal · 2016
Later among the works it cites.
Nonparametric stochastic approximation with large step-sizes
A. Dieuleveut and F. Bach · 2016
Later among the works it cites.
Deep Learning
I. Goodfellow, Y. Bengio, and A. Courville · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Machine Learning: A Probabilistic Perspective
K. P. Murphy · 2012
Cited alongside, same era.
Sharp analysis of low-rank kernel matrix approximations
F. Bach · 2013
Cited alongside, same era.
Bridging the gap between constant step size stochastic gradient descent and markov chains
A. Dieuleveut, A. Durmus, and F. Bach · 2017
Later among the works it cites.
Falkon: An optimal large scale kernel method
A. Rudi, L. Carratino, and L. Rosasco · 2017
Later among the works it cites.