Fetching the paper…
Reading the bibliography…
Stochastic Gradient Descent (SGD) has become the method of choice for solving a broad range of machine learning problems.
Theory of reproducing kernels
Nachman Aronszajn · 1950
Earlier work this paper cites.
Scales of banach spaces
S. G. Krein and Yu I. Petunin · 1966
Earlier work this paper cites.
Error bounds for tikhonov regularization in hilbert scales
Frank Natterer · 1984
Earlier work this paper cites.
Efficient estimations from a slowly convergent Robbins–Monro process
David Ruppert · 1988
Earlier work this paper cites.
Acceleration of stochastic approximation by averaging
B. T. Polyak and A. B. Juditsky · 1992
Earlier work this paper cites.
Tikhonov regularization for finitely and infinitely smoothing operators
Bernard A Mair · 1994
Earlier work this paper cites.
Introduction to partial differential equations
Gerald B Folland · 1995
Earlier work this paper cites.
Regularization of inverse problems
Heinz Werner Engl, Martin Hanke, and Andreas Neubauer · 1996
Earlier work this paper cites.
Error estimates for regularization methods in hilbert scales
Ulrich Tautenhahn · 1996
Earlier work this paper cites.
Effective dimension and generalization of kernel learning
T. Zhang · 2003
Earlier work this paper cites.
Preconditioning landweber iteration in hilbert scales
Herbert Egger and Andreas Neubauer · 2005
Earlier work this paper cites.
Regularization in hilbert scales under general smoothing conditions
M Thamban Nair, Sergei V Pereverzev, and Ulrich Tautenhahn · 2005
Earlier work this paper cites.
Convergence rates for tikhonov regularization from different kinds of smoothness conditions
Albrecht Böttcher, Bernd Hofmann, Ulrich Tautenhahn, and Masahiro Yamamoto · 2006
Earlier work this paper cites.
Optimal rates for regularized least-squares algorithm
Andrea Caponnetto and E. De Vito · 2006
Earlier work this paper cites.
Approximate source conditions in tikhonov–phillips regularization and consequences for inverse problems with multiplication operators
Bernd Hofmann · 2006
Earlier work this paper cites.
Online learning algorithms
Steve Smale and Yuan Yao · 2006
Cited alongside, same era.
On regularization algorithms in learning theory
Frank Bauer, Sergei Pereverzev, and Lorenzo Rosasco · 2007
Cited alongside, same era.
Analysis of profile functions for general linear regularization methods
Bernd Hofmann and Peter Mathé · 2007
Cited alongside, same era.
Error bounds for regularization methods in hilbert scales by using operator monotonicity
Peter Mathe and Ulrich Tautenhahn · 2007
Cited alongside, same era.
Support Vector Machines
I. Steinwart and A. Christmann · 2008
Cited alongside, same era.
Online gradient descent learning algorithms
Yiming Ying and Massimiliano Pontil · 2008
Cited alongside, same era.
Learning with incremental iterative regularization
Lorenzo Rosasco and Silvia Villa · 2015
Later among the works it cites.
Nonparametric stochastic approximation with large step-sizes
Aymeric Dieuleveut and Francis Bach · 2016
Later among the works it cites.
Train faster, generalize better: Stability of stochastic gradient descent
Moritz Hardt, Ben Recht, and Yoram Singer · 2016
Later among the works it cites.
Generalization properties and implicit regularization for multiple passes SGM
Junhong Lin, Raffaello Camoriano, and Lorenzo Rosasco · 2016
Later among the works it cites.
Harder, better, faster, stronger convergence rates for least-squares regression
Aymeric Dieuleveut, Nicolas Flammarion, and Francis Bach · 2017
Later among the works it cites.
Optimal rates for multi-pass stochastic gradient methods
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Robust stochastic approximation approach to stochastic programming
Arkadi Nemirovski, Anatoli Juditsky, Guanghui Lan, and Alexander Shapiro · 2009
Cited alongside, same era.
Better mini-batch algorithms via accelerated gradient methods
Andrew Cotter, Ohad Shamir, Nati Srebro, and Karthik Sridharan · 2011
Cited alongside, same era.
Sharp converse results for the regularization error using distance functions
Jens Flemming, Bernd Hofmann, and Peter Mathé · 2011
Cited alongside, same era.
Making gradient descent optimal for strongly convex stochastic optimization
Alexander Rakhlin, Ohad Shamir, and Karthik Sridharan · 2012
Cited alongside, same era.
Methods of modern mathematical physics: Functional analysis
Michael Reed · 2012
Cited alongside, same era.
Non-strongly-convex smooth stochastic approximation with convergence rate o(1/n)
Francis R. Bach and Eric Moulines · 2013
Cited alongside, same era.
Junhong Lin and Lorenzo Rosasco · 2017
Later among the works it cites.
Optimal rates for the regularized learning algorithms under general source condition
Abhishake Rastogi and Sivananthan Sampath · 2017
Later among the works it cites.
Optimal rates for regularization of statistical inverse learning problems
Gilles Blanchard and Nicole Mücke · 2018
Later among the works it cites.
Optimal rates for spectral algorithms with least-squares regression over hilbert spaces
Junhong Lin, Alessandro Rudi, Lorenzo Rosasco, and Volkan Cevher · 2018
Later among the works it cites.
Statistical optimality of stochastic gradient descent on hard learning problems through multiple passes
Loucas Pillaud-Vivien, Alessandro Rudi, and Francis Bach · 2018
Later among the works it cites.
Lepskii principle in supervised learning
Gilles Blanchard, Peter Mathé, and Nicole Mücke · 2019
Later among the works it cites.
Reproducing kernel hilbert spaces on manifolds: Sobolev and diffusion spaces
Ernesto De Vito, Nicole Mücke, and Lorenzo Rosasco · 2019
Later among the works it cites.
Sobolev norm learning rates for regularized least-squares algorithm
Simon Fischer and Ingo Steinwart · 2019
Later among the works it cites.
Beating sgd saturation with tail-averaging and minibatching
Nicole Mücke, Gergely Neu, and Lorenzo Rosasco · 2019
Later among the works it cites.
Inverse learning in hilbert scales
Abhishake Rastogi and Peter Mathé · 2020
Closest in time.