Fetching the paper…
Reading the bibliography…
Iterative procedures for parameter estimation based on stochastic gradient descent allow the estimation to scale to massive data sets.
A stochastic approximation method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
Efficient recursive estimation; application to estimating the parameters of a covariance function
David J Sakrison · 1965
Earlier work this paper cites.
A learning method for system identification
Jin-Ichi Nagumo and Atsuhiko Noda · 1967
Earlier work this paper cites.
Monotone operators and the proximal point algorithm
R Tyrrell Rockafellar · 1976
Earlier work this paper cites.
Splitting algorithms for the sum of two nonlinear operators
Pierre-Louis Lions and Bertrand Mercier · 1979
Earlier work this paper cites.
Efficient estimators from a slowly convergent robbins-monro process
David Ruppert · 1988
Earlier work this paper cites.
Stochastic approximation: A generalisation of the Robbins-Monro procedure , volume 89
JA Bather · 1989
Earlier work this paper cites.
Adaptive algorithms and stochastic approximations
Albert Benveniste, Pierre Priouret, and Michel Métivier · 1990
Earlier work this paper cites.
Stochastic approximation and optimization of random systems , volume 17
Lennart Ljung, Georg Ch Pflug, and Harro Walk · 1992
Earlier work this paper cites.
Acceleration of stochastic approximation by averaging
Boris T Polyak and Anatoli B Juditsky · 1992
Earlier work this paper cites.
Comparison of neural networks and discriminant analysis in predicting forest cover types
Jock Blackard · 1998
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann Le Cun, Léon Bottou, Yoshua Bengio, and Patrick Haffner · 1998
Earlier work this paper cites.
Stochastic learning
Léon Bottou · 2004
Earlier work this paper cites.
Rcv1: A new benchmark collection for text categorization research
David Lewis, Yiming Yang, Tony Rose, and Fan Li · 2004
Earlier work this paper cites.
Introductory lectures on convex optimization , volume 87
Yurii Nesterov · 2004
Earlier work this paper cites.
The p-norm generalization of the lms algorithm for adaptive filtering
Jyrki Kivinen, Manfred K Warmuth, and Babak Hassibi · 2006
Cited alongside, same era.
A geometric view of non-linear on-line stochastic gradient descent
Krzysztof A Krakowski, Robert E Mahony, Robert C Williamson, and Manfred K Warmuth · 2007
Cited alongside, same era.
Implicit online learning with kernels
Li Cheng SVN Schuurmans and SW Caelli · 2007
Cited alongside, same era.
Pascal large scale learning challenge, 2008
Soeren Sonnenburg, Vojtech Franc, Elad Yom-Tov, and Michele Sebag · 2008
Cited alongside, same era.
A fast iterative shrinkage-thresholding algorithm for linear inverse problems
Amir Beck and Marc Teboulle · 2009
Cited alongside, same era.
Robust stochastic approximation approach to stochastic programming
Arkadi Nemirovski, Anatoli Juditsky, Guanghui Lan, and Alexander Shapiro · 2009
Ohad Shamir and Tong Zhang · 2012
Later among the works it cites.
Lecture 6.5—RmsProp: Divide the Gradient by a Running Average of its Recent Magnitude
T. Tieleman and G. Hinton · 2012
Later among the works it cites.
Non-strongly-convex smooth stochastic approximation with convergence rate O(1/n)
Francis Bach and Eric Moulines · 2013
Later among the works it cites.
Proximal algorithms
Neal Parikh and Stephen Boyd · 2013
Later among the works it cites.
Minimizing finite sums with the stochastic average gradient
Mark Schmidt, Nicolas Le Roux, and Francis Bach · 2013
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Efficient learning using forward-backward splitting
Yoram Singer and John C Duchi · 2009
Cited alongside, same era.
Online importance weight aware updates
Nikos Karampatziakis and John Langford · 2010
Cited alongside, same era.
Implicit online learning
Brian Kulis and Peter L Bartlett · 2010
Cited alongside, same era.
Incremental proximal methods for large scale convex optimization
Dimitri P Bertsekas · 2011
Cited alongside, same era.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Cited alongside, same era.
Non-asymptotic analysis of stochastic approximation algorithms for machine learning
Eric Moulines and Francis R Bach · 2011
Cited alongside, same era.
Lorenzo Rosasco, Silvia Villa, and Bang Công Vũ · 2014
Later among the works it cites.
Stochastic proximal iteration: A non-asymptotic improvement upon stochastic gradient descent
Ernest K Ryu and Stephen Boyd · 2014
Later among the works it cites.
Implicit stochastic gradient descent
Panos Toulis and Edoardo M Airoldi · 2014
Later among the works it cites.
Statistical analysis of stochastic gradient methods for generalized linear models
Panos Toulis, Jason Rennie, and Edoardo Airoldi · 2014
Later among the works it cites.
A proximal stochastic gradient method with progressive variance reduction
Lin Xiao and Tony Zhang · 2014
Later among the works it cites.
Solving large scale linear prediction problems using stochastic gradient descent algorithms
Tong Zhang · 2014
Later among the works it cites.
Averaged least-mean-squares: Bias-variance trade-offs and optimal sampling distributions
Alexandre Défossez and Francis Bach · 2015
Closest in time.
Adam: A Method for Stochastic Optimization
Diederik Kingma and Jimmy Ba · 2015
Closest in time.
Implicit stochastic approximation
Panos Toulis and Edoardo M Airoldi · 2015
Closest in time.