Fetching the paper…
Reading the bibliography…
Stochastic gradient descent procedures have gained popularity for parameter estimation from large data sets.
On the mathematical foundations of theoretical statistics
Ronald A Fisher · 1922
Earlier work this paper cites.
Statistical Methods for Research Workers
R. A. Fisher · 1925
Earlier work this paper cites.
A stochastic approximation method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
On a stochastic approximation method
Kai Lai Chung · 1954
Earlier work this paper cites.
Asymptotic distribution of stochastic approximation procedures
Jerome Sacks · 1958
Earlier work this paper cites.
Multivariate chebyshev inequalities
Albert W Marshall and Ingram Olkin · 1960
Earlier work this paper cites.
Adaptive switching circuits
Bernard Widrow and Marcian E Hoff · 1960
Earlier work this paper cites.
Robust estimation of a location parameter
Peter J Huber et al · 1964
Earlier work this paper cites.
Efficient recursive estimation; application to estimating the parameters of a covariance function
David J Sakrison · 1965
Earlier work this paper cites.
A learning method for system identification
Jin-Ichi Nagumo and Atsuhiko Noda · 1967
Earlier work this paper cites.
On asymptotic normality in stochastic approximation
Vaclav Fabian · 1968
Earlier work this paper cites.
Regression models and life-tables
David R Cox · 1972
Earlier work this paper cites.
Generalized linear models
J.A. Nelder and R.W.M. Wedderburn · 1972
Earlier work this paper cites.
Stochastic approximation and recursive estimation , volume 47
Mikhail Borisovich Nevelʹson and Rafail Zalmanovich Khasʹminskiĭ · 1973
Earlier work this paper cites.
Robust estimation via stochastic approximation
R Douglas Martin and C Johan Masreliez · 1975
Earlier work this paper cites.
Monotone operators and the proximal point algorithm
R Tyrrell Rockafellar · 1976
Earlier work this paper cites.
Maximum likelihood from incomplete data via the EM algorithm
A. Dempster, N. Laird, and D. Rubin · 1977
Earlier work this paper cites.
On asymptotically efficient recursive estimation
Vaclav Fabian · 1978
Earlier work this paper cites.
Splitting algorithms for the sum of two nonlinear operators
Pierre-Louis Lions and Bertrand Mercier · 1979
Earlier work this paper cites.
Adaptive estimation algorithms: convergence, optimality, stability
Boris Teodorovich Polyak and Ya Z Tsypkin · 1979
Earlier work this paper cites.
Robust identification
BT Polyak and Ja Z Tsypkin · 1980
Earlier work this paper cites.
Problem complexity and method efficiency in optimization
DB Nemirovski, Yudin · 1983
Earlier work this paper cites.
Iteratively reweighted least squares for maximum likelihood estimation, and some robust and resistant alternatives
Peter J Green · 1984
Earlier work this paper cites.
Convergence and robustness of the robbins-monro algorithm truncated at randomly varying bounds
Han-Fu Chen, Lei Guo, and Ai-Jun Gao · 1987
Earlier work this paper cites.
Efficient estimations from a slowly convergent robbins-monro process
David Ruppert · 1988
Earlier work this paper cites.
Stochastic approximation: A generalisation of the Robbins-Monro procedure , volume 89
JA Bather · 1989
Earlier work this paper cites.
Adaptive algorithms and stochastic approximations
Albert Benveniste, Pierre Priouret, and Michel Métivier · 1990
Earlier work this paper cites.
Stochastic approximation and optimization of random systems , volume 17
Lennart Ljung, Georg Pflug, and Harro Walk · 1992
Earlier work this paper cites.
The general problem of the stability of motion
Aleksandr Mikhailovich Lyapunov · 1992
Cited alongside, same era.
Algorithm as 274: Least squares routines to supplement those of gentleman
Alan J Miller · 1992
Cited alongside, same era.
Acceleration of stochastic approximation by averaging
Boris T Polyak and Anatoli B Juditsky · 1992
Cited alongside, same era.
On the convergence behavior of the lms and the normalized lms algorithms
Dirk TM Slock · 1993
Cited alongside, same era.
Natural gradient works efficiently in learning
Shun-Ichi Amari · 1998
Cited alongside, same era.
Comparison of neural networks and discriminant analysis in predicting forest cover types
Jock Blackard · 1998
Cited alongside, same era.
Large-scale machine learning with stochastic gradient descent
Léon Bottou · 2010
Later among the works it cites.
Regularization paths for generalized linear models via coordinate descent
Jerome Friedman, Trevor Hastie, and Rob Tibshirani · 2010
Later among the works it cites.
Numerical analysis for statisticians
Kenneth Lange · 2010
Later among the works it cites.
Incremental proximal methods for large scale convex optimization
Dimitri P Bertsekas · 2011
Later among the works it cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Later among the works it cites.
The Elements of Statistical Learning: Data Mining, Inference, and Prediction
T. Hastie, R. Tibshirani, and J. Friedman · 2011
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Gradient-based learning applied to document recognition
Yann Le Cun, Léon Bottou, Yoshua Bengio, and Patrick Haffner · 1998
Cited alongside, same era.
Theory of point estimation , volume 31
Erich Leo Lehmann and George Casella · 1998
Cited alongside, same era.
Adaptive method of realizing natural gradient learning for multilayer perceptrons
Shun-Ichi Amari, Hyeyoung Park, and Kenji Fukumizu · 2000
Cited alongside, same era.
The national morbidity, mortality, and air pollution study
Jonathan M Samet, Scott L Zeger, Francesca Dominici, Frank Curriero, Ivan Coursac, Douglas W Dockery, Joel Schwartz, and Antonella Zanobetti · 2000
Cited alongside, same era.
Numerical methods for engineers and scientists
Joe D Hoffman and Steven Frankel · 2001
Cited alongside, same era.
Air pollution and mortality: estimating regional and national dose-response relationships
Francesca Dominici, Michael Daniels, Scott L Zeger, and Jonathan M Samet · 2002
Cited alongside, same era.
Non-asymptotic analysis of stochastic approximation algorithms for machine learning
Eric Moulines and Francis R Bach · 2011
Later among the works it cites.
Regularization paths for cox’s proportional hazards model via coordinate descent
Noah Simon, Jerome Friedman, Trevor Hastie, Rob Tibshirani, et al · 2011
Later among the works it cites.
Bayesian learning via stochastic gradient langevin dynamics
Max Welling and Yee W Teh · 2011
Later among the works it cites.
Towards optimal one pass large scale learning with averaged stochastic gradient descent
Wei Xu · 2011
Later among the works it cites.
Stochastic Gradient Descent Tricks
Leon Bottou · 2012
Later among the works it cites.
Large scale distributed deep networks
Jeffrey Dean, Greg Corrado, Rajat Monga, Kai Chen, Matthieu Devin, Mark Mao, Andrew Senior, Paul Tucker, Ke Yang, Quoc V Le, et al · 2012
Later among the works it cites.
Ohad Shamir and Tong Zhang · 2012
Later among the works it cites.
Non-strongly-convex smooth stochastic approximation with convergence rate o (1/n)
Francis Bach and Eric Moulines · 2013
Later among the works it cites.
High dimensional robust m-estimation: Asymptotic variance via approximate message passing
David Donoho and Andrea Montanari · 2013
Later among the works it cites.
Quasi-newton methods: A new direction
Philipp Hennig and Martin Kiefel · 2013
Later among the works it cites.
Frontiers in Massive Data Analysis
National Research Council · 2013
Later among the works it cites.
Proximal algorithms
Neal Parikh and Stephen Boyd · 2013
Later among the works it cites.
Minimizing finite sums with the stochastic average gradient
Mark Schmidt, Nicolas Le Roux, and Francis Bach · 2013
Later among the works it cites.
A stochastic quasi-newton method for large-scale optimization
Richard H Byrd, SL Hansen, Jorge Nocedal, and Yoram Singer · 2014
Closest in time.
On iterative hard thresholding methods for high-dimensional m-estimation
Prateek Jain, Ambuj Tewari, and Purushottam Kar · 2014
Closest in time.
Convergence of stochastic proximal gradient algorithm
Lorenzo Rosasco, Silvia Villa, and Bang Công Vũ · 2014
Closest in time.
Statistical analysis of stochastic gradient methods for generalized linear models
Panos Toulis, Edoardo M. Airoldi, and Jason Rennie · 2014
Closest in time.
Generalized additive models for large data sets
Simon N Wood, Yannig Goude, and Simon Shaw · 2014
Closest in time.
A proximal stochastic gradient method with progressive variance reduction
Lin Xiao and Tony Zhang · 2014
Closest in time.
Stochastic gradient descent methods for estimation with large data sets
Dustin Tran, Panos Toulis, and Edoardo M Airoldi · 2015
Closest in time.
Supplement to “asymptotic and finite-sample properties of estimators based on stochastic gradients”
Panos Toulis and Edoardo M Airoldi · 2016
Closest in time.
Towards stability and optimality in stochastic gradient descent
Panos Toulis, Dustin Tran, and Edoardo M. Airoldi · 2016
Closest in time.