Fetching the paper…
Reading the bibliography…
We develop methods for parameter estimation in settings with large-scale data sets, where traditional methods are no longer tenable.
Statistical Methods for Research Workers
Fisher RA (1925) · 1925
Earlier work this paper cites.
“A Stochastic Approximation Method.”
Robbins H, Monro S (1951) · 1951
Earlier work this paper cites.
“Dropout: A Simple Way to Prevent Neural Networks from Overfitting.”
Srivastava N, Hinton G, Krizhevsky A, Sutskever I, Salakhutdinov R (2014) · 1958
Earlier work this paper cites.
“Robust Estimation of a Location Parameter.”
Huber P (1964) · 1964
Earlier work this paper cites.
“Some Methods of Speeding Up the Convergence of Iteration Methods.”
Polyak BT (1964) · 1964
Earlier work this paper cites.
“Efficient Recursive Estimation; Application to Estimating the Parameters of a Covariance Function.”
Sakrison DJ (1965) · 1965
Earlier work this paper cites.
“Generalized Linear Models.”
Nelder J, Wedderburn R (1972) · 1972
Earlier work this paper cites.
“Monotone Operators and the Proximal Point Algorithm.”
Rockafellar RT (1976) · 1976
Earlier work this paper cites.
“Maximum Likelihood from Incomplete Data via the EM Algorithm.”
Dempster A, Laird N, Rubin D (1977) · 1977
Earlier work this paper cites.
“A Method of Solving a Convex Programming Problem with Convergence Rate O(1/k2).”
Nesterov Y (1983) · 1983
Earlier work this paper cites.
“Iteratively Reweighted Least Squares for Maximum Likelihood Estimation, and Some Robust and Resistant Alternatives.”
Green PJ (1984) · 1984
Earlier work this paper cites.
“Fundamentals of Statistical Exponential Families with Applications in Statistical Decision Theory.”
Brown LD (1986) · 1986
Earlier work this paper cites.
“Efficient Estimations from a Slowly Convergent Robbins-Monro Process.”
Ruppert D (1988) · 1988
Earlier work this paper cites.
Stochastic Approximation: A Generalisation of the Robbins-Monro Procedure , volume 89
Bather J (1989) · 1989
Earlier work this paper cites.
“Acceleration of Stochastic Approximation by Averaging.”
Polyak BT, Juditsky AB (1992) · 1992
Earlier work this paper cites.
“Learning Long-term Dependencies with Gradient Descent is Difficult.”
Bengio Y, Simard P, Frasconi P (1994) · 1994
Earlier work this paper cites.
“Natural Gradient Works Efficiently in Learning.”
Amari SI (1998) · 1998
Earlier work this paper cites.
Comparison of Neural Networks and Discriminant Analysis in Predicting Forest Cover Types
Blackard J (1998) · 1998
Earlier work this paper cites.
“The Vanishing Gradient Problem During Learning Recurrent Neural Nets and Problem Solutions.”
Hochreiter S (1998) · 1998
Earlier work this paper cites.
“Gradient-based Learning Applied to Document Recognition.”
Le Cun Y, Bottou L, Bengio Y, Haffner P (1998) · 1998
Earlier work this paper cites.
Theory of Point Estimation , volume 31
Lehmann EL, Casella G (1998) · 1998
Cited alongside, same era.
“Overfitting in Neural Nets: Backpropagation, Conjugate Gradient, and Early Stopping.”
Giles RCSLL (2001) · 2001
Cited alongside, same era.
“Gradient Flow in Recurrent Nets: The Difficulty of Learning Long-term Dependencies.”
Hochreiter S, Bengio Y, Frasconi P, Schmidhuber J (2001) · 2001
Cited alongside, same era.
Modern Applied Statistics with \proglang
Venables WN, Ripley BD (2002) · 2002
Cited alongside, same era.
“Task Clustering and Gating for Bayesian Multitask Learning.”
Bakker B, Heskes T (2003) · 2003
Cited alongside, same era.
“RCV1: A New Benchmark Collection for Text Categorization Research.”
Lewis D, Yang Y, Rose T, Li F (2004) · 2004
Cited alongside, same era.
“Hogwild!: A Lock-Free Approach to Parallelizing Stochastic Gradient Descent.”
Nui F, Recht B, Re C, Wright SJ (2011) · 2011
Later among the works it cites.
“Towards Optimal One Pass Large Scale Learning with Averaged Stochastic Gradient Descent.”
Xu W (2011) · 2011
Later among the works it cites.
“Probabilistic Topic Models.”
Blei DM (2012) · 2012
Later among the works it cites.
“Stochastic Gradient Descent Tricks.”
Bottou L (2012) · 2012
Later among the works it cites.
“On the Difficulty of Training Recurrent Neural Networks.”
Pascanu R, Mikolov T, Bengio Y (2012) · 2012
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Regularization and Variable Selection via the Elastic Net.”
Zou H, Hastie T (2005) · 2005
Cited alongside, same era.
“A Geometric View of Non-Linear On-Line Stochastic Gradient Descent.”
Krakowski KA, Mahony RE, Williamson RC, Warmuth MK (2007) · 2007
Cited alongside, same era.
“A Robust Hybrid of Lasso and Ridge Regression.”
Owen AB (2007) · 2007
Cited alongside, same era.
An Introduction to Generalized Linear Models
Dobson A, Barnett A (2008) · 2008
Cited alongside, same era.
“Compressed Sensing MRI.”
Lustig M, Donoho DL, Santos JM, Pauly JM (2008) · 2008
Cited alongside, same era.
“Pascal Large Scale Learning Challenge.”
Sonnenburg S, Franc V, Yom-Tov E, Sebag M (2008) · 2008
Cited alongside, same era.
Shamir O, Zhang T (2012) · 2012
Later among the works it cites.
“Lecture 6.5—RmsProp: Divide the Gradient by a Running Average of its Recent Magnitude.”
Tieleman T, Hinton G (2012) · 2012
Later among the works it cites.
“Non-strongly-convex Smooth Stochastic Approximation with Convergence Rate O(1/n).”
Bach F, Moulines E (2013) · 2013
Later among the works it cites.
“High Dimensional Robust M-Estimation: Asymptotic Variance via Approximate Message Passing.”
Donoho D, Montanari A (2013) · 2013
Later among the works it cites.
“Scalable Strategies for Computing with Massive Data.”
Kane MJ, Emerson J, Weston S (2013) · 2013
Later among the works it cites.
Frontiers in Massive Data Analysis
National Research Council (2013) · 2013
Later among the works it cites.
“Minimizing Finite Sums with the Stochastic Average Gradient.”
Schmidt M, Le Roux N, Bach F (2013) · 2013
Later among the works it cites.
“Stochastic Dual Coordinate Ascent Methods for Regularized Loss.”
Shalev-Shwartz S, Zhang T (2013) · 2013
Later among the works it cites.
“On Iterative Hard Thresholding Methods for High-dimensional M-Estimation.”
Jain P, Tewari A, Kar P (2014) · 2014
Later among the works it cites.
“Unit Tests for Stochastic Optimization.”
Schaul T, Antonoglou I, Silver D (2014) · 2014
Later among the works it cites.
“Statistical Analysis of Stochastic Gradient Methods for Generalized Linear Models.”
Toulis P, Airoldi E, Rennie J (2014) · 2014
Later among the works it cites.
“Deep Exponential Families.”
Ranganath R, Tang L, Charlin L, Blei DM (2015) · 2015
Closest in time.
“Variational Inference with Normalizing Flows.”
Rezende DJ, Mohamed S (2015) · 2015
Closest in time.
“Stability and Optimality in Stochastic Gradient Descent.”
Toulis P, Tran D, Airoldi EM (2015) · 2015
Closest in time.