Fetching the paper…
Reading the bibliography…
We investigate the statistical behavior of gradient descent iterates with dropout in the linear regression model.
“Dropout: A Simple Way to Prevent Neural Networks from Overfitting”
Nitish Srivastava et al · 1958
Earlier work this paper cites.
“Random Coefficient Autoregressive Models: An Introduction” 11
Des. Nicholls and Barry. Quinn · 1982
Earlier work this paper cites.
“Laws of Large Numbers for Dependent Non-Identically Distributed Random Variables”
Donald.. Andrews · 1988
Earlier work this paper cites.
“Efficient Estimations from a Slowly Convergent Robbins-Monro Process”, Technical Report 781. Cornell University Operations Research and Industrial Engineering, 1988
David Ruppert · 1988
Earlier work this paper cites.
“Approximation by superpositions of a sigmoidal function”
George Cybenko · 1989
Earlier work this paper cites.
“New method of stochastic approximation type”
Boris Polyak · 1990
Earlier work this paper cites.
“Topics in Matrix Analysis”
Roger. Horn and Charles. Johnson · 1991
Earlier work this paper cites.
“Approximation capabilities of multilayer feedforward networks”
Kurt Hornik · 1991
Earlier work this paper cites.
“Acceleration of stochastic approximation by averaging”
Boris Polyak and Anatoli. Juditsky · 1992
Earlier work this paper cites.
“Multilayer feedforward networks with a nonpolynomial activation function can approximate any function”
Moshe Leshno, Vladimir. Lin, Allan Pinkus and Shimon Schocken · 1993
Earlier work this paper cites.
“On the Averaged Stochastic Approximation for Linear Regression”
László Györfi and Harro Walk · 1996
Earlier work this paper cites.
“Lectures and Exercises on Functional Analysis” 233
Aleksandr. Helemskii · 2006
Earlier work this paper cites.
“ImageNet Classification with Deep Convolutional Neural Networks”
Alex Krizhevsky, Ilya Sutskever and Geoffrey Hinton · 2012
Earlier work this paper cites.
“Adaptive dropout for training deep neural networks”
Jimmy Ba and Brendan Frey · 2013
Earlier work this paper cites.
“Understanding Dropout”
Pierre Baldi and Peter. Sadowski · 2013
Earlier work this paper cites.
“Stochastic search for semiparametric linear regression models”
Lutz Dümbgen, Richard. Samworth and Dominic Schuhmacher · 2013
Earlier work this paper cites.
“Matrix Analysis”
Roger. Horn and Charles. Johnson · 2013
Earlier work this paper cites.
“A PAC-Bayesian Tutorial with A Dropout Bound”, arXiv:1307.2118 [cs.LG], 2013
David McAllester · 2013
Earlier work this paper cites.
“Dropout Training as Adaptive Regularization”
Stefan Wager, Sida Wang and Percy Liang · 2013
Earlier work this paper cites.
“Regularization of Neural Networks using DropConnect”
Li Wan et al · 2013
Cited alongside, same era.
“Fast dropout training”
Sida Wang and Christopher Manning · 2013
Cited alongside, same era.
“Unified interval estimation for random coefficient autoregressive models”
Jonathan Hill and Liang Peng · 2014
Cited alongside, same era.
“Caffe: Convolutional Architecture for Fast Feature Embedding”
Yangqing Jia et al · 2014
Cited alongside, same era.
“Keras”, https://keras.io , 2015
François Chollet · 2015
Cited alongside, same era.
“Variational Dropout and the Local Reparameterization Trick”
Diederik. Kingma, Tim Salimans and Max Welling · 2015
Cited alongside, same era.
“PyTorch: An Imperative Style, High-Performance Deep Learning Library”
Adam Paszke et al · 2019
Later among the works it cites.
“On Convergence and Generalization of Dropout Training”
Poorya Mianjy and Raman Arora · 2020
Later among the works it cites.
“A Survey of Regularization Strategies for Deep Models”
Reza Moradi, Reza Berangi and Behrouz Minaei · 2020
Later among the works it cites.
“The Implicit and Explicit Regularization Effects of Dropout”
Colin Wei, Sham Kakade and Tengyu Ma · 2020
Later among the works it cites.
“Dropout: Explicit Forms and Capacity Control”
Raman Arora, Peter Bartlett, Poorya Mianjy and Nathan Srebro · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Haibing Wu and Xiaodong Gu · 2015
Cited alongside, same era.
“TensorFlow: A System for Large-Scale Machine Learning”
Martín Abadi et al · 2016
Cited alongside, same era.
“Computer Age Statistical Inference. Algorithms, Evidence, and Data Science” 5
Bradley Efron and Trevor Hastie · 2016
Cited alongside, same era.
“A Theoretically Grounded Application of Dropout in Recurrent Neural Networks”
Yarin Gal and Zoubin Ghahramani · 2016
Cited alongside, same era.
“Dropout as a Bayesian Approximation: Representing Model Uncertainty in Deep Learning”
Yarin Gal and Zoubin Ghahramani · 2016
Cited alongside, same era.
“Dropout Rademacher complexity of deep neural networks”
Wei Gao and Zhi-Hua Zhou · 2016
Cited alongside, same era.
Gabin Nguegnang, Holger Rauhut and Ulrich Terstiege · 2021
Later among the works it cites.
“Online Covariance Matrix Estimation in Stochastic Gradient Descent”
Wanrong Zhu, Xi Chen and Wei Wu · 2021
Later among the works it cites.
“Learning deep linear neural networks: Riemannian gradient flows and convergence to global minimizers”
Bubacarr Bah, Holger Rauhut, Ulrich Terstiege and Michael Westdickenberg · 2022
Later among the works it cites.
“Universal Approximation in Dropout Neural Networks”
Oxana. Manita et al · 2022
Later among the works it cites.
“Random autoregressive models: a structured overview”
Marta Regis, Paulo Serra and Edwin. van Heuvel · 2022
Later among the works it cites.
“Avoiding Overfitting: A Survey on Regularization Methods for Convolutional Neural Networks”
Claudioçalves Santos and João Papa · 2022
Later among the works it cites.
“Asymptotic Convergence Rate of Dropout on Shallow Linear Neural Networks”
Albert Senen-Cerda and Jaron Sanders · 2022
Later among the works it cites.
“The Dynamics of Sharpness-Aware Minimization: Bouncing Across Ravines and Drifting Towards Wide Minima”
Peter. Bartlett, Philip. Long and Olivier Bousquet · 2023
Closest in time.
Thijs Bos and Johannes Schmidt-Hieber · 2023
Closest in time.
“Central limit theorems for stochastic gradient descent with averaging for stable manifolds”
Steffen Dereich and Sebastian Kassing · 2023
Closest in time.
Johannes Schmidt-Hieber and Wouter. Koolen · 2023
Closest in time.
“Benign Overfitting in Ridge Regression”
Alexander Tsigler and Peter. Bartlett · 2023
Closest in time.
“Trained Transformers Learn Linear Models In-Context”
Ruiqi Zhang, Spencer Frei and Peter. Bartlett · 2024
Closest in time.