Fetching the paper…
Reading the bibliography…
Stochastic Gradient Descent (SGD) based methods have been widely used for training large-scale machine learning models that also generalize well in practice.
A finite sample distribution-free performance bound for local discrimination rules
William H Rogers and Terry J Wagner · 1978
Earlier work this paper cites.
Distribution-free performance bounds with the resubstitution error estimate (corresp.)
Luc Devroye and T Wagner · 1979
Earlier work this paper cites.
Distribution-free inequalities for the deleted and holdout error estimates
Luc Devroye and Terry Wagner · 1979
Earlier work this paper cites.
Robust estimation of a location parameter
Peter J Huber · 1992
Earlier work this paper cites.
A simple weight decay can improve generalization
Anders Krogh and John A Hertz · 1992
Earlier work this paper cites.
Statistical learning theory
Vladimir N Vapnik · 1998
Earlier work this paper cites.
Algorithmic stability and sanity-check bounds for leave-one-out cross-validation
Michael Kearns and Dana Ron · 1999
Earlier work this paper cites.
Rademacher processes and bounding the risk of function learning
Vladimir Koltchinskii and Dmitriy Panchenko · 2000
Earlier work this paper cites.
Stability and generalization
Olivier Bousquet and André Elisseeff · 2002
Earlier work this paper cites.
Stability of randomized learning algorithms
Andre Elisseeff, Theodoros Evgeniou, and Massimiliano Pontil · 2005
Earlier work this paper cites.
Stability results in learning theory
Alexander Rakhlin, Sayan Mukherjee, and Tomaso Poggio · 2005
Earlier work this paper cites.
Calibrating noise to sensitivity in private data analysis
Cynthia Dwork, Frank McSherry, Kobbi Nissim, and Adam Smith · 2006
Earlier work this paper cites.
Learnability, stability and uniform convergence
Shai Shalev-Shwartz, Ohad Shamir, Nathan Srebro, and Karthik Sridharan · 2010
Earlier work this paper cites.
Stochastic gradient descent tricks
Léon Bottou · 2012
Cited alongside, same era.
Making gradient descent optimal for strongly convex stochastic optimization
Alexander Rakhlin, Ohad Shamir, and Karthik Sridharan · 2012
Cited alongside, same era.
Accelerating stochastic gradient descent using predictive variance reduction
Rie Johnson and Tong Zhang · 2013
Cited alongside, same era.
Stochastic gradient descent for non-smooth optimization: Convergence results and optimal averaging schemes
Ohad Shamir and Tong Zhang · 2013
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Cited alongside, same era.
Dropout: a simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Privacy amplification by iteration
Vitaly Feldman, Ilya Mironov, Kunal Talwar, and Abhradeep Thakurta · 2018
Later among the works it cites.
Data-dependent stability of stochastic gradient descent
Ilja Kuzborskij and Christoph Lampert · 2018
Later among the works it cites.
Generalization bounds of sgld for non-convex learning: Two theoretical viewpoints
Wenlong Mou, Liwei Wang, Xiyu Zhai, and Kai Zheng · 2018
Later among the works it cites.
On the convergence of adam and beyond
Sashank J Reddi, Satyen Kale, and Sanjiv Kumar · 2018
Later among the works it cites.
Lower bounds for finding stationary points i
Yair Carmon, John C Duchi, Oliver Hinder, and Aaron Sidford · 2019
Later among the works it cites.
Lower bounds for finding stationary points ii: First-order methods
Yair Carmon, John C Duchi, Oliver Hinder, and Aaron Sidford · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Escaping from saddle points—online stochastic gradient for tensor decomposition
Rong Ge, Furong Huang, Chi Jin, and Yang Yuan · 2015
Cited alongside, same era.
Train faster, generalize better: Stability of stochastic gradient descent
Moritz Hardt, Ben Recht, and Yoram Singer · 2016
Cited alongside, same era.
Spectrally-normalized margin bounds for neural networks
Peter L Bartlett, Dylan J Foster, and Matus J Telgarsky · 2017
Cited alongside, same era.
Exploring generalization in deep learning
Behnam Neyshabur, Srinadh Bhojanapalli, David McAllester, and Nati Srebro · 2017
Cited alongside, same era.
Bolt-on differential privacy for scalable stochastic gradient descent-based analytics
Xi Wu, Fengan Li, Arun Kumar, Kamalika Chaudhuri, Somesh Jha, and Jeffrey Naughton · 2017
Cited alongside, same era.
Stronger generalization bounds for deep nets via a compression approach
Sanjeev Arora, Rong Ge, Behnam Neyshabur, and Yi Zhang · 2018
Cited alongside, same era.
Entropy-sgd: Biasing gradient descent into wide valleys
Pratik Chaudhari, Anna Choromanska, Stefano Soatto, Yann LeCun, Carlo Baldassi, Christian Borgs, Jennifer Chayes, Levent Sagun, and Riccardo Zecchina · 2019
Later among the works it cites.
High probability generalization bounds for uniformly stable algorithms with nearly optimal rate
Vitaly Feldman and Jan Vondrak · 2019
Later among the works it cites.
Uniform convergence may be unable to explain generalization in deep learning
Vaishnavh Nagarajan and J Zico Kolter · 2019
Later among the works it cites.
Lower bounds for smooth nonconvex finite-sum optimization
Dongruo Zhou and Quanquan Gu · 2019
Later among the works it cites.
Stability of stochastic gradient descent on nonsmooth convex losses
Raef Bassily, Vitaly Feldman, Cristóbal Guzmán, and Kunal Talwar · 2020
Later among the works it cites.
Private stochastic convex optimization: optimal rates in linear time
Vitaly Feldman, Tomer Koren, and Kunal Talwar · 2020
Later among the works it cites.
On generalization error bounds of noisy gradient methods for non-convex learning
Jian Li, Xuanyuan Luo, and Mingda Qiao · 2020
Later among the works it cites.