Fetching the paper…
Reading the bibliography…
We study the generalization error of randomized learning algorithms -- focusing on stochastic gradient descent (SGD) -- using a novel combination of PAC-Bayes and algorithmic stability.
Probability inequalities for sums of bounded random variables
W. Hoeffding · 1963
Earlier work this paper cites.
Asymptotic evaluation of certain Markov process expectations for large time
M. Donsker and S. Varadhan · 1975
Earlier work this paper cites.
On the alias method for generating random variables from a discrete distribution
R. Kronmal and A. Peterson · 1979
Earlier work this paper cites.
On the method of bounded differences
C. McDiarmid · 1989
Earlier work this paper cites.
A decision-theoretic generalization of on-line learning and an application to boosting
Y. Freund and R. Schapire · 1995
Earlier work this paper cites.
PAC-Bayesian model averaging
D. McAllester · 1999
Earlier work this paper cites.
Stability and generalization
O. Bousquet and A. Elisseeff · 2002
Earlier work this paper cites.
PAC-Bayes and margins
J. Langford and J. Shawe-Taylor · 2002
Earlier work this paper cites.
PAC-Bayesian generalisation error bounds for Gaussian process classification
M. Seeger · 2002
Earlier work this paper cites.
Stability of randomized learning algorithms
A. Elisseeff, T. Evgeniou, and M. Pontil · 2005
Earlier work this paper cites.
Pac-Bayesian Supervised Classification: The Thermodynamics of Statistical Learning , volume 56 of Institute of Mathematical Statistics Lecture Notes – Monograph Series
O. Catoni · 2007
Earlier work this paper cites.
The tradeoffs of large scale learning
L. Bottou and O. Bousquet · 2008
Cited alongside, same era.
Exponentiated gradient algorithms for conditional random fields and max-margin Markov networks
M. Collins, A. Globerson, T. Koo, X. Carreras, and P. Bartlett · 2008
Cited alongside, same era.
PAC-Bayesian learning of linear classifiers
P. Germain, A. Lacasse, F. Laviolette, and M. Marchand · 2009
Cited alongside, same era.
Learning multiple layers of features from tiny images
A. Krizhevsky and G. Hinton · 2009
Cited alongside, same era.
Learnability, stability and uniform convergence
S. Shalev-Shwartz, O. Shamir, N. Srebro, and K. Sridharan · 2010
Cited alongside, same era.
Adaptive subgradient methods for online learning and stochastic optimization
J. Duchi, E. Hazan, and Y. Singer · 2011
Cited alongside, same era.
Stochastic optimization with importance sampling for regularized loss minimization
P. Zhao and T. Zhang · 2015
Later among the works it cites.
PAC-Bayesian bounds based on the Rényi divergence
L. Bégin, P. Germain, F. Laviolette, and J.-F. Roy · 2016
Later among the works it cites.
Ensemble robustness of deep learning algorithms
J. Feng, T. Zahavy, B. Kang, H. Xu, and S. Mannor · 2016
Later among the works it cites.
Train faster, generalize better: Stability of stochastic gradient descent
M. Hardt, B. Recht, and Y. Singer · 2016
Later among the works it cites.
Optimal learning for multi-pass stochastic gradient methods
J. Lin and L. Rosasco · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Concentration in unbounded metric spaces and algorithmic stability
A. Kontorovich · 2014
Cited alongside, same era.
Selfieboost: A boosting algorithm for deep learning
S. Shalev-Shwartz · 2014
Cited alongside, same era.
Train faster, generalize better: Stability of stochastic gradient descent
M. Hardt, B. Recht, and Y. Singer · 2015
Cited alongside, same era.
Learning with incremental iterative regularization
L. Rosasco and S. Villa · 2015
Cited alongside, same era.
J. Lin, R. Camoriano, and L. Rosasco · 2016
Later among the works it cites.
Stability and generalization in structured prediction
B. London, B. Huang, and L. Getoor · 2016
Later among the works it cites.
Minimizing the maximal loss: How and why
S. Shalev-Shwartz and Y. Wexler · 2016
Later among the works it cites.
Learning with differential privacy: Stability, learnability and the sufficiency and necessity of ERM principle
Y. Wang, J. Lei, and S. Fienberg · 2016
Later among the works it cites.
Data-dependent stability of stochastic gradient descent
I. Kuzborskij and C. Lampert · 2017
Closest in time.