Fetching the paper…
Reading the bibliography…
In statistical learning theory, generalization error is used to quantify the degree to which a supervised machine learning algorithm may overfit to training data.
A method of solving a convex programming problem with convergence rate O ( 1 / k ) O(1/\sqrt{k})
Yurii Nesterov · 1983
Earlier work this paper cites.
Statistical Learning Theory
V. N. Vapnik · 1998
Earlier work this paper cites.
Stability and generalization
O. Bousquet and A. Elisseeff · 2002
Earlier work this paper cites.
Stability of randomized learning algorithms
A. Elisseeff, T. Evgeniou, and M. Pontil · 2005
Earlier work this paper cites.
Learning theory: Stability is sufficient for generalization and necessary and sufficient for consistency of empirical risk minimization
S. Mukherjee, P. Niyogi, T. Poggio, and R. Rifkin · 2006
Earlier work this paper cites.
Learnability, stability and uniform convergence
S. Shalev-Shwartz, O. Shamir, N. Srebro, and K. Sridharan · 2010
Earlier work this paper cites.
Bayesian learning via stochastic gradient Langevin dynamics
M. Welling and Y. W. Teh · 2011
Earlier work this paper cites.
Making gradient descent optimal for strongly convex stochastic optimization
A. Rakhlin, O. Shamir, and K. Sridharan · 2012
Cited alongside, same era.
Concentration Inequalities: A Nonasymptotic Theory of Independence
S. Boucheron, G. Lugosi, and P. Massart · 2013
Cited alongside, same era.
Stochastic gradient descent for non-smooth optimization: Convergence results and optimal averaging schemes
O. Shamir and T. Zhang · 2013
Cited alongside, same era.
Stochastic gradient Hamiltonian Monte Carlo
T. Chen, E. Fox, and C. Guestrin · 2014
Cited alongside, same era.
Understanding Machine Learning: From Theory to Algorithms
S. Shalev-Shwartz and S. Ben-David · 2014
Cited alongside, same era.
Escaping from saddle points—Online stochastic gradient for tensor decomposition
R. Ge, F. Huang, C. Jin, and Y. Yuan · 2015
Train faster, generalize better: Stability of stochastic gradient descent
M. Hardt, B. Recht, and Y. Singer · 2016
Later among the works it cites.
Generalization bounds for randomized learning with application to stochastic gradient descent
B. London · 2016
Later among the works it cites.
Stochastic gradient descent, weighted sampling, and the randomized Kaczmarz algorithm
D. Needell, N. Srebro, and R. Ward · 2016
Later among the works it cites.
Controlling bias in adaptive data analysis using information theory
D. Russo and J. Zou · 2016
Later among the works it cites.
How to escape saddle points efficiently
C. Jin, R. Ge, P. Netrapalli, S. M. Kakade, and M. I. Jordan · 2017
Later among the works it cites.
Generalization bounds of SGLD for non-convex learning: Two theoretical viewpoints
W. Mou, L. Wang, X. Zhai, and K. Zheng · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Stochastic optimization with importance sampling for regularized loss minimization
P. Zhao and T. Zhang · 2015
Cited alongside, same era.
Later among the works it cites.
Information-theoretic analysis of generalization capability of learning algorithms
A. Xu and M. Raginsky · 2017
Later among the works it cites.