Fetching the paper…
Reading the bibliography…
We show that parametric models trained by a stochastic gradient method (SGM) with few iterations have vanishing generalization error.
Monotone operators and the proximal point algorithm
R. T. Rockafellar · 1976
Earlier work this paper cites.
On Cezari’s convergence of the steepest descent method for approximating saddle point of convex-concave functions
A. Nemirovski and D. B. Yudin · 1978
Earlier work this paper cites.
Distribution-free performance bounds for potential function rules
L. P. Devroye and T. Wagner · 1979
Earlier work this paper cites.
Problem complexity and method efficiency in optimization
A. Nemirovski and D. B. Yudin · 1983
Earlier work this paper cites.
Learning internal representations by error propagation
D. E. Rumelhart, G. E. Hinton, and R. J. Williams · 1986
Earlier work this paper cites.
Introduction to optimization
B. T. Polyak · 1987
Earlier work this paper cites.
A simple weight decay can improve generalization
A. Krogh and J. A. Hertz · 1992
Earlier work this paper cites.
Weak convergence and local stability properties of fixed step size recursive algorithms
J. Bucklew, T. G. Kurtz, and W. Sethares · 1993
Earlier work this paper cites.
Building a large annotated corpus of English: The Penn Treebank
M. P. Marcus, M. A. Marcinkiewicz, and B. Santorini · 1993
Earlier work this paper cites.
Online algorithms and stochastic approximations
L. Bottou · 1998
Earlier work this paper cites.
Large margin classification using the perceptron algorithm
Y. Freund and R. E. Schapire · 1999
Earlier work this paper cites.
Algorithmic stability and sanity-check bounds for leave-one-out cross-validation
M. J. Kearns and D. Ron · 1999
Earlier work this paper cites.
Stability and generalization
O. Bousquet and A. Elisseeff · 2002
Earlier work this paper cites.
Stochastic Approximation and Recursive Algorithms and Applications
H. J. Kushner and G. G. Yin · 2003
Earlier work this paper cites.
Introductory lectures on convex optimization
Y. Nesterov · 2004
Cited alongside, same era.
Signal recovery by proximal forward-backward splitting
P. L. Combettes and V. R. Wajs · 2005
Cited alongside, same era.
Stability of randomized learning algorithms
A. Elisseeff, T. Evgeniou, and M. Pontil · 2005
Cited alongside, same era.
Logarithmic regret algorithms for online convex optimization
E. Hazan, A. Kalai, S. Kale, and A. Agarwal · 2006
Cited alongside, same era.
Learning theory: stability is sufficient for generalization and necessary and sufficient for consistency of empirical risk minimization
S. Mukherjee, P. Niyogi, T. Poggio, and R. M. Rifkin · 2006
Cited alongside, same era.
The tradeoffs of large scale learning
L. Bottou and O. Bousquet · 2008
Cited alongside, same era.
First-order methods of smooth convex optimization with inexact oracle
O. Devolder, F. Glineur, and Y. Nesterov · 2014
Later among the works it cites.
Beyond the regret minimization barrier: optimal algorithms for stochastic strongly-convex optimization
E. Hazan and S. Kale · 2014
Later among the works it cites.
Analysis and design of optimization algorithms via integral quadratic constraints
L. Lessard, B. Recht, and A. Packard · 2014
Later among the works it cites.
On the computational efficiency of training neural networks
R. Livni, S. Shalev-Shwartz, and O. Shamir · 2014
Later among the works it cites.
Learning with incremental iterative regularization
L. Rosasco and S. Villa · 2014
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Robust stochastic approximation approach to stochastic programming
A. Nemirovski, A. Juditsky, G. Lan, and A. Shapiro · 2009
Cited alongside, same era.
Learnability, stability and uniform convergence
S. Shalev-Shwartz, O. Shamir, N. Srebro, and K. Sridharan · 2010
Cited alongside, same era.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Cited alongside, same era.
An optimal method for stochastic composite optimization
G. Lan · 2012
Cited alongside, same era.
Making gradient descent optimal for strongly convex stochastic optimization
A. Rakhlin, O. Shamir, and K. Sridharan · 2012
Cited alongside, same era.
Stochastic first-and zeroth-order methods for nonconvex stochastic programming
S. Ghadimi and G. Lan · 2013
Cited alongside, same era.
Dropout: A simple way to prevent neural networks from overfitting
N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov · 2014
Later among the works it cites.
Recurrent neural network regularization
W. Zaremba, I. Sutskever, and O. Vinyals · 2014
Later among the works it cites.
Simple, efficient, and neural algorithms for sparse coding
S. Arora, R. Ge, T. Ma, and A. Moitra · 2015
Closest in time.
Competing with the empirical risk minimizer in a single pass
R. Frostig, R. Ge, S. M. Kakade, and A. Sidford · 2015
Closest in time.
Non-stochastic best arm identification and hyperparameter optimization
K. Jamieson and A. Talwalkar · 2015
Closest in time.
Generalization bounds for neural networks through tensor factorization
M. Janzamin, H. Sedghi, and A. Anandkumar · 2015
Closest in time.
In search of the real inductive bias: On the role of implicit regularization in deep learning
B. Neyshabur, R. Tomioka, and N. Srebro · 2015
Closest in time.
On the generalization properties of differential privacy
K. Nissim and U. Stemmer · 2015
Closest in time.
Going deeper with convolutions
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich · 2015
Closest in time.