Fetching the paper…
Reading the bibliography…
Early stopping is a widely used technique to prevent poor generalization performance when training an over-expressive model by means of gradient-based optimization.
A stochastic approximation method
H. Robbins and S. Monro · 1951
Earlier work this paper cites.
The subgroup algorithm for generating uniform random variables
P. Diaconis and M. Shahshahani · 1987
Earlier work this paper cites.
Generalization and parameter estimation in feedforward nets: Some experiments
N. Morgan and H. Bourlard · 1989
Earlier work this paper cites.
A simple weight decay can improve generalization
A. Krogh and J. A. Hertz · 1991
Earlier work this paper cites.
Creating artificial neural networks that generalize
J. Sietsma and R. J. Dow · 1991
Earlier work this paper cites.
Pruning algorithms-a survey
R. Reed · 1993
Earlier work this paper cites.
Regression shrinkage and selection via the lasso
R. Tibshirani · 1996
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
Pattern Recognition and Machine Learning
C. M. Bishop · 2006
Cited alongside, same era.
Extracting and composing robust features with denoising autoencoders
P. Vincent, H. Larochelle, Y. Bengio, and P.-A. Manzagol · 2008
Cited alongside, same era.
LIBSVM: A library for support vector machines , 2011
C.-C. Chang and C.-J. Lin · 2011
Cited alongside, same era.
UCI Machine Learning Repository: Breast Cancer Wisconsin (Diagnostic) Data Set , Jan. 2011
W. H. Wolberg, W. N. Street, and O. L. Mangasarian · 2011
Cited alongside, same era.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Cited alongside, same era.
Early Stopping — But When? , pages 53–67
L. Prechelt · 2012
Early stopping is nonparametric variational inference
D. Maclaurin, D. Duvenaud, and R. P. Adams · 2015
Later among the works it cites.
Probabilistic line searches for stochastic optimization
M. Mahsereci and P. Hennig · 2015
Later among the works it cites.
Going deeper with convolutions
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich · 2015
Later among the works it cites.
RMSprop Gradient Optimization , 2015
T. Tieleman and G. Hinton · 2015
Later among the works it cites.
Coupling Adaptive Batch Sizes with Learning Rates
L. Balles, J. Romero, and P. Hennig · 2016
Later among the works it cites.
Deep Learning
I. Goodfellow, Y. Bengio, and A. Courville · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
New perspectives on the natural gradient method
J. Martens · 2014
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition"
K. Simonyan and A. Zisserman · 2014
Cited alongside, same era.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Later among the works it cites.
Follow the Signs for Robust Stochastic Optimization
L. Balles and P. Hennig · 2017
Closest in time.