Fetching the paper…
Reading the bibliography…
It is widely observed that deep learning models with learned parameters generalize well, even with much more model parameters than the number of training samples.
Simplifying neural nets by discovering flat minima
S. Hochreiter, J. Schmidhuber, et al · 1995
Earlier work this paper cites.
Flat minima
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Statistical learning theory
Vladimir N. Vapnik · 1998
Earlier work this paper cites.
The importance of complexity in model selection
I. J. Myung · 2000
Earlier work this paper cites.
Rademacher and gaussian complexities: Risk bounds and structural results
P. L Bartlett and S. Mendelson · 2002
Earlier work this paper cites.
Information and complexity in statistical modeling
J. Rissanen · 2007
Earlier work this paper cites.
Differential equations, dynamical systems, and an introduction to chaos
M. W. Hirsch, S. Smale, and R. L. Devaney · 2012
Earlier work this paper cites.
M. Lin, Q. Chen, and S.C. Yan · 2013
Earlier work this paper cites.
Understanding machine learning: From theory to algorithms
S. Shalev-Shwartz and S. Ben-David · 2014
Cited alongside, same era.
Subdominant dense clusters allow for simple learning and high computational performance in neural networks with discrete synapses
C. Baldassi, A. Ingrosso, C. Lucibello, L. Saglietti, and R. Zecchina · 2015
Cited alongside, same era.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
K. M. He, X. Y. Zhang, S. Q. Ren, and J. Sun · 2015
Cited alongside, same era.
Deep learning
Y. LeCun, Y. Bengio, and G. Hinton · 2015
Cited alongside, same era.
Path-sgd: Path-normalized optimization in deep neural networks
B. Neyshabur, R. R. Salakhutdinov, and N. Srebro · 2015
Cited alongside, same era.
Unreasonable effectiveness of learning neural networks: From accessible states and robust ensembles to basic algorithmic schemes
Singularity of the hessian in deep learning
L. Sagun, L. Bottou, and Y. LeCun · 2016
Later among the works it cites.
Entropy-sgd: Biasing gradient descent into wide valleys
P. Chaudhari, A. Choromanska, S. Soatto, and Y. LeCun · 2017
Closest in time.
Sharp minima can generalize for deep nets
L. Dinh, R. Pascanu, S. Bengio, and Y. Bengio · 2017
Closest in time.
Topology and geometry of half-rectified network optimization
J. Freeman, C. D.and Bruna · 2017
Closest in time.
On large-batch training for deep learning: Generalization gap and sharp minima
N. S. Keskar, D. Mudigere, J. Nocedal, M. Smelyanskiy, and P. T. P. Tang · 2017
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
C. Baldassi, C. Borgs, J. Chayes, A. Ingrosso, C. Lucibello, L. Saglietti, and R. Zecchina · 2016
Cited alongside, same era.
Train faster, generalize better: Stability of stochastic gradient descent
M. Hardt, B. Recht, and Y. Singer · 2016
Cited alongside, same era.
N. Y. Ye, Z. X. Zhu, and R. K Mantiuk · 2017
Closest in time.
Understanding deep learning requires rethinking generalization
C. Zhang, S. Bengio, M. Hardt, B. Recht, and O. Vinyals · 2017
Closest in time.