Fetching the paper…
Reading the bibliography…
Numerical Solution of Stochastic Differential Equations
Peter E. Kloeden and Eckhard Platen · 1992
Earlier work this paper cites.
A practical bayesian framework for backpropagation networks
D. J. C. MacKay · 1992
Earlier work this paper cites.
Stochastic Processes in Physics and Chemistry
N.G. Van Kampen · 1992
Earlier work this paper cites.
On-line learning processes in artificial neural networks
T. M. Heskes and B. Kappen · 1993
Earlier work this paper cites.
Bayes factors
R. E. Kass and A. E. Raftery · 1995
Earlier work this paper cites.
Flat minima
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Natural gradient works efficiently in learning
Shun-Ichi Amari · 1998
Earlier work this paper cites.
Online learning and stochastic approximations
L. Bottou · 1998
Earlier work this paper cites.
Bayesian learning via stochastic gradient Langevin dynamics
M. Welling and Y. W. Teh · 2011
Earlier work this paper cites.
Qualitatively characterizing neural network optimization problems
I. J. Goodfellow, O. Vinyals, and A. M. Saxe · 2014
Earlier work this paper cites.
New insights and perspectives on the natural gradient method
James Martens · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2014
Cited alongside, same era.
Cyclical Learning Rates for Training Neural Networks
L. N. Smith · 2015
Cited alongside, same era.
On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima
N. Shirish Keskar, D. Mudigere, J. Nocedal, M. Smelyanskiy, and P. T. P. Tang · 2016
Cited alongside, same era.
Understanding deep learning requires rethinking generalization
C. Zhang, S. Bengio, M. Hardt, B. Recht, and Oriol Vinyals · 2016
Cited alongside, same era.
High-dimensional dynamics of generalization error in neural networks
M. S. Advani and A. M. Saxe · 2017
Cited alongside, same era.
Stochastic gradient descent as approximate Bayesian inference
S. Mandt, M. D. Hoffman, and D. M. Blei · 2017
Closest in time.
Empirical Analysis of the Hessian of Over-Parametrized Neural Networks
L. Sagun, U. Evci, V. Ugur Guney, Y. Dauphin, and L. Bottou · 2017
Closest in time.
Opening the Black Box of Deep Neural Networks via Information
R. Shwartz-Ziv and N. Tishby · 2017
Closest in time.
Understanding generalization and stochastic gradient descent
S.L. Smith and Q.V. Le · 2017
Closest in time.
Towards understanding generalization of deep learning: Perspective of loss landscapes
Lei Wu, Zhanxing Zhu, et al · 2017
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A closer look at memorization in deep networks
D. Arpit and et al · 2017
Cited alongside, same era.
P. Chaudhari and S. Soatto · 2017
Cited alongside, same era.
Sharp Minima Can Generalize For Deep Nets
L. Dinh, R. Pascanu, S. Bengio, and Y. Bengio · 2017
Cited alongside, same era.
Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour
P. Goyal and et al · 2017
Cited alongside, same era.
Batch Size Matters: A Diffusion Approximation Framework on Nonconvex Stochastic Gradient Descent
C. Junchi Li and et al · 2017
Cited alongside, same era.
Stochastic modified equations and adaptive stochastic gradient algorithms
Q. Li, C. Tai, and Weinan E · 2017
Cited alongside, same era.
Stochastic Methods: A Handbook for the Natural and Social Sciences
C. Gardiner
Cited in the paper.
H. Xiao, K. Rasul, and R. Vollgraf · 2017
Closest in time.
Theory of Deep Learning III: explaining the non-overfitting puzzle
T. Poggio and et al · 2018
Closest in time.
On the information bottleneck theory of deep learning
Andrew Michael Saxe, Yamini Bansal, Joel Dapello, Madhu Advani, Artemy Kolchinsky, Brendan Daniel Tracey, and David Daniel Cox · 2018
Closest in time.
Energy-entropy competition and the effectiveness of stochastic gradient descent in machine learning
Yao Zhang, Andrew M. Saxe, Madhu S. Advani, and Alpha A. Lee · 2018
Closest in time.
The Regularization Effects of Anisotropic Noise in Stochastic Gradient Descent
Z. Zhu, J. Wu, B. Yu, L. Wu, and J. Ma · 2018
Closest in time.