Fetching the paper…
Reading the bibliography…
Numerous empirical evidence has corroborated that the noise plays a crucial rule in effective and efficient training of neural networks.
Learning representations by back-propagating errors
Rumelhart, D. E · 1986
Earlier work this paper cites.
Stochastic gradient learning in neural networks
Bottou, L · 1991
Earlier work this paper cites.
Restricted boltzmann machines for collaborative filtering
Salakhutdinov, R · 2007
Earlier work this paper cites.
Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups
Hinton, G · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A · 2012
Earlier work this paper cites.
Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
Dauphin, Y. N · 2014
Earlier work this paper cites.
Thermalisation for stochastic small random perturbations of hyperbolic dynamical systems
Barrera, G · 2015
Earlier work this paper cites.
The loss surfaces of multilayer networks
Choromanska, A · 2015
Earlier work this paper cites.
Adding gradient noise improves learning for very deep networks
Neelakantan, A · 2015
Cited alongside, same era.
Identity matters in deep learning
Hardt, M · 2016
Cited alongside, same era.
Deep learning without poor local minima
Kawaguchi, K · 2016
Cited alongside, same era.
On large-batch training for deep learning: Generalization gap and sharp minima
Keskar, N. S · 2016
Cited alongside, same era.
Understanding deep learning requires rethinking generalization
Zhang, C · 2016
Cited alongside, same era.
How to escape saddle points efficiently
Jin, C · 2017
Later among the works it cites.
Convergence analysis of two-layer neural networks with relu activation
Li, Y · 2017
Later among the works it cites.
Tian, Y · 2017
Later among the works it cites.
Recovery guarantees for one-hidden-layer neural networks
Zhong, K · 2017
Later among the works it cites.
Stochastic mirror descent in variationally coherent optimization problems
Zhou, Z · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Brutzkus, A · 2017
Cited alongside, same era.
Gradient descent learns one-hidden-layer cnn: Don’t be afraid of spurious local minima
Du, S. S · 2017
Cited alongside, same era.
On the local minima of the empirical risk
Jin, C · 2018
Later among the works it cites.
An alternative view: When does sgd escape local minima?
Kleinberg, R · 2018
Later among the works it cites.