Fetching the paper…
Reading the bibliography…
Adaptive optimization algorithms, such as Adam and RMSprop, have shown better optimization performance than stochastic gradient descent (SGD) in some scenarios.
Flat minima
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Algorithms for manifold learning
Lawrence Cayton · 2005
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky and Geoffrey Hinton · 2009
Earlier work this paper cites.
Sample complexity of testing the manifold hypothesis
Hariharan Narayanan and Sanjoy Mitter · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Earlier work this paper cites.
Deep sparse rectifier neural networks
Xavier Glorot, Antoine Bordes, and Yoshua Bengio · 2011
Earlier work this paper cites.
The manifold tangent classifier
Salah Rifai, Yann N Dauphin, Pascal Vincent, Yoshua Bengio, and Xavier Muller · 2011
Earlier work this paper cites.
Lecture 6.5—RmsProp: Divide the gradient by a running average of its recent magnitude
Tijmen Tieleman and Geoffrey Hinton · 2012
Earlier work this paper cites.
Adadelta: an adaptive learning rate method
Matthew D Zeiler · 2012
Earlier work this paper cites.
Beyond convexity: Stochastic quasi-convex optimization
Elad Hazan, Kfir Levy, and Shai Shalev-Shwartz · 2015
Cited alongside, same era.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba · 2015
Cited alongside, same era.
Path-sgd: Path-normalized optimization in deep neural networks
Behnam Neyshabur, Ruslan R Salakhutdinov, and Nati Srebro · 2015
Cited alongside, same era.
Going deeper with convolutions
Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich · 2015
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
A closer look at memorization in deep networks
Devansh Arpit, Stanisław Jastrzębski, Nicolas Ballas, David Krueger, Emmanuel Bengio, Maxinder S Kanwal, Tegan Maharaj, Asja Fischer, Aaron Courville, Yoshua Bengio, et al · 2017
Closest in time.
Shake-shake regularization of 3-branch residual networks
Xavier Gastaldi · 2017
Closest in time.
Be careful what you backpropagate: A case for linear output activations & gradient boosting
Anders Oland, Aayush Bansal, Roger B Dannenberg, and Bhiksha Raj · 2017
Closest in time.
The marginal value of adaptive gradient methods in machine learning
Ashia C Wilson, Rebecca Roelofs, Mitchell Stern, Nathan Srebro, and Benjamin Recht · 2017
Closest in time.
Normalized gradient with adaptive stepsize method for deep neural network training
Adams Wei Yu, Qihang Lin, Ruslan Salakhutdinov, and Jaime Carbonell · 2017
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Weight normalization: A simple reparameterization to accelerate training of deep neural networks
Tim Salimans and Diederik P Kingma · 2016
Cited alongside, same era.
A tensorflow implementation of wide residual networks, 2016
Neal Wu · 2016
Cited alongside, same era.
Weight uncertainty in neural networks
Charles Blundell, Julien Cornebise, Koray Kavukcuoglu, and Daan Wierstra
Cited in the paper.
Fixing weight decay regularization in adam
Ilya Loshchilov and Frank Hutter
Cited in the paper.
Sgdr: stochastic gradient descent with restarts
Ilya Loshchilov and Frank Hutter
Cited in the paper.
A pytorch implementation of wide residual networks, 2016a
Sergey Zagoruyko and Nikos Komodakis
Cited in the paper.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2017
Closest in time.
Squeeze-and-excitation networks
Jie Hu, Li Shen, and Gang Sun · 2018
Closest in time.
A bayesian perspective on generalization and stochastic gradient descent
Samuel L Smith and Quoc V Le · 2018
Closest in time.