Fetching the paper…
Reading the bibliography…
Despite superior training outcomes, adaptive optimization methods such as Adam, Adagrad or RMSprop have been found to generalize poorly compared to Stochastic gradient descent (SGD).
A stochastic approximation method
Robbins, Herbert and Monro, Sutton · 1951
Earlier work this paper cites.
Long short-term memory
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
Numerical optimization
Nocedal, J. and Wright, S · 2006
Earlier work this paper cites.
ImageNet: A Large-Scale Hierarchical Image Database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A. and Hinton, G · 2009
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
Duchi, J., Hazan, E., and Singer, Y · 2011
Earlier work this paper cites.
RNNLM-recurrent neural network language modeling toolkit
Mikolov, T., Kombrink, S., Deoras, A., Burget, L., and Cernocky, J · 2011
Earlier work this paper cites.
Lecture 6.5-RMSProp: Divide the gradient by a running average of its recent magnitude
Tieleman, T. and Hinton, G · 2012
Earlier work this paper cites.
On the importance of initialization and momentum in deep learning
Sutskever, I., Martens, J., Dahl, G., and Hinton, G · 2013
Earlier work this paper cites.
Regularization of neural networks using dropconnect
Wan, L., Zeiler, M., Zhang, S., LeCun, Y, and Fergus, R · 2013
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Bahdanau, D., Cho, K., and Bengio, Y · 2014
Earlier work this paper cites.
Densenet: Implementing efficient convnet descriptor pyramids
Iandola, F., Moskewicz, M., Karayev, S., Girshick, R., Darrell, T., and Keutzer, K · 2014
Cited alongside, same era.
Train faster, generalize better: Stability of stochastic gradient descent
Hardt, M., Recht, B., and Singer, Y · 2015
Cited alongside, same era.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2015
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, D. and Ba, J · 2015
Cited alongside, same era.
Quasi-Recurrent Neural Networks
Bradbury, J., Merity, S., Xiong, C., and Socher, R · 2016
Cited alongside, same era.
Extremely large minibatch SGD: Training resnet-50 on ImageNet in 15 minutes
Akiba, T., Suzuki, S., and Fukuda, K · 2017
Closest in time.
Squeeze-and-excitation networks
Hu, J., Shen, L., and Sun, G · 2017
Closest in time.
A Peek at Trends in Machine Learning
Karpathy, A · 2017
Closest in time.
Fixing Weight Decay Regularization in Adam
Loshchilov, I. and Hutter, F · 2017
Closest in time.
Stochastic Gradient Descent as Approximate Bayesian Inference
Mandt, S., Hoffman, M. D., and Blei, D. M · 2017
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Deep pyramidal residual networks
Han, D., Kim, J., and Kim, J · 2016
Cited alongside, same era.
Gradient descent converges to minimizers
Lee, J., Simchowitz, M., Jordan, M. I, and Recht, B · 2016
Cited alongside, same era.
SGDR: Stochastic gradient descent with warm restarts
Loshchilov, I. and Hutter, F · 2016
Cited alongside, same era.
Pointer sentinel mixture models
Merity, S., Xiong, C., Bradbury, J., and Socher, R · 2016
Cited alongside, same era.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Wu, Y., Schuster, M., Chen, Z., Le, Q., Norouzi, M., Macherey, W., Krikun, M., Cao, Y., Gao, Q., Macherey, K., et al · 2016
Cited alongside, same era.
Understanding deep learning requires rethinking generalization
Zhang, C., Bengio, S., Hardt, M., Recht, B., and Vinyals, O · 2016
Cited alongside, same era.
Merity, S., Keskar, N., and Socher, R · 2017
Closest in time.
The implicit bias of gradient descent on separable data
Soudry, D., Hoffer, E., and Srebro, N · 2017
Closest in time.
The Marginal Value of Adaptive Gradient Methods in Machine Learning
Wilson, A. C., Roelofs, R., Stern, M., Srebro, N., and Recht, B · 2017
Closest in time.
Normalized direction-preserving Adam
Zhang, Z., Ma, L., Li, Z., and Wu, C · 2017
Closest in time.
On the convergence of Adam and beyond
Anonymous · 2018
Closest in time.