Fetching the paper…
Reading the bibliography…
In the context of stochastic gradient descent(SGD) and adaptive moment estimation (Adam),researchers have recently proposed optimization techniques that transition from Adam to SGD with the goal of improving both convergence and generalization performance.
Long short-term memory
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
Stability and generalization
Bousquet, O. and Elisseeff, A · 2002
Earlier work this paper cites.
A friendly introduction to numerical analysis
Bradie, B · 2006
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A. and Hinton, G · 2009
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
Duchi, J., Hazan, E., and Singer, Y · 2011
Earlier work this paper cites.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
Tieleman, T. and Hinton, G · 2012
Earlier work this paper cites.
Introduction to online convex optimization
Hazan, E · 2014
Earlier work this paper cites.
Convex optimization: Algorithms and complexity
Bubeck, S · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Earlier work this paper cites.
Show, attend and tell: Neural image caption generation with visual attention
Xu, K., Ba, J., Kiros, R., Cho, K., Courville, A., Salakhutdinov, R., Zemel, R., and Bengio, Y · 2015
Earlier work this paper cites.
Deep learning , chapter 7, pp. 239–245
Goodfellow, I., Bengio, Y., and Courville, A · 2016
Earlier work this paper cites.
Train faster, generalize better: Stability of stochastic gradient descent
Hardt, M., Recht, B., and Singer, Y · 2016
Earlier work this paper cites.
Mimic-iii, a freely accessible critical care database
Johnson, A. E., Pollard, T. J., Shen, L., Li-wei, H. L., Feng, M., Ghassemi, M., Moody, B., Szolovits, P., Celi, L. A., and Mark, R. G · 2016
Earlier work this paper cites.
Pointer sentinel mixture models
Merity, S., Xiong, C., Bradbury, J., and Socher, R · 2016
Cited alongside, same era.
An overview of gradient descent optimization algorithms
Ruder, S · 2016
Cited alongside, same era.
Gastaldi, X · 2017
Cited alongside, same era.
Multitask learning and benchmarking with clinical time series data
Harutyunyan, H., Khachatrian, H., Kale, D. C., Steeg, G. V., and Galstyan, A · 2017
Cited alongside, same era.
Improving generalization performance by switching from adam to sgd
Closing the generalization gap of adaptive gradient methods in training deep neural networks
Chen, J. and Gu, Q · 2018
Later among the works it cites.
Autoaugment: Learning augmentation policies from data
Cubuk, E. D., Zoph, B., Mane, D., Vasudevan, V., and Le, Q. V · 2018
Later among the works it cites.
Gradient descent happens in a tiny subspace
Gur-Ari, G., Roberts, D. A., and Dyer, E · 2018
Later among the works it cites.
Gpipe: Efficient training of giant neural networks using pipeline parallelism
Huang, Y., Cheng, Y., Chen, D., Lee, H., Ngiam, J., Le, Q. V., and Chen, Z · 2018
Later among the works it cites.
On the convergence of stochastic gradient descent with adaptive stepsizes
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Keskar, N. S. and Socher, R · 2017
Cited alongside, same era.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2017
Cited alongside, same era.
Value prediction network
Oh, J., Singh, S., and Lee, H · 2017
Cited alongside, same era.
Automatic differentiation in pytorch
Paszke, A., Gross, S., Chintala, S., Chanan, G., Yang, E., DeVito, Z., Lin, Z., Desmaison, A., Antiga, L., and Lerer, A · 2017
Cited alongside, same era.
Relections on random kitchen sinks
Rahimi, A. and Recht, B · 2017
Cited alongside, same era.
The marginal value of adaptive gradient methods in machine learning
Wilson, A. C., Roelofs, R., Stern, M., Srebro, N., and Recht, B · 2017
Cited alongside, same era.
Fashion-mnist: a novel image dataset for benchmarking machine learning algorithms
Xiao, H., Rasul, K., and Vollgraf, R · 2017
Cited alongside, same era.
Understanding deep learning requires rethinking generalization
Zhang, C., Bengio, S., Hardt, M., Recht, B., and Vinyals, O · 2017
Cited alongside, same era.
Li, X. and Orabona, F · 2018
Later among the works it cites.
On the convergence of adam and beyond
Reddi, S. J., Kale, S., and Kumar, S · 2018
Later among the works it cites.
Adashift: Decorrelation and convergence of adaptive learning rate methods
Zhou, Z., Zhang, Q., Lu, G., Wang, H., Zhang, W., and Yu, Y · 2018
Later among the works it cites.
On the convergence of a class of adam-type algorithms for non-convex optimization
Chen, X., Liu, S., Sun, R., and Hong, M · 2019
Later among the works it cites.
On empirical comparisons of optimizers for deep learning
Choi, D., Shallue, C. J., Nado, Z., Lee, J., Maddison, C. J., and Dahl, G. E · 2019
Later among the works it cites.
An investigation into neural net optimization via hessian eigenvalue density
Ghorbani, B., Krishnan, S., and Xiao, Y · 2019
Later among the works it cites.
A recipe for training neural networks
Karpathy, A · 2019
Later among the works it cites.
Adaptive gradient methods with dynamic bound of learning rate
Luo, L., Xiong, Y., Liu, Y., and Sun, X · 2019
Later among the works it cites.
Painless stochastic gradient: Interpolation, line-search, and convergence rates
Vaswani, S., Mishkin, A., Laradji, I., Schmidt, M., Gidel, G., and Lacoste-Julien, S · 2019
Later among the works it cites.