Fetching the paper…
Reading the bibliography…
The performance of optimizers, particularly in deep learning, depends considerably on their chosen hyperparameter configuration.
Schwartz, R., Dodge, J., Smith, N. A., and Etzioni, O · 1907
Earlier work this paper cites.
A stochastic approximation method
Robbins, H. and Monro, S · 1951
Earlier work this paper cites.
An introduction to the bootstrap
Tibshirani, R. J. and Efron, B · 1993
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
Duchi, J., Hazan, E., and Singer, Y · 2011
Earlier work this paper cites.
Learning word vectors for sentiment analysis
Maas, A. L., Daly, R. E., Pham, P. T., Huang, D., Ng, A. Y., and Potts, C · 2011
Earlier work this paper cites.
Random search for hyper-parameter optimization
Bergstra, J. and Bengio, Y · 2012
Earlier work this paper cites.
Lecture 6.5—RmsProp: Divide the gradient by a running average of its recent magnitude
Tieleman, T. and Hinton, G · 2012
Earlier work this paper cites.
On the importance of initialization and momentum in deep learning
Sutskever, I., Martens, J., Dahl, G., and Hinton, G · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. and Ba, J · 2015
Earlier work this paper cites.
Parsenet: Looking wider to see better
Liu, W., Rabinovich, A., and Berg, A. C · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Cited alongside, same era.
Closing the generalization gap of adaptive gradient methods in training deep neural networks
Chen, J. and Gu, Q · 2018
Cited alongside, same era.
Deep reinforcement learning that matters
Henderson, P., Islam, R., Bachman, P., Pineau, J., Precup, D., and Meger, D · 2018
Cited alongside, same era.
Are gans created equal? a large-scale study
Lucic, M., Kurach, K., Michalski, M., Gelly, S., and Bousquet, O · 2018
Cited alongside, same era.
An empirical study on hyperparameter tuning of decision trees
Mantovani, R. G., Horváth, T., Cerri, R., Junior, S. B., Vanschoren, J., de Carvalho, A. C. P. d., and Ferreira, L · 2018
Cited alongside, same era.
On empirical comparisons of optimizers for deep learning
Choi, D., Shallue, C. J., Nado, Z., Lee, J., Maddison, C. J., and Dahl, G. E · 2019
Closest in time.
Show your work: Improved reporting of experimental results
Dodge, J., Gururangan, S., Card, D., Schwartz, R., and Smith, N. A · 2019
Closest in time.
Pitfalls and best practices in algorithm configuration
Eggensperger, K., Lindauer, M., and Hutter, F · 2019
Closest in time.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2019
Closest in time.
Tunability: Importance of hyperparameters of machine learning algorithms
Probst, P., Boulesteix, A.-L., and Bischl, B · 2019
Closest in time.
DeepOBS: A deep learning optimizer benchmark suite
Schneider, F., Balles, L., and Hennig, P · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
On the state of the art of evaluation in neural language models
Melis, G., Dyer, C., and Blunsom, P · 2018
Cited alongside, same era.
Winner’s curse? on pace, progress, and empirical rigor
Sculley, D., Snoek, J., Wiltschko, A. B., and Rahimi, A · 2018
Cited alongside, same era.
Minimum norm solutions do not always generalize well for over-parameterized problems
Shah, V., Kyrillidis, A., and Sanghavi, S · 2018
Cited alongside, same era.
Don’t decay the learning rate, increase the batch size
Smith, S. L., Kindermans, P.-J., and Le, Q. V · 2018
Cited alongside, same era.
The importance of better models in stochastic optimization
Asi, H. and Duchi, J. C · 2019
Cited alongside, same era.
Stochastic (approximate) proximal point methods: Convergence, optimality, and adaptivity
Asi, H. and Duchi, J. C
Cited in the paper.
Automated machine learning-methods, systems, challenges, 2019a
Hutter, F., Kotthoff, L., and Vanschoren, J
Cited in the paper.
Closest in time.
Measuring the effects of data parallelism on neural network training
Shallue, C. J., Lee, J., Antognini, J., Sohl-Dickstein, J., Frostig, R., and Dahl, G. E · 2019
Closest in time.
Energy and policy considerations for deep learning in NLP
Strubell, E., Ganesh, A., and McCallum, A · 2019
Closest in time.
The marginal value of adaptive gradient methods in machine learning
Wilson, A. C., Roelofs, R., Stern, M., Srebro, N., and Recht, B · 2019
Closest in time.
Using a thousand optimization tasks to learn hyperparameter search strategies
Metz, L., Maheswaranathan, N., Sun, R., Freeman, C. D., Poole, B., and Sohl-Dickstein, J · 2020
Closest in time.