Fetching the paper…
Reading the bibliography…
Choosing the optimizer is considered to be among the most crucial design decisions in deep learning, and it is not an easy one.
Chen, Y., Jing, H., Zhao, W., Liu, Z., Li, O., Qiao, L., Xue, W., Fu, H., and Yang, G · 1905
Earlier work this paper cites.
NAMSG: An Efficient Method For Training Neural Networks. arXiv preprint: 1905.01422 , 2019c
Chen, Y., Jing, H., Zhao, W., Liu, Z., Qiao, L., Xue, W., Fu, H., and Yang, G · 1905
Earlier work this paper cites.
A Stochastic Approximation Method
Robbins, H. and Monro, S · 1951
Earlier work this paper cites.
Some methods of speeding up the convergence of iteration methods
Polyak, B. T · 1964
Earlier work this paper cites.
A method for solving the convex programming problem with convergence rate O ( 1 / k 2 ) O(1/k^{2})
Nesterov, Y · 1983
Earlier work this paper cites.
No free lunch theorems for optimization
Wolpert, D. H. and Macready, W. G · 1997
Earlier work this paper cites.
Wang, B., Nguyen, T. M., Bertozzi, A. L., Baraniuk, R. G., and Osher, S. J · 2002
Earlier work this paper cites.
Li, W., Zhang, Z., Wang, X., and Luo, P · 2004
Earlier work this paper cites.
Li, Z., Bao, H., Zhang, X., and Richtárik, P · 2008
Earlier work this paper cites.
ImageNet: A Large-Scale Hierarchical Image Database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L · 2009
Earlier work this paper cites.
Chen, R. T. Q., Choi, D., Balles, L., Duvenaud, D., and Hennig, P · 2011
Earlier work this paper cites.
Adaptive Subgradient Methods for Online Learning and Stochastic Optimization
Duchi, J., Hazan, E., and Singer, Y · 2011
Earlier work this paper cites.
Adam + : A Stochastic Method with Adaptive Variance Reduction. arXiv preprint: 2011.11985 , 2020b
Liu, M., Zhang, W., Orabona, F., and Yang, T · 2011
Earlier work this paper cites.
Random Search for Hyper-Parameter Optimization
Bergstra, J. and Bengio, Y · 2012
Earlier work this paper cites.
Stochastic gradient descent tricks
Bottou, L · 2012
Earlier work this paper cites.
Lecture 6.5—RMSProp: Divide the gradient by a running average of its recent magnitude, 2012
Tieleman, T. and Hinton, G · 2012
Earlier work this paper cites.
ADADELTA: An Adaptive Learning Rate Method. arXiv preprint: 1212.5701 , 2012
Zeiler, M. D · 2012
Earlier work this paper cites.
Adaptive learning rates and parallelization for stochastic, sparse, non-smooth gradients
Schaul, T. and LeCun, Y · 2013
Earlier work this paper cites.
No more pesky learning rates
Schaul, T., Zhang, S., and LeCun, Y · 2013
Earlier work this paper cites.
Generative Adversarial Nets
Goodfellow, I. J., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y · 2014
Earlier work this paper cites.
TensorFlow: Large-Scale Machine Learning on Heterogeneous Systems, 2015
Abadi, M., Agarwal, A., Barham, P., Brevdo, E., Chen, Z., Citro, C., Corrado, G. S., Davis, A., Dean, J., Devin, M., Ghemawat, S., Goodfellow, I., Harp, A., Irving, G., Isard, M., Jia, Y., Jozefowicz, R., Kaiser, L., Kudlur, M., Levenberg, J., Mané, D., Monga, R., Moore, S., Murray, D., Olah, C., Schuster, M., Shlens, J., Steiner, B., Sutskever, I., Talwar, K., Tucker, P., Vanhoucke, V., Vasudevan, V., Viégas, F., Vinyals, O., Warden, P., Wattenberg, M., Wicke, M., Yu, Y., and Zheng, X · 2015
Earlier work this paper cites.
Adam: A Method for Stochastic Optimization
Kingma, D. P. and Ba, J · 2015
Earlier work this paper cites.
Optimizing Neural Networks with Kronecker-Factored Approximate Curvature
Martens, J. and Grosse, R · 2015
Earlier work this paper cites.
Scale-Free Algorithms for Online Linear Optimization
Orabona, F. and Pál, D · 2015
Earlier work this paper cites.
Deep Learning
Goodfellow, I., Bengio, Y., and Courville, A · 2016
Earlier work this paper cites.
Incorporating Nesterov Momentum into Adam
Dozat, T · 2016
Earlier work this paper cites.
Deep Learning
Goodfellow, I., Bengio, Y., and Courville, A · 2016
Earlier work this paper cites.
Deep Residual Learning for Image Recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Barzilai-Borwein Step Size for Stochastic Gradient Descent
Tan, C., Ma, S., Dai, Y., and Qian, Y · 2016
Earlier work this paper cites.
Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour. arXiv preprint: 1706.02677 , 2017
Goyal, P., Dollár, P., Girshick, R., Noordhuis, P., Wesolowski, L., Kyrola, A., Tulloch, A., Jia, Y., and He, K · 2017
Earlier work this paper cites.
SGDR: Stochastic Gradient Descent with Warm Restarts
Loshchilov, I. and Hutter, F · 2017
Earlier work this paper cites.
Attention Is All You Need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
The Marginal Value of Adaptive Gradient Methods in Machine Learning
Wilson, A. C., Roelofs, R., Stern, M., Srebro, N., and Recht, B · 2017
Earlier work this paper cites.
Neural Optimizer Search with Reinforcement Learning
Bello, I., Zoph, B., Vasudevan, V., and Le, Q. V · 2017
Earlier work this paper cites.
Practical Gauss-Newton Optimisation for Deep Learning
Botev, A., Ritter, H., and Barber, D · 2017
Earlier work this paper cites.
AdaBatch: Adaptive Batch Sizes for Training Deep Neural Networks. arXiv preprint: 1712.02029 , 2017
Devarakonda, A., Naumov, M., and Garland, M · 2017
Earlier work this paper cites.
Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour. arXiv preprint: 1706.02677 , 2017
Goyal, P., Dollár, P., Girshick, R., Noordhuis, P., Wesolowski, L., Kyrola, A., Tulloch, A., Jia, Y., and He, K · 2017
Earlier work this paper cites.
Adaptive Learning Rate via Covariance Matrix Based Preconditioning for Deep Neural Networks
Ida, Y., Fujiwara, Y., and Iwamura, S · 2017
Earlier work this paper cites.
Keskar, N. S. and Socher, R · 2017
Earlier work this paper cites.
SGDR: Stochastic Gradient Descent with Warm Restarts
Loshchilov, I. and Hutter, F · 2017
Earlier work this paper cites.
Probabilistic Line Searches for Stochastic Optimization
Mahsereci, M. and Hennig, P · 2017
Earlier work this paper cites.
Variants of RMSProp and Adagrad with Logarithmic Regret Bounds
Mukkamala, M. C. and Hein, M · 2017
Earlier work this paper cites.
Cyclical Learning Rates for Training Neural Networks
Smith, L. N · 2017
Earlier work this paper cites.
Smith, L. N. and Topin, N · 2017
Earlier work this paper cites.
Attention Is All You Need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Large Batch Training of Convolutional Networks. arXiv preprint: 1708.03888 , 2017
You, Y., Gitman, I., and Ginsburg, B · 2017
Earlier work this paper cites.
Loss Surfaces, Mode Connectivity, and Fast Ensembling of DNNs
Garipov, T., Izmailov, P., Podoprikhin, D., Vetrov, D. P., and Wilson, A. G · 2018
Earlier work this paper cites.
Averaging weights leads to wider optima and better generalization
Izmailov, P., Podoprikhin, D., Garipov, T., Vetrov, D., and G, W. A · 2018
Earlier work this paper cites.
A Walk with SGD. arXiv preprint: 1802.08770 , 2018
Xing, C., Arpit, D., Tsirigotis, C., and Bengio, Y · 2018
Earlier work this paper cites.
Dissecting Adam: The Sign, Magnitude and Variance of Stochastic Gradients
Balles, L. and Hennig, P · 2018
Earlier work this paper cites.
SIGNSGD: Compressed Optimisation for Non-Convex Problems
Bernstein, J., Wang, Y., Azizzadenesheli, K., and Anandkumar, A · 2018
Earlier work this paper cites.
AdaComp: Adaptive Residual Gradient Compression for Data-Parallel Distributed Training
Chen, C., Choi, J., Brand, D., Agrawal, A., Zhang, W., and Gopalakrishnan, K · 2018
Earlier work this paper cites.
Fast Approximate Natural Gradient Descent in a Kronecker Factored Eigenbasis
George, T., Laurent, C., Bouthillier, X., Ballas, N., and Vincent, P · 2018
Earlier work this paper cites.
Shampoo: Preconditioned Stochastic Tensor Optimization
Gupta, V., Koren, T., and Singer, Y · 2018
Earlier work this paper cites.
Hayashi, H., Koushik, J., and Neubig, G · 2018
Earlier work this paper cites.
Universal Language Model Fine-tuning for Text Classification
Howard, J. and Ruder, S · 2018
Cited alongside, same era.
Fast and Scalable Bayesian Deep Learning by Weight-Perturbation in Adam
Khan, M. E., Nielsen, D., Tangkaratt, V., Lin, W., Gal, Y., and Srivastava, A · 2018
Cited alongside, same era.
Online Adaptive Methods, Universality and Acceleration
Levy, K. Y., Yurtsever, A., and Cevher, V · 2018
Cited alongside, same era.
On the Convergence of Adam and Beyond
Reddi, S. J., Kale, S., and Kumar, S · 2018
Cited alongside, same era.
L4: Practical loss-based stepsize adaptation for deep learning
Rolínek, M. and Martius, G · 2018
Cited alongside, same era.
Daley, B. and Amato, C · 2020
Closest in time.
diffGrad: An Optimization Method for Convolutional Neural Networks
Dubey, S. R., Chakraborty, S., Roy, S. K., Mukherjee, S., Singh, S. K., and Chaudhuri, B. B · 2020
Closest in time.
ADAS Optimzier
Eliyahu, Y · 2020
Closest in time.
Gao, K.-X., Liu, X.-L., Huang, Z.-H., Wang, M., Wang, S., Wang, Z., Xu, D., and Yu, F · 2020
Closest in time.
Practical Quasi-Newton Methods for Training Deep Neural Networks
Goldfarb, D., Ren, Y., and Bahamou, A · 2020
Closest in time.
RangerLars
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Salas, A., Kessler, S., Zohren, S., and Roberts, S · 2018
Cited alongside, same era.
Adafactor: Adaptive Learning Rates with Sublinear Memory Cost
Shazeer, N. and Stern, M · 2018
Cited alongside, same era.
WNGrad: Learn the Learning Rate in Gradient Descent. arXiv preprint: 1803.02865 , 2018
Wu, X., Ward, R., and Bottou, L · 2018
Cited alongside, same era.
A Walk with SGD. arXiv preprint: 1802.08770 , 2018
Xing, C., Arpit, D., Tsirigotis, C., and Bengio, Y · 2018
Cited alongside, same era.
Adaptive Methods for Nonconvex Optimization
Zaheer, M., Reddi, S. J., Sachan, D. S., Kale, S., and Kumar, S · 2018
Cited alongside, same era.
Noisy Natural Gradient as Variational Inference
Zhang, G., Sun, S., Duvenaud, D., and Grosse, R · 2018
Cited alongside, same era.
Zhang, J. and Gouza, F. B · 2018
Cited alongside, same era.
Grankin, M · 2020
Closest in time.
Granziol, D., Wan, X., and Roberts, S · 2020
Closest in time.
AdaS: Adaptive Scheduling of Stochastic Gradients. arXiv preprint: 2006.06587 , 2020
Hosseini, M. S. and Plataniotis, K. N · 2020
Closest in time.
Biased Stochastic First-Order Methods for Conditional Stochastic Optimization and Applications in Meta Learning
Hu, Y., Zhang, S., Chen, X., and He, N · 2020
Closest in time.
Huang, X., Zhou, H., Xu, R., Wang, Z., and Li, L · 2020
Closest in time.
TAdam: A Robust Stochastic Gradient Optimizer. arXiv preprint: 2003.00179 , 2020
Ilboudo, W. E. L., Kobayashi, T., and Sugimoto, K · 2020
Closest in time.
AdaScale SGD: A User-Friendly Algorithm for Distributed Training
Johnson, T. B., Agrawal, P., Gu, H., and Guestrin, C · 2020
Closest in time.
Kelterborn, C., Mazur, M., and Petrenko, B. V · 2020
Closest in time.
Mixing ADAM and SGD: a Combined Optimization Method. arXiv preprint: 2011.08042 , 2020
Landro, N., Gallo, I., and Grassa, R. L · 2020
Closest in time.
AEGD: Adaptive Gradient Decent with Energy. arXiv preprint: 2010.05109 , 2020
Liu, H. and Tian, X · 2020
Closest in time.
A New Accelerated Stochastic Gradient Method with Momentum. arXiv preprint: 2006.00423 , 2020
Liu, L. and Luo, X · 2020
Closest in time.
MTAdam: Automatic Balancing of Multiple Training Loss Terms. arXiv preprint: 2006.14683 , 2020
Malkiel, I. and Wolf, L · 2020
Closest in time.
Parabolic Approximation Line Search for DNNs
Mutschler, M. and Zell, A · 2020
Closest in time.
Purkayastha, S. and Purkayastha, S · 2020
Closest in time.
VR-SGD: A Simple Stochastic Variance Reduction Method for Machine Learning
Shang, F., Zhou, K., Liu, H., Cheng, J., Tsang, I. W., Zhang, L., Tao, D., and Jiao, L · 2020
Closest in time.
Sung, W., Choi, I., Park, J., Choi, S., and Shin, S · 2020
Closest in time.
Shuffling Gradient-Based Methods with Momentum. arXiv preprint: 2011.11884 , 2020
Tran, T. H., Nguyen, L. M., and Tran-Dinh, Q · 2020
Closest in time.
Compositional ADAM: An Adaptive Compositional Solver. arXiv preprint: 2002.03755 , 2020
Tutunov, R., Li, M., Cowen-Rivers, A. I., Wang, J., and Bou-Ammar, H · 2020
Closest in time.
Wang, B. and Ye, Q · 2020
Closest in time.
AdaSGD: Bridging the gap between SGD and Adam. arXiv preprint: 2006.16541 , 2020
Wang, J. and Wiens, J · 2020
Closest in time.
Xie, Z., Wang, X., Zhang, H., Sato, I., and Sugiyama, M · 2020
Closest in time.
Xu, Y · 2020
Closest in time.
Yang, M., Xu, D., Li, Y., Wen, Z., and Chen, M · 2020
Closest in time.
Yao, Z., Gholami, A., Shen, S., Keutzer, K., and Mahoney, M. W · 2020
Closest in time.
Large Batch Optimization for Deep Learning: Training BERT in 76 minutes
You, Y., Li, J., Reddi, S., Hseu, J., Kumar, S., Bhojanapalli, S., Song, X., Demmel, J., Keutzer, K., and Hsieh, C.-J · 2020
Closest in time.
EAdam Optimizer: How ϵ \epsilon Impact Adam. arXiv preprint: 2011.02150 , 2020
Yuan, W. and Gao, K.-X · 2020
Closest in time.
SALR: Sharpness-aware Learning Rates for Improved Generalization. arXiv preprint: 2011.05348 , 2020
Yue, X., Nouiehed, M., and Kontar, R. A · 2020
Closest in time.
Why are adaptive methods good for attention models?
Zhang, J., Karimireddy, S. P., Veit, A., Kim, S., Reddi, S. J., Kumar, S., and Sra, S · 2020
Closest in time.
Zhao, S.-Y., Xie, Y.-P., and Li, W.-J · 2020
Closest in time.
On the Trend-corrected Variant of Adaptive Stochastic Optimization Methods
Zhou, B., Zheng, X., and Gao, J · 2020
Closest in time.
AdaBelief Optimizer: Adapting Stepsizes by the Belief in Observed Gradients
Zhuang, J., Tang, T., Ding, Y., Tatikonda, S., Dvornek, N., Papademetris, X., and Duncan, J. S · 2020
Closest in time.
LaProp: a Better Way to Combine Momentum with Adaptive Gradient. arXiv preprint: 2002.04839 , 2020
Ziyin, L., Wang, Z. T., and Ueda, M · 2020
Closest in time.
A Generalizable Approach to Learning Optimizers. arXiv preprint: 2106.00958 , 2021
Almeida, D., Winter, C., Tang, J., and Zaremba, W · 2021
Closest in time.
Bahrami, D. and Zadeh, S. P · 2021
Closest in time.
Second-order step-size tuning of SGD for non-convex optimization. arXiv preprint: 2103.03570 , 2021
Castera, C., Bolte, J., Févotte, C., and Pauwels, E · 2021
Closest in time.
Chae, Y., Wilke, D. N., and Kafka, D · 2021
Closest in time.
Chakrabarti, K. and Chopra, N · 2021
Closest in time.
CADA: Communication-Adaptive Distributed Adam
Chen, T., Guo, Z., Sun, Y., and Yin, W · 2021
Closest in time.
de Roos, F., Jidling, C., Wills, A., Schön, T., and Hennig, P · 2021
Closest in time.
Defazio, A. and Jelassi, S · 2021
Closest in time.
Dellaferrera, G., Wozniak, S., Indiveri, G., Pantazi, A., and Eleftheriou, E · 2021
Closest in time.
Sharpness-aware Minimization for Efficiently Improving Generalization
Foret, P., Kleiner, A., Mobahi, H., and Neyshabur, B · 2021
Closest in time.
AdamP: Slowing Down the Weight Norm Increase in Momentum-based Optimizers
Heo, B., Chun, S., Oh, S. J., Han, D., Yun, S., Uh, Y., and Ha, J.-W · 2021
Closest in time.
AutoLRS: Automatic Learning-Rate Schedule by Bayesian Optimization on the Fly
Jin, Y., Zhou, T., Zhao, L., Zhu, Y., Guo, C., Canini, M., and Krishnamurthy, A · 2021
Closest in time.
Kwon, J., Kim, J., Park, H., and Choi, I. K · 2021
Closest in time.
Ramezani-Kebrya, A., Khisti, A., and Liang, B · 2021
Closest in time.
Ren, Y. and Goldfarb, D · 2021
Closest in time.
Roy, S. K., Paoletti, M. E., Haut, J. M., Dubey, S. R., Kar, P., Plaza, A., and Chaudhuri, B. B · 2021
Closest in time.
Correcting Momentum with Second-order Information. arXiv preprint: 2103.03265 , 2021
Tran, H. and Cutkosky, A · 2021
Closest in time.