Fetching the paper…
Reading the bibliography…
In deep learning, different kinds of deep networks typically need different optimizers, which have to be chosen after multiple trials, making the training process inefficient.
H. Robbins and S. Monro, “A stochastic approximation method,” The Annals of Mathematical Statistics , pp. 400–407, 1951
1951
Earlier work this paper cites.
B. T. Polyak, “Some methods of speeding up the convergence of iteration methods,” Ussr Computational Mathematics and Mathematical Physics , vol. 4, no. 5, pp. 1–17, 1964
1964
Earlier work this paper cites.
Y. Nesterov, “A method for solving the convex programming problem with convergence rate 𝒪 ( 1 / k 2 ) \order{1/k^{2}} ,” in Doklady Akademii Nauk , vol. 269, no. 3. Russian Academy of Sciences, 1983, pp. 543–547
1983
Earlier work this paper cites.
Y. Nesterov, “On an approach to the construction of optimal methods of minimization of smooth convex functions,” Ekonomika i Mateaticheskie Metody , vol. 24, no. 3, pp. 509–517, 1988
1988
Earlier work this paper cites.
M. A. Marcinkiewicz, “Building a large annotated corpus of english: The penn treebank,” Using Large Corpora , vol. 273, 1994
1994
Earlier work this paper cites.
J. Schmidhuber, S. Hochreiter et al. , “Long short-term memory,” Neural Comput , vol. 9, no. 8, pp. 1735–1780, 1997
1997
Earlier work this paper cites.
D. Saad, “Online algorithms and stochastic approximations,” Online Learning , vol. 5, pp. 6–3, 1998
1998
Earlier work this paper cites.
Y. Nesterov, Introductory lectures on convex optimization: A basic course . Springer Science & Business Media, 2003, vol. 87
2003
Earlier work this paper cites.
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” in Proceedings of the IEEE conference on Computer Vision and Pattern Recognition , 2009, pp. 248–255
2009
Earlier work this paper cites.
A. Krizhevsky, G. Hinton et al. , “Learning multiple layers of features from tiny images,” 2009
2009
Earlier work this paper cites.
J. Duchi, E. Hazan, and Y. Singer, “Adaptive subgradient methods for online learning and stochastic optimization,” Journal of Machine Learning Research , vol. 12, no. 7, 2011
2011
Earlier work this paper cites.
T. Tijmen and H. Geoffrey, “Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude,” COURSERA: Neural Networks for Machine Learning , vol. 4, 2012
2012
Earlier work this paper cites.
E. Todorov, T. Erez, and Y. Tassa, “Mujoco: A physics engine for model-based control,” in 2012 IEEE/RSJ International Conference on Intelligent Robots and Systems . IEEE, 2012, pp. 5026–5033
2012
Earlier work this paper cites.
T. N. Sainath, B. Kingsbury, A.-r. Mohamed, G. E. Dahl, G. Saon, H. Soltau et al. , “Improvements to deep convolutional neural networks for LVCSR,” in 2013 IEEE Workshop on Automatic Speech Recognition and Understanding . IEEE, 2013, pp. 315–320
2013
Earlier work this paper cites.
R. Johnson and T. Zhang, “Accelerating stochastic gradient descent using predictive variance reduction,” Advances in neural information processing systems , vol. 26, 2013
2013
Earlier work this paper cites.
O. Abdel-Hamid, A. Mohamed, H. Jiang, L. Deng, G. Penn, and D. Yu, “Convolutional neural networks for speech recognition,” IEEE Trans. on Audio, Speech, and Language Processing , vol. 22, no. 10, pp. 1533–1545, 2014
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
N. Parikh and S. Boyd, “Proximal algorithms,” Foundations and Trends in optimization , vol. 1, no. 3, pp. 127–239, 2014
2014
Earlier work this paper cites.
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan et al. , “Microsoft coco: Common objects in context,” in European Conference on Computer Vision , 2014, pp. 740–755
2014
Earlier work this paper cites.
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov et al. , “Going deeper with convolutions,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2015, pp. 1–9
2015
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2016, pp. 770–778
2016
Earlier work this paper cites.
T. Dozat, “Incorporating nesterov momentum into Adam,” in International Conference on Learning Representations Workshops , 2016
2016
Earlier work this paper cites.
G. Huang, Y. Sun, Z. Liu, D. Sedra, and K. Q. Weinberger, “Deep networks with stochastic depth,” in European Conference on Computer Vision , 2016, pp. 646–661
2016
Earlier work this paper cites.
Y. Duan, X. Chen, R. Houthooft, J. Schulman, and P. Abbeel, “Benchmarking deep reinforcement learning for continuous control,” in International conference on machine learning . PMLR, 2016, pp. 1329–1338
2016
Earlier work this paper cites.
2017
Earlier work this paper cites.
B. Xie, Y. Liang, and L. Song, “Diverse neural network learns true target functions,” in Artificial Intelligence and Statistics . PMLR, 2017, pp. 1216–1224
2017
Earlier work this paper cites.
Y. Li and Y. Yuan, “Convergence analysis of two-layer neural networks with ReLU activation,” Advances in Neural Information Processing Systems , vol. 30, 2017
2017
Earlier work this paper cites.
K. He, G. Gkioxari, P. Dollár, and R. Girshick, “Mask R-CNN,” in Proceedings of the IEEE International Conference on Computer Vision , 2017, pp. 2961–2969
2017
Earlier work this paper cites.
S. J. Reddi, S. Kale, and S. Kumar, “On the convergence of Adam and beyond,” in International Conference on Learning Representations , 2018
2018
Earlier work this paper cites.
X. Chen, S. Liu, R. Sun, and M. Hong, “On the convergence of a class of Adam-type algorithms for non-convex optimization,” in International Conference on Learning Representations , 2018
2018
Earlier work this paper cites.
L. Luo, Y. Xiong, Y. Liu, and X. Sun, “Adaptive gradient methods with dynamic bound of learning rate,” in International Conference on Learning Representations , 2018
2018
Earlier work this paper cites.
I. Loshchilov and F. Hutter, “Decoupled weight decay regularization,” in International Conference on Learning Representations , 2018
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
H. Zhang, M. Cisse, Y. N. Dauphin, and D. Lopez-Paz, “Mixup: Beyond empirical risk minimization,” in International Conference on Learning Representations , 2018
2018
Earlier work this paper cites.
Z. Charles and D. Papailiopoulos, “Stability and generalization of learning algorithms that converge to global optima,” in International Conference on Machine Learning . PMLR, 2018, pp. 745–754
2018
Earlier work this paper cites.
C. Jin, P. Netrapalli, and M. I. Jordan, “Accelerated gradient descent escapes saddle points faster than gradient descent,” in Conference On Learning Theory . PMLR, 2018, pp. 1042–1085
2018
Cited alongside, same era.
M. Zaheer, S. Reddi, D. Sachan, S. Kale, and S. Kumar, “Adaptive methods for nonconvex optimization,” Advances in Neural Information Processing Systems , vol. 31, 2018
2018
Cited alongside, same era.
S. S. Du, W. Hu, and J. D. Lee, “Algorithmic regularization in learning deep homogeneous models: Layers are automatically balanced,” Advances in Neural Information Processing Systems , vol. 31, 2018
2018
Cited alongside, same era.
L. Liu, H. Jiang, P. He, W. Chen, X. Liu, J. Gao et al. , “On the variance of the adaptive learning rate and beyond,” in International Conference on Learning Representations , 2019
2019
Cited alongside, same era.
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y. You, J. Li, S. Reddi, J. Hseu, S. Kumar, S. Bhojanapalli et al. , “Large batch optimization for deep learning: Training bert in 76 minutes,” in International Conference on Learning Representations , 2019
2019
Cited alongside, same era.
Z. Dai, Z. Yang, Y. Yang, J. G. Carbonell, Q. Le, and R. Salakhutdinov, “Transformer-xl: Attentive language models beyond a fixed-length context,” in Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics , 2019, pp. 2978–2988
2019
Cited alongside, same era.
J. D. M.-W. C. Kenton and L. K. Toutanova, “BERT: Pre-training of deep bidirectional transformers for language understanding,” in Proceedings of NAACL-HLT , 2019, pp. 4171–4186
2019
Cited alongside, same era.
R. Wightman, “Pytorch image models,” https://github.com/rwightman/pytorch-image-models , 2019
2019
Cited alongside, same era.
S. Yun, D. Han, S. J. Oh, S. Chun, J. Choe, and Y. Yoo, “Cutmix: Regularization strategy to train strong classifiers with localizable features,” in Proceedings of the IEEE International Conference on Computer Vision , 2019, pp. 6023–6032
2019
Cited alongside, same era.
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever et al. , “Language models are unsupervised multitask learners,” OpenAI blog , vol. 1, no. 8, p. 9, 2019
2019
Cited alongside, same era.
2019
Cited alongside, same era.
M. Ott, S. Edunov, A. Baevski, A. Fan, S. Gross, N. Ng, D. Grangier, and M. Auli, “fairseq: A fast, extensible toolkit for sequence modeling,” in Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics (Demonstrations) , 2019, pp. 48–53
2019
Cited alongside, same era.
P. Foret, A. Kleiner, H. Mobahi, and B. Neyshabur, “Sharpness-aware minimization for efficiently improving generalization,” in International Conference on Learning Representations , 2021
2021
Later among the works it cites.
J. Chen, D. Zhou, Y. Tang, Z. Yang, Y. Cao, and Q. Gu, “Closing the generalization gap of adaptive gradient methods in training deep neural networks,” in Proceedings of the Twenty-Ninth International Conference on International Joint Conferences on Artificial Intelligence , 2021, pp. 3267–3275
2021
Later among the works it cites.
J. Kwon, J. Kim, H. Park, and I. K. Choi, “Asam: Adaptive sharpness-aware minimization for scale-invariant learning of deep neural networks,” in International Conference on Machine Learning . PMLR, 2021, pp. 5905–5914
2021
Later among the works it cites.
J. Wang, H. Chen, L. Ma, L. Chen, X. Gong, and W. Liu, “Sphere loss: Learning discriminative features for scene classification in a hyperspherical feature space,” IEEE Transactions on Geoscience and Remote Sensing , vol. 60, pp. 1–19, 2021
2021
Later among the works it cites.
P. Zhou, H. Yan, X. Yuan, J. Feng, and S. Yan, “Towards understanding why Lookahead generalizes better than SGD and beyond,” in Advances in Neural Information Processing Systems , 2021
2021
Later among the works it cites.
Q. Nguyen, M. Mondelli, and G. F. Montufar, “Tight bounds on the smallest eigenvalue of the neural tangent kernel for deep ReLU networks,” in International Conference on Machine Learning , 2021, pp. 8119–8129
2021
Later among the works it cites.
2021
Later among the works it cites.
X. Zhu, W. Su, L. Lu, B. Li, X. Wang, and J. Dai, “Deformable DETR: Deformable transformers for end-to-end object detection,” in International Conference on Learning Representations , 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
T. Xiao, M. Singh, E. Mintun, T. Darrell, P. Dollár, and R. Girshick, “Early convolutions help transformers see better,” Advances in Neural Information Processing Systems , vol. 34, pp. 30 392–30 400, 2021
2021
Later among the works it cites.
Z. Liu, H. Mao, C.-Y. Wu, C. Feichtenhofer, T. Darrell, and S. Xie, “A convnet for the 2020s,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2022, pp. 11 976–11 986
2022
Closest in time.
W. Yu, M. Luo, P. Zhou, C. Si, Y. Zhou, X. Wang et al. , “Metaformer is actually what you need for vision,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2022, pp. 10 819–10 829
2022
Closest in time.
Y. Liu, S. Mai, X. Chen, C.-J. Hsieh, and Y. You, “Towards efficient and scalable sharpness-aware minimization,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2022, pp. 12 360–12 370
2022
Closest in time.
Y. Arjevani, Y. Carmon, J. C. Duchi, D. J. Foster, N. Srebro, and B. Woodworth, “Lower bounds for non-convex stochastic optimization,” Mathematical Programming , pp. 1–50, 2022
2022
Closest in time.
K. He, X. Chen, S. Xie, Y. Li, P. Dollár, and R. Girshick, “Masked autoencoders are scalable vision learners,” in Proceedings of the IEEE conference on Computer Vision and Pattern Recognition , 2022
2022
Closest in time.
J. Du, H. Yan, J. Feng, J. T. Zhou, L. Zhen, R. S. M. Goh et al. , “Efficient sharpness-aware minimization for improved training of neural networks,” in International Conference on Learning Representations , 2022
2022
Closest in time.
Z. Xie, X. Wang, H. Zhang, I. Sato, and M. Sugiyama, “Adaptive inertia: Disentangling the effects of adaptive learning rate and momentum,” in International Conference on Machine Learning . PMLR, 2022, pp. 24 430–24 459
2022
Closest in time.
X. Xie, Q. Wang, Z. Ling, X. Li, G. Liu, and Z. Lin, “Optimization induced equilibrium networks: An explicit optimization perspective for understanding equilibrium models,” IEEE Transactions on Pattern Analysis and Machine Intelligence , 2022
2022
Closest in time.
Z. Zhuang, M. Liu, A. Cutkosky, and F. Orabona, “Understanding AdamW through proximal methods and scale-freeness,” Transactions on Machine Learning Research , 2022
2022
Closest in time.
H. Li and Z. Lin, “Restarted nonconvex accelerated gradient descent: No more polylogarithmic factor in the 𝒪 ( ϵ − 7 / 4 ) \order{\epsilon^{-7/4}} complexity,” in International Conference on Machine Learning . PMLR, 2022, pp. 12 901–12 916
2022
Closest in time.
2022
Closest in time.
X. Chen, C.-J. Hsieh, and B. Gong, “When vision transformers outperform resnets without pre-training or strong data augmentations,” in International Conference on Learning Representation , 2022
2022
Closest in time.
D. Kocetkov, R. Li, L. Ben Allal, J. Li, C. Mou, C. Muñoz Ferrandis et al. , “The Stack: 3 tb of permissively licensed source code,” Preprint , 2022
2022
Closest in time.
J. Weng, H. Chen, D. Yan, K. You, A. Duburcq, M. Zhang et al. , “Tianshou: A highly modularized deep reinforcement learning library,” Journal of Machine Learning Research , vol. 23, no. 267, pp. 1–6, 2022
2022
Closest in time.
T. Computer, “Redpajama: an open dataset for training large language models,” 2023. [Online]. Available: https://github.com/togethercomputer/RedPajama-Data
2023
Closest in time.
H. Touvron, M. Cord, and H. Jégou, “Deit III: Revenge of the ViT,” in European Conference on Computer Vision . Springer, 2022, pp. 516–533
2023
Closest in time.
P. Zhou, X. Xie, Z. Lin, K.-C. Toh, and S. Yan, “Win: Weight-decay-integrated nesterov acceleration for faster network training,” Journal of Machine Learning Research , vol. 25, no. 83, pp. 1–74, 2024
2024
Closest in time.
X. Xie, J. Wu, G. Liu, and Z. Lin, “Sscnet: learning-based subspace clustering,” Visual Intelligence , vol. 2, no. 1, p. 11, 2024
2024
Closest in time.
P. Zhou, X. Xie, Z. Lin, and S. Yan, “Towards understanding convergence and generalization of adamw,” IEEE Transactions on Pattern Analysis and Machine Intelligence , 2024
2024
Closest in time.