Fetching the paper…
Reading the bibliography…
Adam with decoupled weight decay, also known as AdamW, is widely acclaimed for its superior performance in language modeling tasks, surpassing Adam with $\ell_2$ regularization in terms of generalization and optimization.
A stochastic approximation method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
An algorithm for quadratic programming
Marguerite Frank, Philip Wolfe, et al · 1956
Earlier work this paper cites.
Learning representations by back-propagating errors
David E Rumelhart, Geoffrey E Hinton, and Ronald J Williams · 1986
Earlier work this paper cites.
Learning translation invariant recognition in a massively parallel networks
Geoffrey E Hinton · 1987
Earlier work this paper cites.
A simple weight decay can improve generalization
Anders Krogh and John Hertz · 1991
Earlier work this paper cites.
Building a large annotated corpus of English: The Penn Treebank
Mitch Marcus, Beatrice Santorini, and Mary Ann Marcinkiewicz · 1993
Earlier work this paper cites.
Using weight decay to optimize the generalization ability of a perceptron
Siegfried Bos and E Chug · 1996
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Lecture 6.5-rmsprop, coursera: Neural networks for machine learning
Tijmen Tieleman and Geoffrey Hinton · 2012
Earlier work this paper cites.
ADADELTA: an adaptive learning rate method
Matthew D Zeiler · 2012
Earlier work this paper cites.
Revisiting Frank-Wolfe: Projection-free sparse convex optimization
Martin Jaggi · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Implicit regularization in matrix factorization
Suriya Gunasekar, Blake E Woodworth, Srinadh Bhojanapalli, Behnam Neyshabur, and Nati Srebro · 2017
Earlier work this paper cites.
The marginal value of adaptive gradient methods in machine learning
Ashia C Wilson, Rebecca Roelofs, Mitchell Stern, Nati Srebro, and Benjamin Recht · 2017
Earlier work this paper cites.
Theoretical Analysis of Auto Rate-Tuning by Batch Normalization
Sanjeev Arora, Zhiyuan Li, and Kaifeng Lyu · 2018
Earlier work this paper cites.
Dissecting adam: The sign, magnitude and variance of stochastic gradients
Lukas Balles and Philipp Hennig · 2018
Earlier work this paper cites.
signSGD: Compressed optimisation for non-convex problems
Jeremy Bernstein, Yu-Xiang Wang, Kamyar Azizzadenesheli, and Animashree Anandkumar · 2018
Cited alongside, same era.
On the Convergence of A Class of Adam-Type Algorithms for Non-Convex Optimization
Xiangyi Chen, Sijia Liu, Ruoyu Sun, and Mingyi Hong · 2018
Cited alongside, same era.
Characterizing implicit bias in terms of optimization geometry
Suriya Gunasekar, Jason Lee, Daniel Soudry, and Nathan Srebro · 2018
Cited alongside, same era.
Norm matters: efficient and accurate normalization schemes in deep networks
Elad Hoffer, Ron Banner, Itay Golan, and Daniel Soudry · 2018
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
A novel convergence analysis for algorithms of the adam family
Zhishuai Guo, Yi Xu, Wotao Yin, Rong Jin, and Tianbao Yang · 2021
Later among the works it cites.
What Happens after SGD Reaches Zero Loss?–A Mathematical Framework
Zhiyuan Li, Tianhao Wang, and Sanjeev Arora · 2021
Later among the works it cites.
Gradient Descent on Two-layer Nets: Margin Maximization and Simplicity Bias
Kaifeng Lyu, Zhiyuan Li, Runzhe Wang, and Sanjeev Arora · 2021
Later among the works it cites.
Rmsprop converges with proper hyperparameter
Naichen Shi and Dawei Li · 2021
Later among the works it cites.
The implicit bias for adaptive optimization algorithms on homogeneous neural networks
Bohan Wang, Qi Meng, Wei Chen, and Tie-Yan Liu · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Cited alongside, same era.
Gradient descent aligns the layers of deep linear networks
Ziwei Ji and Matus Telgarsky · 2018
Cited alongside, same era.
Decoupled Weight Decay Regularization
Ilya Loshchilov and Frank Hutter · 2018
Cited alongside, same era.
Adaptive Gradient Methods with Dynamic Bound of Learning Rate
Liangchen Luo, Yuanhao Xiong, Yan Liu, and Xu Sun · 2018
Cited alongside, same era.
On the Convergence of Adam and Beyond
Sashank J Reddi, Satyen Kale, and Sanjiv Kumar · 2018
Cited alongside, same era.
Adafactor: Adaptive learning rates with sublinear memory cost
Noam Shazeer and Mitchell Stern · 2018
Cited alongside, same era.
The Implicit Bias of Gradient Descent on Separable Data
Daniel Soudry, Elad Hoffer, Mor Shpigel Nacson, and Nathan Srebro · 2018
Cited alongside, same era.
Three Mechanisms of Weight Decay Regularization
Guodong Zhang, Chaoqi Wang, Bowen Xu, and Roger Grosse · 2018
Cited alongside, same era.
Later among the works it cites.
Adahessian: An adaptive second order optimizer for machine learning
Zhewei Yao, Amir Gholami, Sheng Shen, Mustafa Mustafa, Kurt Keutzer, and Michael Mahoney · 2021
Later among the works it cites.
Understanding gradient descent on the edge of stability in deep learning
Sanjeev Arora, Zhiyuan Li, and Abhishek Panigrahi · 2022
Later among the works it cites.
Robustness to Unbounded Smoothness of Generalized SignSGD
Michael Crawshaw, Mingrui Liu, Francesco Orabona, Wei Zhang, and Zhenxun Zhuang · 2022
Later among the works it cites.
Self-Stabilization: The Implicit Bias of Gradient Descent at the Edge of Stability
Alex Damian, Eshaan Nichani, and Jason D Lee · 2022
Later among the works it cites.
A Simple Convergence Proof of Adam and Adagrad
Alexandre Défossez, Leon Bottou, Francis Bach, and Nicolas Usunier · 2022
Later among the works it cites.
Frederik Kunstner, Jacques Chen, Jonathan Wilder Lavington, and Mark Schmidt · 2022
Later among the works it cites.
On the SDEs and Scaling Rules for Adaptive Gradient Algorithms
Sadhika Malladi, Kaifeng Lyu, Abhishek Panigrahi, and Sanjeev Arora · 2022
Later among the works it cites.
How Sharpness-Aware Minimization Minimizes Sharpness?
Kaiyue Wen, Tengyu Ma, and Zhiyuan Li · 2022
Later among the works it cites.
Adam Can Converge Without Any Modification On Update Rules
Yushun Zhang, Congliang Chen, Naichen Shi, Ruoyu Sun, and Zhi-Quan Luo · 2022
Later among the works it cites.
Understanding adamw through proximal methods and scale-freeness
Zhenxun Zhuang, Mingrui Liu, Ashok Cutkosky, and Francesco Orabona · 2022
Later among the works it cites.
The iterates of the Frank–Wolfe algorithm may not converge
Jérôme Bolte, Cyrille W Combettes, and Edouard Pauwels · 2023
Later among the works it cites.
Matias D Cattaneo, Jason M Klusowski, and Boris Shigida · 2023
Later among the works it cites.
Lion Secretly Solves a Constrained Optimization: As Lyapunov Predicts
Lizhang Chen, Bo Liu, Kaizhao Liang, et al · 2023
Later among the works it cites.
Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training
Hong Liu, Zhiyuan Li, David Leo Wright Hall, Percy Liang, and Tengyu Ma · 2023
Later among the works it cites.
Symbolic discovery of optimization algorithms
Xiangning Chen, Chen Liang, Da Huang, Esteban Real, Kaiyuan Wang, Hieu Pham, Xuanyi Dong, Thang Luong, Cho-Jui Hsieh, Yifeng Lu, et al · 2024
Closest in time.