Fetching the paper…
Reading the bibliography…
As optimizers are critical to the performances of neural networks, every year a large number of papers innovating on the subject are published.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Cifar-10 (canadian institute for advanced research)
Alex Krizhevsky, Vinod Nair, and Geoffrey Hinton · 2010
Earlier work this paper cites.
The mnist database of handwritten images, 2012
Y LeCun, C Cortes, and CJC Burgess · 2012
Earlier work this paper cites.
Combinatorial bandits
Nicolo Cesa-Bianchi and Gábor Lugosi · 2012
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Rethinking the inception architecture for computer vision
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna · 2016
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2017
Earlier work this paper cites.
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Ranger - a synergistic optimizer
Less Wright · 2019
Cited alongside, same era.
On the variance of the adaptive learning rate and beyond
Liyuan Liu, Haoming Jiang, Pengcheng He, Weizhu Chen, Xiaodong Liu, Jianfeng Gao, and Jiawei Han · 2019
Cited alongside, same era.
Lookahead optimizer: k steps forward, 1 step back
Michael Zhang, James Lucas, Jimmy Ba, and Geoffrey E Hinton · 2019
Cited alongside, same era.
Why gradient clipping accelerates training: A theoretical justification for adaptivity
Jingzhao Zhang, Tianxing He, Suvrit Sra, and Ali Jadbabaie · 2019
Cited alongside, same era.
Adaptive learning rate clipping stabilizes learning
Jeffrey M Ede and Richard Beanland · 2020
Later among the works it cites.
Gradient centralization: A new optimization technique for deep neural networks
Hongwei Yong, Jianqiang Huang, Xiansheng Hua, and Lei Zhang · 2020
Later among the works it cites.
Stable weight decay regularization
Zeke Xie, Issei Sato, and Masashi Sugiyama · 2020
Later among the works it cites.
Wide-minima density hypothesis and the explore-exploit learning rate schedule
Nikhil Iyer, V Thejas, Nipun Kwatra, Ramachandran Ramjee, and Muthian Sivathanu · 2020
Later among the works it cites.
Optax: composable gradient transformation and optimisation, in jax!
Matteo Hessel, David Budden, Fabio Viola, Mihaela Rosca, Eren Sezener, and Tom Hennigan · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
On the adequacy of untuned warmup for adaptive optimization
Jerry Ma and Denis Yarats · 2019
Cited alongside, same era.
How we beat the fastai leaderboard score by +19.77%… a synergy of new deep learning techniques for your consideration
Less Wright · 2019
Cited alongside, same era.
Descending through a crowded valley–benchmarking deep learning optimizers
Robin M Schmidt, Frank Schneider, and Philipp Hennig · 2020
Cited alongside, same era.
Andrew Brock, Soham De, Samuel L Smith, and Karen Simonyan · 2021
Closest in time.
Positive-negative momentum: Manipulating stochastic gradient noise to improve generalization
Zeke Xie, li Yuan, Zhanxing Zhu, and Masashi Sugiyama · 2021
Closest in time.
Norm loss: An efficient yet effective regularization method for deep neural networks
Theodoros Georgiou, Sebastian Schmitt, Thomas Bäck, Wei Chen, and Michael Lew · 2021
Closest in time.