Fetching the paper…
Reading the bibliography…
Approximating Stochastic Gradient Descent (SGD) as a Stochastic Differential Equation (SDE) has allowed researchers to enjoy the benefits of studying a continuous optimization trajectory while carefully preserving the stochasticity of SGD.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Earlier work this paper cites.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude., 2012
Tijmen Tieleman and Geoffrey Hinton · 2012
Earlier work this paper cites.
One weird trick for parallelizing convolutional neural networks
Alex Krizhevsky · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books
Yukun Zhu, Ryan Kiros, Rich Zemel, Ruslan Salakhutdinov, Raquel Urtasun, Antonio Torralba, and Sanja Fidler · 2015
Earlier work this paper cites.
Incorporating Nesterov momentum into Adam
Timothy Dozat · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Accurate, large minibatch sgd: Training imagenet in 1 hour
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He · 2017
Earlier work this paper cites.
Train longer, generalize better: closing the generalization gap in large batch training of neural networks
Elad Hoffer, Itay Hubara, and Daniel Soudry · 2017
Earlier work this paper cites.
Three factors influencing minima in SGD
Stanisław Jastrzębski, Zachary Kenton, Devansh Arpit, Nicolas Ballas, Asja Fischer, Yoshua Bengio, and Amos Storkey · 2017
Earlier work this paper cites.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher · 2017
Earlier work this paper cites.
Searching for activation functions
Prajit Ramachandran, Barret Zoph, and Quoc V Le · 2017
Earlier work this paper cites.
On the convergence of Adam and beyond
Sashank J. Reddi, Satyen Kale, and Sanjiv Kumar · 2018
Cited alongside, same era.
Adaptive methods for nonconvex optimization
Manzil Zaheer, Sashank Reddi, Devendra Sachan, Satyen Kale, and Sanjiv Kumar · 2018
Cited alongside, same era.
Stochastic modified equations and dynamics of stochastic gradient algorithms I: Mathematical foundations
Qianxiao Li, Cheng Tai, and Weinan E · 2019
Cited alongside, same era.
RoBERTa: A robustly optimized BERT pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 2019
Cited alongside, same era.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2019
Cited alongside, same era.
Adaptive gradient methods with dynamic bound of learning rate
Large batch optimization for deep learning: Training BERT in 76 minutes
Yang You, Jing Li, Sashank Reddi, Jonathan Hseu, Sanjiv Kumar, Srinadh Bhojanapalli, Xiaodan Song, James Demmel, Kurt Keutzer, and Cho-Jui Hsieh · 2020
Later among the works it cites.
Why are adaptive methods good for attention models?
Jingzhao Zhang, Sai Praneeth Karimireddy, Andreas Veit, Seungyeon Kim, Sashank Reddi, Sanjiv Kumar, and Suvrit Sra · 2020
Later among the works it cites.
Towards theoretically understanding why SGD generalizes better than Adam in deep learning
Pan Zhou, Jiashi Feng, Chao Ma, Caiming Xiong, Steven Chu Hong Hoi, and Weinan E · 2020
Later among the works it cites.
AdaBelief optimizer: Adapting stepsizes by the belief in observed gradients
Juntang Zhuang, Tommy Tang, Yifan Ding, Sekhar C Tatikonda, Nicha Dvornek, Xenophon Papademetris, and James Duncan · 2020
Later among the works it cites.
Closing the generalization gap of adaptive gradient methods in training deep neural networks
Jinghui Chen, Dongruo Zhou, Yiqi Tang, Ziyan Yang, Yuan Cao, and Quanquan Gu · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Liangchen Luo, Yuanhao Xiong, and Yan Liu · 2019
Cited alongside, same era.
Quasi-hyperbolic momentum and Adam for deep learning
Jerry Ma and Denis Yarats · 2019
Cited alongside, same era.
A tail-index analysis of stochastic gradient noise in deep neural networks
Umut Simsekli, Levent Sagun, and Mert Gurbuzbalaban · 2019
Cited alongside, same era.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman · 2019
Cited alongside, same era.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Cited alongside, same era.
An exponential learning rate schedule for deep learning
Zhiyuan Li and Sanjeev Arora · 2020
Cited alongside, same era.
Reconciling modern deep learning with traditional optimization analyses: The intrinsic learning rate
Zhiyuan Li, Kaifeng Lyu, and Sanjeev Arora · 2020
Cited alongside, same era.
How to train BERT with an academic budget
Peter Izsak, Moshe Berchansky, and Omer Levy · 2021
Later among the works it cites.
Neural mechanics: Symmetry and broken conservation laws in deep learning dynamics
Daniel Kunin, Javier Sagastuy-Brena, Surya Ganguli, Daniel LK Yamins, and Hidenori Tanaka · 2021
Later among the works it cites.
On the validity of modeling sgd with stochastic differential equations (sdes)
Zhiyuan Li, Sadhika Malladi, and Sanjeev Arora · 2021
Later among the works it cites.
A diffusion theory for deep learning dynamics: Stochastic gradient descent exponentially favors flat minima
Zeke Xie, Issei Sato, and Masashi Sugiyama · 2021
Later among the works it cites.
Learning rates as a function of batch size: A random matrix theory approach to neural network training
Diego Granziol, Stefan Zohren, and Stephen Roberts · 2022
Closest in time.
A qualitative study of the dynamic behavior for adaptive gradient algorithms
Chao Ma, Lei Wu, and Weinan E · 2022
Closest in time.
Should you mask 15% in masked language modeling?
Alexander Wettig, Tianyu Gao, Zexuan Zhong, and Danqi Chen · 2022
Closest in time.
Adaptive inertia: Disentangling the effects of adaptive learning rate and momentum
Zeke Xie, Xinrui Wang, Huishuai Zhang, Issei Sato, and Masashi Sugiyama · 2022
Closest in time.