Fetching the paper…
Reading the bibliography…
The learning rate warmup heuristic achieves remarkable success in stabilizing training, accelerating convergence and improving generalization for adaptive stochastic optimization algorithms like RMSprop and Adam.
Taylor series methods
Kirk M Wolter · 2007
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2010
Earlier work this paper cites.
Neural networks for machine learning lecture 6a overview of mini-batch gradient descent
Geoffrey Hinton, Nitish Srivastava, and Kevin Swersky · 2012
Earlier work this paper cites.
Adadelta: an adaptive learning rate method
Matthew D Zeiler · 2012
Earlier work this paper cites.
Advances in optimizing recurrent networks
Yoshua Bengio, Nicolas Boulanger-Lewandowski, and Razvan Pascanu · 2013
Earlier work this paper cites.
One billion word benchmark for measuring progress in statistical language modeling
Ciprian Chelba, Tomas Mikolov, Michael Schuster, Qi Ge, Thorsten Brants, Phillipp Koehn, and Tony Robinson · 2013
Earlier work this paper cites.
Report on the 11th iwslt evaluation campaign, iwslt 2014
Mauro Cettolo, Jan Niehues, Sebastian Stüker, Luisa Bentivogli, and Marcello Federico · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Forecasting with moving averages
Robert Nau · 2014
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Cited alongside, same era.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton · 2016
Cited alongside, same era.
Incorporating nesterov momentum into adam
Timothy Dozat · 2016
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Rethinking the inception architecture for computer vision
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna · 2016
Cited alongside, same era.
The shattered gradients problem: If resnets are the answer, then what is the question?
Efficient contextualized representation: Language model pruning for sequence labeling
Liyuan Liu, Xiang Ren, Jingbo Shang, Jian Peng, and Jiawei Han · 2018
Later among the works it cites.
Fixing weight decay regularization in adam
Ilya Loshchilov and Frank Hutter · 2018
Later among the works it cites.
Training tips for the transformer model
Martin Popel and Ondřej Bojar · 2018
Later among the works it cites.
On the convergence of adam and beyond
Sashank J Reddi, Satyen Kale, and Sanjiv Kumar · 2018
Later among the works it cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Closest in time.
A closer look at deep learning heuristics: Learning rate restarts, warmup and distillation
Akhilesh Gotmare, Nitish Shirish Keskar, Caiming Xiong, and Richard Socher · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
David Balduzzi, Marcus Frean, Lennox Leary, JP Lewis, Kurt Wan-Duo Ma, and Brian McWilliams · 2017
Cited alongside, same era.
Accurate, large minibatch sgd: Training imagenet in 1 hour
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Dscovr: Randomized primal-dual block coordinate algorithms for asynchronous distributed optimization
Lin Xiao, Adams Wei Yu, Qihang Lin, and Weizhu Chen · 2017
Cited alongside, same era.
signsgd: Compressed optimisation for non-convex problems
Jeremy Bernstein, Yu-Xiang Wang, Kamyar Azizzadenesheli, and Anima Anandkumar · 2018
Cited alongside, same era.
Closing the generalization gap of adaptive gradient methods in training deep neural networks
Jinghui Chen, Dongruo Zhou, Yiqi Tang, Ziyan Yang, and Quanquan Gu · 2018
Cited alongside, same era.
Méthode générale pour la résolution des systemes d’équations simultanées
Augustin Cauchy
Cited in the paper.
Closest in time.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 2019
Closest in time.
Adaptive gradient methods with dynamic bound of learning rate
Liangchen Luo, Yuanhao Xiong, Yan Liu, and Xu Sun · 2019
Closest in time.
fairseq: A fast, extensible toolkit for sequence modeling
Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli · 2019
Closest in time.
Fixup initialization: Residual learning without normalization
Hongyi Zhang, Yann N Dauphin, and Tengyu Ma · 2019
Closest in time.