Fetching the paper…
Reading the bibliography…
Complex learning rate schedules have become an integral part of deep learning.
Online learning and stochastic approximations, 1998
Léon Bottou · 1998
Earlier work this paper cites.
Gradient-based hyperparameter optimization through reversible learning, 2015
Dougal Maclaurin, David Duvenaud, and Ryan P. Adams · 2015
Earlier work this paper cites.
Learning to optimize, 2016
Ke Li and Jitendra Malik · 2016
Earlier work this paper cites.
Meta-sgd: Learning to learn quickly for few-shot learning, 2017
Zhenguo Li, Fengwei Zhou, Fei Chen, and Hang Li · 2017
Earlier work this paper cites.
L2 regularization versus batch and weight normalization, 2017
Twan van Laarhoven · 2017
Earlier work this paper cites.
Attention is all you need, 2017
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Learned optimizers that scale and generalize, 2017
Olga Wichrowska, Niru Maheswaranathan, Matthew W. Hoffman, Sergio Gomez Colmenarejo, Misha Denil, Nando de Freitas, and Jascha Sohl-Dickstein · 2017
Earlier work this paper cites.
JAX: composable transformations of Python+NumPy programs
James Bradbury, Roy Frostig, Peter Hawkins, Matthew James Johnson, Chris Leary, Dougal Maclaurin, and Skye Wanderman-Milne · 2018
Earlier work this paper cites.
L4: Practical loss-based stepsize adaptation for deep learning, 2018
Michal Rolinek and Georg Martius · 2018
Earlier work this paper cites.
Fluctuation-dissipation relations for stochastic gradient descent, 2018
Sho Yaida · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding, 2019
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Cited alongside, same era.
An exponential learning rate schedule for deep learning, 2019
Zhiyuan Li and Sanjeev Arora · 2019
Cited alongside, same era.
Measuring the effects of data parallelism on neural network training, 2019
Christopher J. Shallue, Jaehoon Lee, Joseph Antognini, Jascha Sohl-Dickstein, Roy Frostig, and George E. Dahl · 2019
Cited alongside, same era.
Disentangling adaptive gradient methods from learning rates, 2020
Naman Agarwal, Rohan Anil, Elad Hazan, Tomer Koren, and Cyril Zhang · 2020
Cited alongside, same era.
Language models are few-shot learners, 2020
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Neural mechanics: Symmetry and broken conservation laws in deep learning dynamics, 2020
Daniel Kunin, Javier Sagastuy-Brena, Surya Ganguli, Daniel L. K. Yamins, and Hidenori Tanaka · 2020
Later among the works it cites.
On the training dynamics of deep networks with l 2 l_{2} regularization, 2020
Aitor Lewkowycz and Guy Gur-Ari · 2020
Later among the works it cites.
The large learning rate phase of deep learning: the catapult mechanism, 2020
Aitor Lewkowycz, Yasaman Bahri, Ethan Dyer, Jascha Sohl-Dickstein, and Guy Gur-Ari · 2020
Later among the works it cites.
Gradient descent maximizes the margin of homogeneous neural networks, 2020
Kaifeng Lyu and Jian Li · 2020
Later among the works it cites.
Learning rate annealing can provably help generalization, even for convex problems, 2020
Preetum Nakkiran · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Sharpness-aware minimization for efficiently improving generalization, 2020
Pierre Foret, Ariel Kleiner, Hossein Mobahi, and Behnam Neyshabur · 2020
Cited alongside, same era.
Scaling laws for neural language models, 2020
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2020
Cited alongside, same era.
Towards explaining the regularization effect of initial large learning rate in training neural networks, 2020a
Yuanzhi Li, Colin Wei, and Tengyu Ma
Cited in the paper.
Reconciling modern deep learning with traditional optimization analyses: The intrinsic learning rate, 2020b
Zhiyuan Li, Kaifeng Lyu, and Sanjeev Arora
Cited in the paper.
Xiaoman Qi, PanPan Zhu, Yuebin Wang, Liqiang Zhang, Junhuan Peng, Mengfan Wu, Jialong Chen, Xudong Zhao, Ning Zang, and P. Takis Mathiopoulos · 2020
Later among the works it cites.
Efficientnet: Rethinking model scaling for convolutional neural networks, 2020
Mingxing Tan and Quoc V. Le · 2020
Later among the works it cites.
Spherical motion dynamics: Learning dynamics of neural network with normalization, weight decay, and sgd, 2020
Ruosi Wan, Zhanxing Zhu, Xiangyu Zhang, and Jian Sun · 2020
Later among the works it cites.