Fetching the paper…
Reading the bibliography…
The learning rate (LR) schedule is one of the most important hyper-parameters needing careful tuning in training DNNs.
A statistical method for global optimization
Dennis D Cox and Susan John · 1992
Earlier work this paper cites.
Parameter adaptation in stochastic optimization
Luís B Almeida, Thibault Langlois, José D Amaral, and Alexander Plakhov · 1998
Earlier work this paper cites.
Classes of kernels for machine learning: a statistics perspective
Marc G Genton · 2001
Earlier work this paper cites.
Using confidence bounds for exploitation-exploration trade-offs
Peter Auer · 2002
Earlier work this paper cites.
Automatically constructing a corpus of sentential paraphrases
William B Dolan and Chris Brockett · 2005
Earlier work this paper cites.
Gaussian Processes for Machine Learning
CE. Rasmussen and CKI. Williams · 2006
Earlier work this paper cites.
Learning multiple layers of features from tiny images
A. Krizhevsky and G. Hinton · 2009
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Earlier work this paper cites.
Sequential model-based optimization for general algorithm configuration
Frank Hutter, Holger H Hoos, and Kevin Leyton-Brown · 2011
Earlier work this paper cites.
Practical bayesian optimization of machine learning algorithms
Jasper Snoek, Hugo Larochelle, and Ryan P Adams · 2012
Earlier work this paper cites.
Lecture 6.5—RmsProp: Divide the gradient by a running average of its recent magnitude
T. Tieleman and G. Hinton · 2012
Earlier work this paper cites.
Adadelta: an adaptive learning rate method
Matthew D Zeiler · 2012
Earlier work this paper cites.
Hyperopt: A python library for optimizing the hyperparameters of machine learning algorithms
James Bergstra, Dan Yamins, and David D Cox · 2013
Earlier work this paper cites.
No more pesky learning rates
Tom Schaul, Sixin Zhang, and Yann LeCun · 2013
Earlier work this paper cites.
One weird trick for parallelizing convolutional neural networks
Alex Krizhevsky · 2014
Earlier work this paper cites.
Freeze-thaw bayesian optimization
Kevin Swersky, Jasper Snoek, and Ryan Prescott Adams · 2014
Earlier work this paper cites.
Speeding up automatic hyperparameter optimization of deep neural networks by extrapolation of learning curves
Tobias Domhan, Jost Tobias Springenberg, and Frank Hutter · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederick P Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Probabilistic line searches for stochastic optimization
Maren Mahsereci and Philipp Hennig · 2015
Cited alongside, same era.
ImageNet Large Scale Visual Recognition Challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li Fei-Fei · 2015
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2015
Cited alongside, same era.
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books
Yukun Zhu, Ryan Kiros, Rich Zemel, Ruslan Salakhutdinov, Raquel Urtasun, Antonio Torralba, and Sanja Fidler · 2015
Cited alongside, same era.
Deep Learning
Ian Goodfellow, Yoshua Bengio, and Aaron Courville · 2016
Cited alongside, same era.
Non-stochastic best arm identification and hyperparameter optimization
Cyclical learning rates for training neural networks
Leslie N Smith · 2017
Later among the works it cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Later among the works it cites.
Online learning rate adaptation with hypergradient descent
Atılım Güneş Baydin, Robert Cornish, David Martínez Rubio, Mark Schmidt, and Frank Wood · 2018
Later among the works it cites.
BOHB: Robust and efficient hyperparameter optimization at scale
Stefan Falkner, Aaron Klein, and Frank Hutter · 2018
Later among the works it cites.
Mixed precision training
Paulius Micikevicius, Sharan Narang, Jonah Alben, Gregory Diamos, Erich Elsen, David Garcia, Boris Ginsburg, Michael Houston, Oleksii Kuchaiev, Ganesh Venkatesh, and Hao Wu · 2018
Later among the works it cites.
Step size matters in deep learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kevin Jamieson and Ameet Talwalkar · 2016
Cited alongside, same era.
Deep learning without poor local minima
Kenji Kawaguchi · 2016
Cited alongside, same era.
SQuAD: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang · 2016
Cited alongside, same era.
Taking the human out of the loop: A review of Bayesian optimization
Bobak Shahriari, Kevin Swersky, Ziyu Wang, Ryan P. Adams, and Nando De Freitas · 2016
Cited alongside, same era.
Yukun Zhu, Ryan Kiros, Richard Zemel, Ruslan Salakhutdinov, Raquel Urtasun, Antonio Torralba, and Sanja Fidler · 2016
Cited alongside, same era.
Forward and reverse gradient-based hyperparameter optimization
Luca Franceschi, Michele Donini, Paolo Frasconi, and Massimiliano Pontil · 2017
Cited alongside, same era.
Accurate, large minibatch SGD: Training ImageNet in 1 hour
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He · 2017
Cited alongside, same era.
Kamil Nar and Shankar Sastry · 2018
Later among the works it cites.
Leslie N Smith · 2018
Later among the works it cites.
Flipout: Efficient pseudo-independent weight perturbations on mini-batches
Yeming Wen, Paul Vicol, Jimmy Ba, Dustin Tran, and Roger Grosse · 2018
Later among the works it cites.
A broad-coverage challenge corpus for sentence understanding through inference
Adina Williams, Nikita Nangia, and Samuel Bowman · 2018
Later among the works it cites.
Understanding short-horizon bias in stochastic meta-optimization
Yuhuai Wu, Mengye Ren, Renjie Liao, and Roger Grosse · 2018
Later among the works it cites.
Towards automated deep learning: Efficient joint neural architecture and hyperparameter search
Arber Zela, Aaron Klein, Stefan Falkner, and Frank Hutter · 2018
Later among the works it cites.
Bayesian optimization meets Bayesian optimal stopping
Zhongxiang Dai, Haibin Yu, Bryan Kian Hsiang Low, and Patrick Jaillet · 2019
Later among the works it cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Later among the works it cites.
A closer look at deep learning heuristics: Learning rate restarts, warmup and distillation
Deepak Akhilesh Gotmare, Shirish Nitish Keskar, Caiming Xiong, and Richard Socher · 2019
Later among the works it cites.
Neural network acceptability judgments
Alex Warstadt, Amanpreet Singh, and Samuel R Bowman · 2019
Later among the works it cites.
MARTHE: Scheduling the learning rate via online hypergradients
Michele Donini, Luca Franceschi, Orchid Majumder, Massimiliano Pontil, and Paolo Frasconi · 2020
Later among the works it cites.
Provably efficient online hyperparameter optimization with population-based bandits
Jack Parker-Holder, Vu Nguyen, and Stephen Roberts · 2020
Later among the works it cites.