Fetching the paper…
Reading the bibliography…
Hyperparameter tuning is a bothersome step in the training of deep learning models.
Optimal brain damage
Yann LeCun, John S. Denker, and Sara A. Solla · 1990
Earlier work this paper cites.
Backpropagation through time: what it does and how to do it
Paul J Werbos · 1990
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Natural gradient works efficiently in learning
Shun-ichi Amari · 1998
Earlier work this paper cites.
Tracking the best expert
Mark Herbster and Manfred K Warmuth · 1998
Earlier work this paper cites.
Efficient backprop
Yann LeCun, Leon Bottou, Genevieve B Orr, and Klaus-Robert Müller · 1998
Earlier work this paper cites.
Switching between two universal source coding algorithms
Paul AJ Volf and Frans MJ Willems · 1998
Earlier work this paper cites.
Local gain adaptation in stochastic gradient descent
Nicol N Schraudolph · 1999
Earlier work this paper cites.
Bayesian Model Selection and Model Averaging
Larry Wasserman · 2000
Earlier work this paper cites.
Neural architecture search with reinforcement learning
Barret Zoph and Quoc V Le · 2000
Earlier work this paper cites.
Evolving neural networks through augmenting topologies
Kenneth O Stanley and Risto Miikkulainen · 2002
Earlier work this paper cites.
Combining expert advice efficiently
Wouter Koolen and Steven De Rooij · 2008
Earlier work this paper cites.
Catching up faster in Bayesian model selection and model averaging
Tim Van Erven, Steven D. Rooij, and Peter Grünwald · 2008
Earlier work this paper cites.
ImageNet: A Large-Scale Hierarchical Image Database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Learning Multiple Layers of Features from Tiny Images
Alex Krizhevsky · 2009
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Earlier work this paper cites.
Practical recommendations for gradient-based training of deep architectures
Yoshua Bengio · 2012
Earlier work this paper cites.
Tuning-free step-size adaptation
Ashique Rupam Mahmood, Richard S Sutton, Thomas Degris, and Patrick M Pilarski · 2012
Earlier work this paper cites.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
Tijmen Tieleman and Geoffrey Hinton · 2012
Earlier work this paper cites.
Catching up faster by switching sooner: A predictive approach to adaptive estimation with an application to the AIC-BIC dilemma
Tim Van Erven, Peter Grünwald, and Steven De Rooij · 2012
Earlier work this paper cites.
Making a science of model search: Hyperparameter optimization in hundreds of dimensions for vision architectures
James Bergstra, Daniel Yamins, and David Daniel Cox · 2013
Cited alongside, same era.
No more pesky learning rates
Tom Schaul, Sixin Zhang, and Yann LeCun · 2013
Cited alongside, same era.
Gradient descent: Convergence analysis, 2013
Ryan Tibshirani and Micol Marchetti-Bowick · 2013
Cited alongside, same era.
One weird trick for parallelizing convolutional neural networks
Alex Krizhevsky · 2014
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2014
Cited alongside, same era.
Compression of Neural Machine Translation Models via Pruning
Abigail See, Minh-Thang Luong, and Christopher D Manning · 2016
Later among the works it cites.
Stronger baselines for trustable results in neural machine translation
Michael Denkowski and Graham Neubig · 2017
Later among the works it cites.
Mobilenets: Efficient convolutional neural networks for mobile vision applications
Andrew G Howard, Menglong Zhu, Bo Chen, Dmitry Kalenichenko, Weijun Wang, Tobias Weyand, Marco Andreetto, and Hartwig Adam · 2017
Later among the works it cites.
Densely connected convolutional networks
Gao Huang, Zhuang Liu, Laurens Van Der Maaten, and Kilian Q Weinberger · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Rafal Jozefowicz, Wojciech Zaremba, and Ilya Sutskever · 2015
Cited alongside, same era.
Adam: A Method for Stochastic Optimization
Diederik P. Kingma and Jimmy Ba · 2015
Cited alongside, same era.
Continuous control with deep reinforcement learning
Timothy P. Lillicrap, Jonathan J. Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2015
Cited alongside, same era.
Gradient-based hyperparameter optimization through reversible learning
Dougal Maclaurin, David Duvenaud, and Ryan Adams · 2015
Cited alongside, same era.
Pierre-Yves Massé and Yann Ollivier · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Cited alongside, same era.
Going deeper with convolutions
Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich · 2015
Cited alongside, same era.
Stanislaw Jastrzebski, Zachary Kenton, Devansh Arpit, Nicolas Ballas, Asja Fischer, Yoshua Bengio, and Amos Storkey · 2017
Later among the works it cites.
Improving generalization performance by switching from Adam to SGD
Nitish Shirish Keskar and Richard Socher · 2017
Later among the works it cites.
Hyperband: A novel bandit-based approach to hyperparameter optimization
Lisha Li, Kevin Jamieson, Giulia DeSalvo, Afshin Rostamizadeh, and Ameet Talwalkar · 2017
Later among the works it cites.
Progressive neural architecture search
Chenxi Liu, Barret Zoph, Jonathon Shlens, Wei Hua, Li-Jia Li, Li Fei-Fei, Alan Yuille, Jonathan Huang, and Kevin Murphy · 2017
Later among the works it cites.
Building a large annotated corpus of english: The penn treebank
Mitchell P. Marcus, Mary Ann Marcinkiewicz, and Beatrice Santorini · 2017
Later among the works it cites.
Training deep networks without learning rates through coin betting
Francesco Orabona and Tatiana Tommasi · 2017
Later among the works it cites.
Automatic differentiation in pytorch
Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer · 2017
Later among the works it cites.
Large-scale evolution of image classifiers
Esteban Real, Sherry Moore, Andrew Selle, Saurabh Saxena, Yutaka Leon Suematsu, Jie Tan, Quoc Le, and Alex Kurakin · 2017
Later among the works it cites.
Estimating an Optimal Learning Rate For a Deep Neural Network, 2017
Pavel Surmenok · 2017
Later among the works it cites.
The marginal value of adaptive gradient methods in machine learning
Ashia C Wilson, Rebecca Roelofs, Mitchell Stern, Nati Srebro, and Benjamin Recht · 2017
Later among the works it cites.
The Lottery Ticket Hypothesis: Finding Small, Trainable Neural Networks
Jonathan Frankle and Michael Carbin · 2018
Closest in time.
pytorch-cifar, 2018
Kianglu · 2018
Closest in time.
Learning Rate Tuning in Deep Learning: A Practical Guide — Machine Learning Explained, 2018
Keita Kurita · 2018
Closest in time.
Measuring the Intrinsic Dimension of Objective Landscapes
Chunyuan Li, Heerad Farkhoor, Rosanne Liu, and Jason Yosinski · 2018
Closest in time.
Darts: Differentiable architecture search
Hanxiao Liu, Karen Simonyan, and Yiming Yang · 2018
Closest in time.