Fetching the paper…
Reading the bibliography…
Although deep learning has produced dazzling successes for applications of image, speech, and video processing in the past few years, most trainings are with suboptimal hyper-parameters, requiring unnecessarily long training times.
Simulated annealing and boltzmann machines
Emile Aarts and Jan Korst · 1988
Earlier work this paper cites.
Neural networks: tricks of the trade
Genevieve B Orr and Klaus-Robert Müller · 2003
Earlier work this paper cites.
The general inefficiency of batch training for gradient descent learning
D Randall Wilson and Tony R Martinez · 2003
Earlier work this paper cites.
Curriculum learning
Yoshua Bengio, Jerome Louradour, Ronan Collobert, and Jason Weston · 2009
Earlier work this paper cites.
Practical recommendations for gradient-based training of deep architectures
Yoshua Bengio · 2012
Earlier work this paper cites.
Random search for hyper-parameter optimization
James Bergstra and Yoshua Bengio · 2012
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Dropout: A simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Cited alongside, same era.
No more pesky learning rate guessing games
Leslie N Smith · 2015
Cited alongside, same era.
Deep learning , volume 1
Ian Goodfellow, Yoshua Bengio, Aaron Courville, and Yoshua Bengio · 2016
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Regularization for deep learning: A taxonomy
Jan Kukačka, Vladimir Golkov, and Daniel Cremers · 2017
Cited alongside, same era.
Cyclical learning rates for training neural networks
Understanding generalization and stochastic gradient descent
Samuel L Smith and Quoc V Le · 2017
Later among the works it cites.
Don’t decay the learning rate, increase the batch size
Samuel L Smith, Pieter-Jan Kindermans, and Quoc V Le · 2017
Later among the works it cites.
Inception-v4, inception-resnet and the impact of residual connections on learning
Christian Szegedy, Sergey Ioffe, Vincent Vanhoucke, and Alexander A Alemi · 2017
Later among the works it cites.
Do deep nets really need weight decay and dropout?
Alex Hernández-García and Peter König · 2018
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Leslie N Smith · 2017
Cited alongside, same era.
Super-convergence: Very fast training of residual networks using large learning rates
Leslie N Smith and Nicholay Topin · 2017
Cited alongside, same era.
Residual connections encourage iterative inference
Stanisław Jastrzebski, Devansh Arpit, Nicolas Ballas, Vikas Verma, Tong Che, and Yoshua Bengio
Cited in the paper.
Three factors influencing minima in sgd
Stanisław Jastrzebski, Zachary Kenton, Devansh Arpit, Nicolas Ballas, Asja Fischer, Yoshua Bengio, and Amos Storkey
Cited in the paper.
Tianyi Liu, Zhehui Chen, Enlu Zhou, and Tuo Zhao · 2018
Closest in time.
Stochastic hyperparameter optimization through hypernetworks
Jonathan Lorraine and David Duvenaud · 2018
Closest in time.
Chen Xing, Devansh Arpit, Christos Tsirigotis, and Yoshua Bengio · 2018
Closest in time.