Fetching the paper…
Reading the bibliography…
Hyperparameter tuning is one of the most time-consuming workloads in deep learning.
Some methods of speeding up the convergence of iteration methods
Boris T Polyak · 1964
Earlier work this paper cites.
A method of solving a convex programming problem with convergence rate o (1/k2)
Yurii Nesterov · 1983
Earlier work this paper cites.
Accelerated training of backpropagation networks by using adaptive momentum step
G Qiu, MR Varley, and TJ Terrell · 1992
Earlier work this paper cites.
Building a large annotated corpus of english: The penn treebank
Mitchell P Marcus, Mary Ann Marcinkiewicz, and Beatrice Santorini · 1993
Earlier work this paper cites.
Simple adaptive momentum: new algorithm for training multilayer perceptrons
DJ Swanston, JM Bishop, and Richard James Mitchell · 1994
Earlier work this paper cites.
Optimal stochastic search and adaptive momentum
Todd K Leen and Genevieve B Orr · 1994
Earlier work this paper cites.
Levenberg-marquardt algorithm with adaptive momentum for the efficient training of feedforward networks
Nikolaos Ampazis and Stavros J Perantonis · 2000
Earlier work this paper cites.
Stable adaptive momentum for rapid online learning in nonlinear systems
Thore Graepel and Nicol N Schraudolph · 2002
Earlier work this paper cites.
Neural networks: tricks of the trade
Genevieve B Orr and Klaus-Robert Müller · 2003
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Earlier work this paper cites.
Hogwild: A lock-free approach to parallelizing stochastic gradient descent
Feng Niu, Benjamin Recht, Christopher Re, and Stephen Wright · 2011
Earlier work this paper cites.
Sequential model-based optimization for general algorithm configuration
Frank Hutter, Holger H Hoos, and Kevin Leyton-Brown · 2011
Earlier work this paper cites.
The effect of adaptive momentum in improving the accuracy of gradient descent back propagation algorithm on classification problems
Mohammad Zubair Rehman and Nazri Mohd Nawi · 2011
Earlier work this paper cites.
Practical recommendations for gradient-based training of deep architectures
Yoshua Bengio · 2012
Earlier work this paper cites.
Deep learning of representations for unsupervised and transfer learning
Yoshua Bengio et al · 2012
Earlier work this paper cites.
Stochastic gradient descent tricks
Léon Bottou · 2012
Cited alongside, same era.
Random search for hyper-parameter optimization
James Bergstra and Yoshua Bengio · 2012
Cited alongside, same era.
Practical bayesian optimization of machine learning algorithms
Jasper Snoek, Hugo Larochelle, and Ryan P Adams · 2012
Cited alongside, same era.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
Tijmen Tieleman and Geoffrey Hinton · 2012
Cited alongside, same era.
University Lecture, 2012
Simon Foucart · 2012
Cited alongside, same era.
Estimating the hessian by back-propagating curvature
James Martens, Ilya Sutskever, and Kevin Swersky · 2012
Cited alongside, same era.
Visualizing and understanding recurrent networks
Andrej Karpathy, Justin Johnson, and Li Fei-Fei · 2015
Later among the works it cites.
A thorough examination of the cnn/daily mail reading comprehension task
Danqi Chen, Jason Bolton, and Christopher D Manning · 2016
Later among the works it cites.
Asynchrony begets momentum, with an application to deep learning
Ioannis Mitliagkas, Ce Zhang, Stefan Hadjis, and Christopher Ré · 2016
Later among the works it cites.
Analysis and design of optimization algorithms via integral quadratic constraints
Laurent Lessard, Benjamin Recht, and Andrew Packard · 2016
Later among the works it cites.
Fisher information., 2016
John Duchi · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
On the importance of initialization and momentum in deep learning
Ilya Sutskever, James Martens, George Dahl, and Geoffrey Hinton · 2013
Cited alongside, same era.
On the difficulty of training recurrent neural networks
Razvan Pascanu, Tomas Mikolov, and Yoshua Bengio · 2013
Cited alongside, same era.
Introductory lectures on convex optimization: A basic course , volume 87
Yurii Nesterov · 2013
Cited alongside, same era.
Revisiting natural gradient for deep networks
Razvan Pascanu and Yoshua Bengio · 2013
Cited alongside, same era.
No more pesky learning rates
Tom Schaul, Sixin Zhang, and Yann LeCun · 2013
Cited alongside, same era.
Intriguing properties of neural networks
Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian Goodfellow, and Rob Fergus · 2013
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Later among the works it cites.
Back-propagation algorithm with variable adaptive momentum
Alaa Ali Hameed, Bekir Karlik, and Mohammad Shukri Salman · 2016
Later among the works it cites.
Chenzhuo Zhu, Song Han, Huizi Mao, and William J Dally · 2016
Later among the works it cites.
Deep Learning
Ian Goodfellow, Yoshua Bengio, and Aaron Courville · 2016
Later among the works it cites.
Omnivore: An optimizer for multi-device deep learning on cpus and gpus
Stefan Hadjis, Ce Zhang, Ioannis Mitliagkas, Dan Iter, and Christopher Ré · 2016
Later among the works it cites.
Aggregated residual transformations for deep neural networks
Saining Xie, Ross Girshick, Piotr Dollár, Zhuowen Tu, and Kaiming He · 2016
Later among the works it cites.
Using the output embedding to improve language models
Ofir Press and Lior Wolf · 2016
Later among the works it cites.
Convolutional sequence to sequence learning
Jonas Gehring, Michael Auli, David Grangier, Denis Yarats, and Yann N Dauphin · 2017
Closest in time.
The marginal value of adaptive gradient methods in machine learning
Ashia C Wilson, Rebecca Roelofs, Mitchell Stern, Nathan Srebro, and Benjamin Recht · 2017
Closest in time.