Fetching the paper…
Reading the bibliography…
Standard first-order stochastic optimization algorithms base their updates solely on the average mini-batch gradient, and it has been shown that tracking additional quantities such as the curvature can help de-sensitize common hyperparameters.
A new approach to linear filtering and prediction problems
Rudolph Emil Kalman · 1960
Earlier work this paper cites.
A new method of locating the maximum point of an arbitrary multipeak curve in the presence of noise
Harold J Kushner · 1964
Earlier work this paper cites.
Some methods of speeding up the convergence of iteration methods
Boris T Polyak · 1964
Earlier work this paper cites.
Historical development of the newton–raphson method
Tjalling J Ypma · 1995
Earlier work this paper cites.
Natural gradient works efficiently in learning
Shun-Ichi Amari · 1998
Earlier work this paper cites.
Efficient global optimization of expensive black-box functions
Donald R Jones, Matthias Schonlau, and William J Welch · 1998
Earlier work this paper cites.
Parameter adaptation in stochastic optimization
Luís B Almeida, Thibault Langlois, José D Amaral, and Alexander Plakhov · 1999
Earlier work this paper cites.
Local gain adaptation in stochastic gradient descent
Nicol N Schraudolph · 1999
Earlier work this paper cites.
Learning rate adaptation in stochastic gradient descent
VP Plagianakos, GD Magoulas, and MN Vrahatis · 2001
Earlier work this paper cites.
Kalman filtering in stochastic gradient algorithms: construction of a stopping rule
Barbara Bittner and Luc Pronzato · 2004
Earlier work this paper cites.
Fast online policy gradient learning with smd gain vector adaptation
Jin Yu, Douglas Aberdeen, and Nicol N Schraudolph · 2006
Earlier work this paper cites.
Deep learning via Hessian-free optimization
James Martens · 2010
Earlier work this paper cites.
Rectified linear units improve restricted boltzmann machines
Vinod Nair and Geoffrey E Hinton · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Earlier work this paper cites.
Parameter tuning for configuring and analyzing evolutionary algorithms
Agoston E Eiben and Selmar K Smit · 2011
Cited alongside, same era.
Adadelta: an adaptive learning rate method
Matthew D Zeiler · 2012
Cited alongside, same era.
Accelerating stochastic gradient descent using predictive variance reduction
Rie Johnson and Tong Zhang · 2013
Cited alongside, same era.
Bayesian filtering and smoothing , volume 3
Simo Särkkä · 2013
Cited alongside, same era.
A framework for self-tuning optimization algorithm
Xin-She Yang, Suash Deb, Martin Loomes, and Mehmet Karamanoglu · 2013
Cited alongside, same era.
Optimization methods for large-scale machine learning
Léon Bottou, Frank E Curtis, and Jorge Nocedal · 2018
Later among the works it cites.
JAX: composable transformations of Python+NumPy programs, 2018
James Bradbury, Roy Frostig, Peter Hawkins, Matthew James Johnson, Chris Leary, Dougal Maclaurin, and Skye Wanderman-Milne · 2018
Later among the works it cites.
Probabilistic Approaches to Stochastic Optimization
Maren Mahsereci · 2018
Later among the works it cites.
L4: Practical loss-based stepsize adaptation for deep learning
Michal Rolinek and Georg Martius · 2018
Later among the works it cites.
Kalman gradient descent: Adaptive variance reduction in stochastic optimization, 2018
James Vuckovic · 2018
Later among the works it cites.
Understanding short-horizon bias in stochastic meta-optimization
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Diederik P Kingma and Jimmy Ba · 2014
Cited alongside, same era.
New insights and perspectives on the natural gradient method
James Martens · 2014
Cited alongside, same era.
Kalman-based stochastic gradient method with stop condition and insensitivity to conditioning
Vivak Patel · 2016
Cited alongside, same era.
Taking the human out of the loop: A review of bayesian optimization
B. Shahriari, K. Swersky, Z. Wang, R. P. Adams, and N. de Freitas · 2016
Cited alongside, same era.
Online learning rate adaptation with hypergradient descent
Atilim Gunes Baydin, Robert Cornish, David Martinez Rubio, Mark Schmidt, and Frank Wood · 2017
Cited alongside, same era.
Tracking the gradients using the hessian: A new look at variance reducing stochastic methods
Robert M Gower, Nicolas Le Roux, and Francis Bach · 2017
Cited alongside, same era.
Probabilistic line searches for stochastic optimization
Maren Mahsereci and Philipp Hennig · 2017
Cited alongside, same era.
Yuhuai Wu, Mengye Ren, Renjie Liao, and Roger Grosse · 2018
Later among the works it cites.
Group normalization
Yuxin Wu and Kaiming He · 2018
Later among the works it cites.
Auto-vectorizing TensorFlow graphs: Jacobians, auto-batching and beyond
Ashish Agarwal and Igor Ganichev · 2019
Later among the works it cites.
Training neural networks for and by interpolation
Leonard Berrada, Andrew Zisserman, and M Pawan Kumar · 2019
Later among the works it cites.
On empirical comparisons of optimizers for deep learning
Dami Choi, Christopher J Shallue, Zachary Nado, Jaehoon Lee, Chris J Maddison, and George E Dahl · 2019
Later among the works it cites.
Deepobs: A deep learning optimizer benchmark suite
Frank Schneider, Lukas Balles, and Philipp Hennig · 2019
Later among the works it cites.
Painless stochastic gradient: Interpolation, line-search, and convergence rates
Sharan Vaswani, Aaron Mishkin, Issam Laradji, Mark Schmidt, Gauthier Gidel, and Simon Lacoste-Julien · 2019
Later among the works it cites.
BackPACK: Packing more into backprop
Felix Dangel, Frederik Kunstner, and Philipp Hennig · 2020
Closest in time.