Fetching the paper…
Reading the bibliography…
Optimizers like Adam and AdaGrad have been very successful in training large-scale neural networks.
On information and sufficiency
Solomon Kullback and Richard A Leibler · 1951
Earlier work this paper cites.
A method for solving the convex programming problem with convergence rate of ( 1 / k 2 1/k^{2} )
Yurii E. Nesterov · 1983
Earlier work this paper cites.
Increased rates of convergence through learning rate adaptation
Robert A Jacobs · 1988
Earlier work this paper cites.
Learning quickly when irrelevant attributes abound: A new linear-threshold algorithm
Nick Littlestone · 1988
Earlier work this paper cites.
Back-propagation heuristics: a study of the extended delta-bar-delta algorithm
A. A. Minai and R. D. Williams · 1990
Earlier work this paper cites.
A direct adaptive method for faster backpropagation learning: The rprop algorithm
Martin Riedmiller and Heinrich Braun · 1993
Earlier work this paper cites.
Adapting bias by gradient descent: An incremental version of delta-bar-delta
Richard Sutton · 1995
Earlier work this paper cites.
A decision-theoretic generalization of on-line learning and an application to boosting
Yoav Freund and Robert E Schapire · 1997
Earlier work this paper cites.
Exponentiated gradient versus gradient descent for linear predictors
Jyrki Kivinen and Manfred K. Warmuth · 1997
Earlier work this paper cites.
The perceptron algorithm versus winnow: linear versus logarithmic mistake bounds when few input variables are relevant
Jyrki Kivinen, Manfred K Warmuth, and Peter Auer · 1997
Cited alongside, same era.
Tracking the best expert
Mark Herbster and Manfred K. Warmuth · 1998
Cited alongside, same era.
Local gain adaptation in stochastic gradient descent
Nicol N Schraudolph · 1999
Cited alongside, same era.
Batch and on-line parameter estimation of Gaussian mixtures based on the joint entropy
Yoram Singer and Manfred K Warmuth · 1999
Cited alongside, same era.
Tracking a small set of experts by mixing past posteriors
Olivier Bousquet and Manfred K. Warmuth · 2002
Cited alongside, same era.
The p-norm generalization of the LMS algorithm for adaptive filtering
Jyrki Kivinen, Manfred K Warmuth, and Babak Hassibi · 2006
Randomized online PCA algorithms with regret bounds that are logarithmic in the dimension
Manfred K Warmuth and Dima Kuzmin · 2008
Later among the works it cites.
Mnist handwritten digit database
Yann LeCun, Corinna Cortes, and CJ Burges · 2010
Later among the works it cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Later among the works it cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Later among the works it cites.
Online learning rate adaptation with hypergradient descent
Atilim Gunes Baydin, Robert Cornish, David Martínez-Rubio, Mark Schmidt, and Frank Wood · 2018
Later among the works it cites.
An implicit form of Krasulina’s k-PCA update without the orthonormality constraint
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Winnowing subspaces
Manfred K. Warmuth · 2007
Cited alongside, same era.
Explicit update vs implicit update
Wenwu He and Hui Jiang · 2008
Cited alongside, same era.
Neural networks for machine learning lecture 6a overview of mini-batch gradient descent
Geoffrey Hinton, Nitish Srivastava, and Kevin Swersky
Cited in the paper.
Ehsan Amid and Manfred K Warmuth · 2020
Later among the works it cites.
Winnowing with gradient descent
Ehsan Amid and Manfred K. Warmuth · 2020
Later among the works it cites.
First-order preconditioning via hypergradient descent
Ted Moskovitz, Rui Wang, Janice Lan, Sanyam Kapoor, Thomas Miconi, Jason Yosinski, and Aditya Rawal · 2020
Later among the works it cites.