Fetching the paper…
Reading the bibliography…
Learned optimizers are algorithms that can themselves be trained to solve optimization problems.
A stochastic approximation method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
An automatic method for finding the greatest or least value of a function
HoHo Rosenbrock · 1960
Earlier work this paper cites.
Some methods of speeding up the convergence of iteration methods
Boris T Polyak · 1964
Earlier work this paper cites.
Numerical optimization
Jorge Nocedal and Stephen Wright · 2006
Earlier work this paper cites.
Matplotlib: A 2d graphics environment
John D Hunter · 2007
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Earlier work this paper cites.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
Tijmen Tieleman and Geoffrey Hinton · 2012
Earlier work this paper cites.
Opening the black box: low-dimensional dynamics in high-dimensional recurrent neural networks
David Sussillo and Omri Barak · 2013
Earlier work this paper cites.
On the difficulty of training recurrent neural networks
Razvan Pascanu, Tomas Mikolov, and Yoshua Bengio · 2013
Earlier work this paper cites.
No more pesky learning rates
Tom Schaul, Sixin Zhang, and Yann LeCun · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
A differential equation for modeling nesterov’s accelerated gradient method: Theory and insights
Weijie Su, Stephen Boyd, and Emmanuel Candes · 2014
Earlier work this paper cites.
Learning phrase representations using rnn encoder-decoder for statistical machine translation
Kyunghyun Cho, Bart Van Merriënboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio · 2014
Earlier work this paper cites.
Optimizing neural networks with kronecker-factored approximate curvature
James Martens and Roger Grosse · 2015
Earlier work this paper cites.
Learning to learn by gradient descent by gradient descent
Marcin Andrychowicz, Misha Denil, Sergio Gomez, Matthew W Hoffman, David Pfau, Tom Schaul, Brendan Shillingford, and Nando De Freitas · 2016
Earlier work this paper cites.
Ke Li and Jitendra Malik · 2016
Earlier work this paper cites.
A lyapunov analysis of momentum methods in optimization
Ashia C Wilson, Benjamin Recht, and Michael I Jordan · 2016
Cited alongside, same era.
Sgdr: Stochastic gradient descent with warm restarts
Ilya Loshchilov and Frank Hutter · 2016
Cited alongside, same era.
Discovering governing equations from data by sparse identification of nonlinear dynamical systems
Steven L Brunton, Joshua L Proctor, and J Nathan Kutz · 2016
Cited alongside, same era.
Learned optimizers that scale and generalize
Olga Wichrowska, Niru Maheswaranathan, Matthew W Hoffman, Sergio Gomez Colmenarejo, Misha Denil, Nando de Freitas, and Jascha Sohl-Dickstein · 2017
Cited alongside, same era.
Learning gradient descent: Better generalization and longer horizons
On the convergence of adam and beyond
Sashank J Reddi, Satyen Kale, and Sanjiv Kumar · 2019
Later among the works it cites.
Acceleration via symplectic discretization of high-resolution differential equations
Bin Shi, Simon S Du, Weijie Su, and Michael I Jordan · 2019
Later among the works it cites.
Gated recurrent units viewed through the lens of continuous time dynamical systems
Ian D Jordan, Piotr Aleksander Sokol, and Il Memming Park · 2019
Later among the works it cites.
The step decay schedule: A near optimal, geometrically decaying learning rate procedure for least squares
Rong Ge, Sham M Kakade, Rahul Kidambi, and Praneeth Netrapalli · 2019
Later among the works it cites.
On empirical comparisons of optimizers for deep learning
Dami Choi, Christopher J Shallue, Zachary Nado, Jaehoon Lee, Chris J Maddison, and George E Dahl · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kaifeng Lv, Shunhua Jiang, and Jian Li · 2017
Cited alongside, same era.
Neural optimizer search with reinforcement learning
Irwan Bello, Barret Zoph, Vijay Vasudevan, and Quoc V Le · 2017
Cited alongside, same era.
Why momentum really works
Gabriel Goh · 2017
Cited alongside, same era.
Don’t decay the learning rate, increase the batch size
Samuel L Smith, Pieter-Jan Kindermans, Chris Ying, and Quoc V Le · 2017
Cited alongside, same era.
Accurate, large minibatch sgd: Training imagenet in 1 hour
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He · 2017
Cited alongside, same era.
Cyclical learning rates for training neural networks
Leslie N Smith · 2017
Cited alongside, same era.
Fixing weight decay regularization in adam
Ilya Loshchilov and Frank Hutter · 2018
Cited alongside, same era.
Nonlinear dynamics and chaos with student solutions manual: With applications to physics, biology, chemistry, and engineering
Steven H Strogatz · 2018
Cited alongside, same era.
Later among the works it cites.
An exponential learning rate schedule for deep learning
Zhiyuan Li and Sanjeev Arora · 2019
Later among the works it cites.
Data-driven discovery of coordinates and governing equations
Kathleen Champion, Bethany Lusch, J Nathan Kutz, and Steven L Brunton · 2019
Later among the works it cites.
Tasks, stability, architecture, and compute: Training more effective learned optimizers, and using them to train themselves, 2020
Luke Metz, Niru Maheswaranathan, C. Daniel Freeman, Ben Poole, and Jascha Sohl-Dickstein · 2020
Closest in time.
Why gradient clipping accelerates training: A theoretical justification for adaptivity
Jingzhao Zhang, Tianxing He, Suvrit Sra, and Ali Jadbabaie · 2020
Closest in time.
How recurrent networks implement contextual processing in sentiment analysis
Niru Maheswaranathan and David Sussillo · 2020
Closest in time.
Reverse-engineering recurrent neural network solutions to a hierarchical inference task for mice
Rylan Schaeffer, Mikail Khona, Leenoy Meshulam, Ila Rani Fiete, et al · 2020
Closest in time.
Theory of gating in recurrent neural networks
Kamesh Krishnamurthy, Tankut Can, and David J Schwab · 2020
Closest in time.
Gating creates slow modes and controls phase-space complexity in grus and lstms
Tankut Can, Kamesh Krishnamurthy, and David J Schwab · 2020
Closest in time.
Discovering symbolic models from deep learning with inductive biases
Miles Cranmer, Alvaro Sanchez-Gonzalez, Peter Battaglia, Rui Xu, Kyle Cranmer, David Spergel, and Shirley Ho · 2020
Closest in time.
Array programming with numpy
Charles R Harris, K Jarrod Millman, Stéfan J van der Walt, Ralf Gommers, Pauli Virtanen, David Cournapeau, Eric Wieser, Julian Taylor, Sebastian Berg, Nathaniel J Smith, et al · 2020
Closest in time.
Scipy 1.0: fundamental algorithms for scientific computing in python
Pauli Virtanen, Ralf Gommers, Travis E Oliphant, Matt Haberland, Tyler Reddy, David Cournapeau, Evgeni Burovski, Pearu Peterson, Warren Weckesser, Jonathan Bright, et al · 2020
Closest in time.