Fetching the paper…
Reading the bibliography…
Fast gradient-based optimization algorithms have become increasingly essential for the computationally efficient training of machine learning models.
A survey of optimization methods from a machine learning perspective, 2019a
Shiliang Sun, Zehui Cao, Han Zhu, and Jing Zhao · 1906
Earlier work this paper cites.
First-order preconditioning via hypergradient descent, 2019
Ted Moskovitz, Rui Wang, Janice Lan, Sanyam Kapoor, Thomas Miconi, Jason Yosinski, and Aditya Rawal · 1910
Earlier work this paper cites.
Non-convex optimization for machine learning
Prateek Jain and Purushottam Kar · 1935
Earlier work this paper cites.
A method for the solution of certain non-linear problems in least squares
Kenneth Levenberg · 1944
Earlier work this paper cites.
The perceptron: a probabilistic model for information storage and organization in the brain
Frank Rosenblatt · 1958
Earlier work this paper cites.
The Convergence of a Class of Double-rank Minimization Algorithms 1. General Considerations
C. G. Broyden · 1970
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Y. Lecun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
Large-Scale Quasi-Newton and Partially Separable Optimization , pp. 222–249
Jorge Nocedal and Stephen J. Wright (eds.) · 1999
Earlier work this paper cites.
On the distance between two neural networks and the stability of learning, 2020
Jeremy Bernstein, Arash Vahdat, Yisong Yue, and Ming-Yu Liu · 2002
Earlier work this paper cites.
A stochastic quasi-newton method for online convex optimization
Nicol N. Schraudolph, Jin Yu, and Simon Günter · 2007
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
Luke Metz, Niru Maheswaranathan, C. Daniel Freeman, Ben Poole, and Jascha Sohl-Dickstein · 2009
Earlier work this paper cites.
Ridge rider: Finding diverse solutions by following eigenvectors of the hessian, 2020
Jack Parker-Holder, Luke Metz, Cinjon Resnick, Hengyuan Hu, Adam Lerer, Alistair Letcher, Alex Peysakhovich, Aldo Pacchiano, and Jakob Foerster · 2011
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Cited alongside, same era.
Learning to learn by gradient descent by gradient descent
Marcin Andrychowicz, Misha Denil, Sergio Gomez, Matthew W Hoffman, David Pfau, Tom Schaul, Brendan Shillingford, and Nando De Freitas · 2016
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Ke Li and Jitendra Malik · 2016
Cited alongside, same era.
Conditional image generation with pixelcnn decoders
Aäron van den Oord, Nal Kalchbrenner, Oriol Vinyals, Lasse Espeholt, Alex Graves, and Koray Kavukcuoglu · 2016
Cited alongside, same era.
An intriguing failing of convolutional neural networks and the coordconv solution
Rosanne Liu, Joel Lehman, Piero Molino, Felipe Petroski Such, Eric Frank, Alex Sergeev, and Jason Yosinski · 2018
Later among the works it cites.
Understanding and correcting pathologies in the training of learned optimizers, 2018
Luke Metz, Niru Maheswaranathan, Jeremy Nixon, C. Daniel Freeman, and Jascha Sohl-Dickstein · 2018
Later among the works it cites.
Adaptive methods for nonconvex optimization
Manzil Zaheer, Sashank J. Reddi, Devendra Sachan, Satyen Kale, and Sanjiv Kumar · 2018
Later among the works it cites.
Understanding and correcting pathologies in the training of learned optimizers
Luke Metz, Niru Maheswaranathan, Jeremy Nixon, Daniel Freeman, and Jascha Sohl-Dickstein · 2019
Later among the works it cites.
Meta-curvature
Eunbyung Park and Junier B Oliva · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Atilim Gunes Baydin, Robert Cornish, David Martinez Rubio, Mark Schmidt, and Frank Wood · 2017
Cited alongside, same era.
Tunable efficient unitary neural networks (EUNN) and their application to RNNs
Li Jing, Yichen Shen, Tena Dubcek, John Peurifoy, Scott Skirlo, Yann LeCun, Max Tegmark, and Marin Soljačić · 2017
Cited alongside, same era.
Self-normalizing neural networks
Günter Klambauer, Thomas Unterthiner, Andreas Mayr, and Sepp Hochreiter · 2017
Cited alongside, same era.
Learning to optimize neural nets
Ke Li and Jitendra Malik · 2017
Cited alongside, same era.
Learning gradient descent: Better generalization and longer horizons
Kaifeng Lv, Shunhua Jiang, and Jian Li · 2017
Cited alongside, same era.
Large batch training of convolutional networks, 2017
Yang You, Igor Gitman, and Boris Ginsburg · 2017
Cited alongside, same era.
Fast approximate natural gradient descent in a kronecker-factored eigenbasis
Thomas George, César Laurent, Xavier Bouthillier, Nicolas Ballas, and Pascal Vincent · 2018
Cited alongside, same era.
Practical quasi-newton methods for training deep neural networks
Donald Goldfarb, Yi Ren, and Achraf Bahamou · 2020
Later among the works it cites.
Learning to optimize: A primer and a benchmark, 2021
Tianlong Chen, Xiaohan Chen, Wuyang Chen, Howard Heaton, Jialin Liu, Zhangyang Wang, and Wotao Yin · 2021
Later among the works it cites.
Adahessian: An adaptive second order optimizer for machine learning
Zhewei Yao, Amir Gholami, Sheng Shen, Mustafa Mustafa, Kurt Keutzer, and Michael Mahoney · 2021
Later among the works it cites.
Step-size adaptation using exponentiated gradient updates, 2022
Ehsan Amid, Rohan Anil, Christopher Fifty, and Manfred K. Warmuth · 2022
Closest in time.
Amortized proximal optimization, 2022
Juhan Bae, Paul Vicol, Jeff Z. HaoChen, and Roger Grosse · 2022
Closest in time.
Neural networks for machine learning - lecture 6e
Geoffrey Hinton, Nitish Srivastava, and Kevin Swersky · 2022
Closest in time.
Spectrum of the transposition graph, 2022
Elena V. Konstantinova and Artem Kravchuk · 2022
Closest in time.