Fetching the paper…
Reading the bibliography…
Optimization plays a costly and crucial role in developing machine learning systems.
Using learned optimizers to make models robust to input noise
Luke Metz, Niru Maheswaranathan, Jonathon Shlens, Jascha Sohl-Dickstein, and Ekin D Cubuk · 1906
Earlier work this paper cites.
A stochastic approximation method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
Some methods of speeding up the convergence of iteration methods
Boris T Polyak · 1964
Earlier work this paper cites.
The convergence of a class of double-rank minimization algorithms 1. general considerations
Charles George Broyden · 1970
Earlier work this paper cites.
Quasi-newton methods, motivation and theory
John E Dennis, Jr and Jorge J Moré · 1977
Earlier work this paper cites.
A method for unconstrained convex minimization problem with the rate of convergence o (1/kˆ 2)
Yurii Nesterov · 1983
Earlier work this paper cites.
On the limited memory bfgs method for large scale optimization
Dong C Liu and Jorge Nocedal · 1989
Earlier work this paper cites.
On the optimization of a synaptic learning rule
Samy Bengio, Yoshua Bengio, Jocelyn Cloutier, and Jan Gecsei · 1992
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
The new Palgrave dictionary of economics and the law
Peter Newman · 1998
Earlier work this paper cites.
Evolution and design of distributed learning rules
Thomas Philip Runarsson and Magnus Thor Jonsson · 2000
Earlier work this paper cites.
Learning to learn using gradient descent
Sepp Hochreiter, A Steven Younger, and Peter R Conwell · 2001
Earlier work this paper cites.
Using a thousand optimization tasks to learn hyperparameter search strategies
Luke Metz, Niru Maheswaranathan, Ruoxi Sun, C Daniel Freeman, Ben Poole, and Jascha Sohl-Dickstein · 2002
Earlier work this paper cites.
Matplotlib: A 2d graphics environment
J. D. Hunter · 2007
Earlier work this paper cites.
Cifar-10 and cifar-100 datasets
Alex Krizhevsky, Vinod Nair, and Geoffrey Hinton · 2009
Earlier work this paper cites.
Luke Metz, Niru Maheswaranathan, C Daniel Freeman, Ben Poole, and Jascha Sohl-Dickstein · 2009
Earlier work this paper cites.
Large-scale machine learning with stochastic gradient descent
Léon Bottou · 2010
Earlier work this paper cites.
Algorithms for hyper-parameter optimization
James Bergstra, Rémi Bardenet, Yoshua Bengio, and Balázs Kégl · 2011
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Earlier work this paper cites.
On optimization methods for deep learning
Quoc V Le, Jiquan Ngiam, Adam Coates, Ahbik Lahiri, Bobby Prochnow, and Andrew Y Ng · 2011
Earlier work this paper cites.
The numpy array: a structure for efficient numerical computation
Stefan Van Der Walt, S Chris Colbert, and Gael Varoquaux · 2011
Earlier work this paper cites.
Random search for hyper-parameter optimization
James Bergstra and Yoshua Bengio · 2012
Earlier work this paper cites.
Practical bayesian optimization of machine learning algorithms
Jasper Snoek, Hugo Larochelle, and Ryan P Adams · 2012
Earlier work this paper cites.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
Tijmen Tieleman and Geoffrey Hinton · 2012
Earlier work this paper cites.
Adadelta: an adaptive learning rate method
Matthew D Zeiler · 2012
Earlier work this paper cites.
One billion word benchmark for measuring progress in statistical language modeling
Ciprian Chelba, Tomas Mikolov, Mike Schuster, Qi Ge, Thorsten Brants, and Phillipp Koehn · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2014
Cited alongside, same era.
Deep learning
Yann LeCun, Yoshua Bengio, and Geoffrey Hinton · 2015
Cited alongside, same era.
Gradient-based hyperparameter optimization through reversible learning
Dougal Maclaurin, David Duvenaud, and Ryan Adams · 2015
Cited alongside, same era.
Optimizing neural networks with kronecker-factored approximate curvature
James Martens and Roger Grosse · 2015
Cited alongside, same era.
Learning to learn by gradient descent by gradient descent
Shampoo: Preconditioned stochastic tensor optimization
Vineet Gupta, Tomer Koren, and Yoram Singer · 2018
Later among the works it cites.
Aggregated momentum: Stability through passive damping
James Lucas, Richard Zemel, and Roger Grosse · 2018
Later among the works it cites.
Learning unsupervised learning rules
Luke Metz, Niru Maheswaranathan, Brian Cheung, and Jascha Sohl-Dickstein · 2018
Later among the works it cites.
Adaptive methods for nonconvex optimization
S Reddi, Manzil Zaheer, Devendra Sachan, Satyen Kale, and Sanjiv Kumar · 2018
Later among the works it cites.
Adafactor: Adaptive learning rates with sublinear memory cost
Noam Shazeer and Mitchell Stern · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Marcin Andrychowicz, Misha Denil, Sergio Gomez, Matthew W Hoffman, David Pfau, Tom Schaul, and Nando de Freitas · 2016
Cited alongside, same era.
Learning to learn without gradient descent by gradient descent
Yutian Chen, Matthew W Hoffman, Sergio Gómez Colmenarejo, Misha Denil, Timothy P Lillicrap, Matt Botvinick, and Nando de Freitas · 2016
Cited alongside, same era.
Learning step size controllers for robust neural network training
Christian Daniel, Jonathan Taylor, and Sebastian Nowozin · 2016
Cited alongside, same era.
Incorporating nesterov momentum into adam
Timothy Dozat · 2016
Cited alongside, same era.
Using deep q-learning to control optimization hyperparameters
Samantha Hansen · 2016
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Optimization as a model for few-shot learning
Sachin Ravi and Hugo Larochelle · 2016
Cited alongside, same era.
Optuna: A next-generation hyperparameter optimization framework
Takuya Akiba, Shotaro Sano, Toshihiko Yanase, Takeru Ohta, and Masanori Koyama · 2019
Later among the works it cites.
Memory-efficient adaptive optimization
Rohan Anil, Vineet Gupta, Tomer Koren, and Yoram Singer · 2019
Later among the works it cites.
Learning to optimize in swarms
Yue Cao, Tianlong Chen, Zhangyang Wang, and Yang Shen · 2019
Later among the works it cites.
On the variance of the adaptive learning rate and beyond
Liyuan Liu, Haoming Jiang, Pengcheng He, Weizhu Chen, Xiaodong Liu, Jianfeng Gao, and Jiawei Han · 2019
Later among the works it cites.
Guided evolutionary strategies: Augmenting random search with surrogate gradients
Niru Maheswaranathan, Luke Metz, George Tucker, Dami Choi, and Jascha Sohl-Dickstein · 2019
Later among the works it cites.
Efficientnet: Rethinking model scaling for convolutional neural networks
Mingxing Tan and Quoc Le · 2019
Later among the works it cites.
Learning an adaptive learning rate schedule
Zhen Xu, Andrew M Dai, Jonas Kemp, and Luke Metz · 2019
Later among the works it cites.
Large batch optimization for deep learning: Training bert in 76 minutes
Yang You, Jing Li, Sashank Reddi, Jonathan Hseu, Sanjiv Kumar, Srinadh Bhojanapalli, Xiaodan Song, James Demmel, Kurt Keutzer, and Cho-Jui Hsieh · 2019
Later among the works it cites.
Second order optimization made practical
Rohan Anil, Vineet Gupta, Tomer Koren, Kevin Regan, and Yoram Singer · 2020
Later among the works it cites.
Training stronger baselines for learning to optimize
Tianlong Chen, Weiyi Zhang, Zhou Jingyang, Shiyu Chang, Sijia Liu, Lisa Amini, and Zhangyang Wang · 2020
Later among the works it cites.
Haiku: Sonnet for JAX, 2020
Tom Hennigan, Trevor Cai, Tamara Norman, and Igor Babuschkin · 2020
Later among the works it cites.
Reverse engineering learned optimizers reveals known and novel mechanisms
Niru Maheswaranathan, David Sussillo, Luke Metz, Ruoxi Sun, and Jascha Sohl-Dickstein · 2020
Later among the works it cites.
Google’s training chips revealed: Tpuv2 and tpuv3
Thomas Norrie, Nishant Patil, Doe Hyun Yoon, George Kurian, Sheng Li, James Laudon, Cliff Young, Norman P Jouppi, and David Patterson · 2020
Later among the works it cites.
Descending through a crowded valley–benchmarking deep learning optimizers
Robin M Schmidt, Frank Schneider, and Philipp Hennig · 2020
Later among the works it cites.
Improved adversarial training via learned optimizer
Yuanhao Xiong and Cho-Jui Hsieh · 2020
Later among the works it cites.
A generalizable approach to learning optimizers
Diogo Almeida, Clemens Winter, Jie Tang, and Wojciech Zaremba · 2021
Later among the works it cites.
Learn2hop: Learned optimization on rough landscapes
Amil Merchant, Luke Metz, Sam Schoenholz, and Ekin Dogus Cubuk · 2021
Later among the works it cites.
Training learned optimizers with randomly initialized learned optimizers
Luke Metz, C Daniel Freeman, Niru Maheswaranathan, and Jascha Sohl-Dickstein · 2021
Later among the works it cites.
Learning a minimax optimizer: A pilot study
Jiayi Shen, Xiaohan Chen, Howard Heaton, Tianlong Chen, Jialin Liu, Wotao Yin, and Zhangyang Wang · 2021
Later among the works it cites.
Unbiased gradient estimation in unrolled computation graphs with persistent evolution strategies
Paul Vicol, Luke Metz, and Jascha Sohl-Dickstein · 2021
Later among the works it cites.
Symbolic learning to optimize: Towards interpretability and scalability
Wenqing Zheng, Tianlong Chen, Ting-Kuei Hu, and Zhangyang Wang · 2022
Closest in time.