Fetching the paper…
Reading the bibliography…
While deep learning models have replaced hand-designed features across many domains, these models are still trained with hand-designed optimizers.
Using learned optimizers to make models robust to input noise
Luke Metz, Niru Maheswaranathan, Jonathon Shlens, Jascha Sohl-Dickstein, and Ekin D Cubuk · 1906
Earlier work this paper cites.
AI memo 39: The new compiler
M Hart, T Levin, and Mike Levin · 1962
Earlier work this paper cites.
Evolutionsstrategie–Optimierung technisher Systeme nach Prinzipien der biologischen Evolution
Ingo Rechenberg · 1973
Earlier work this paper cites.
On the optimization of a synaptic learning rule
Samy Bengio, Yoshua Bengio, Jocelyn Cloutier, and Jan Gecsei · 1992
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
The MNIST database of handwritten digits
Yann LeCun · 1998
Earlier work this paper cites.
Evolution and design of distributed learning rules
Thomas Philip Runarsson and Magnus Thor Jonsson · 2000
Earlier work this paper cites.
Using a thousand optimization tasks to learn hyperparameter search strategies
Luke Metz, Niru Maheswaranathan, Ruoxi Sun, C Daniel Freeman, Ben Poole, and Jascha Sohl-Dickstein · 2002
Earlier work this paper cites.
Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks
Alex Graves, Santiago Fernández, Faustino Gomez, and Jürgen Schmidhuber · 2006
Earlier work this paper cites.
Visualizing data using t-SNE
Laurens van der Maaten and Geoffrey Hinton · 2008
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky · 2009
Earlier work this paper cites.
Luke Metz, Niru Maheswaranathan, C Daniel Freeman, Ben Poole, and Jascha Sohl-Dickstein · 2009
Earlier work this paper cites.
Evolution and future directions of large-scale storage and computation systems at Google
Jeffrey Dean · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Earlier work this paper cites.
Random gradient-free minimization of convex functions
Yurii Nesterov and Vladimir Spokoiny · 2011
Earlier work this paper cites.
Reading digits in natural images with unsupervised feature learning
Yuval Netzer, Tao Wang, Adam Coates, Alessandro Bissacco, Bo Wu, and Andrew Y Ng · 2011
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Lecture 6.5—RmsProp: Divide the gradient by a running average of its recent magnitude
T. Tieleman and G. Hinton · 2012
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Marc G Bellemare, Yavar Naddaf, Joel Veness, and Michael Bowling · 2013
Earlier work this paper cites.
One billion word benchmark for measuring progress in statistical language modeling
Ciprian Chelba, Tomas Mikolov, Mike Schuster, Qi Ge, Thorsten Brants, Phillipp Koehn, and Tony Robinson · 2013
Earlier work this paper cites.
Commentary: The materials project: A materials genome approach to accelerating materials innovation
Anubhav Jain, Shyue Ping Ong, Geoffroy Hautier, Wei Chen, William Davidson Richards, Stephen Dacek, Shreyas Cholia, Dan Gunter, David Skinner, Gerbrand Ceder, et al · 2013
Earlier work this paper cites.
Auto-encoding variational Bayes
Diederik P Kingma and Max Welling · 2013
Earlier work this paper cites.
Rectifier nonlinearities improve neural network acoustic models
Andrew L Maas, Awni Y Hannun, and Andrew Y Ng · 2013
Earlier work this paper cites.
Empirical evaluation of gated recurrent neural networks on sequence modeling
Junyoung Chung, Caglar Gulcehre, KyungHyun Cho, and Yoshua Bengio · 2014
Earlier work this paper cites.
Deep speech: Scaling up end-to-end speech recognition
Awni Hannun, Carl Case, Jared Casper, Bryan Catanzaro, Greg Diamos, Erich Elsen, Ryan Prenger, Sanjeev Satheesh, Shubho Sengupta, Adam Coates, et al · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
One weird trick for parallelizing convolutional neural networks
Alex Krizhevsky · 2014
Earlier work this paper cites.
Microsoft COCO: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Earlier work this paper cites.
Monte Carlo Theory, Methods and Examples (book draft), 2014
Art B Owen · 2014
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on Imagenet classification
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
Gradient-based hyperparameter optimization through reversible learning
Dougal Maclaurin, David Duvenaud, and Ryan Adams · 2015
Earlier work this paper cites.
Optimizing neural networks with Kronecker-factored approximate curvature
James Martens and Roger Grosse · 2015
Earlier work this paper cites.
Librispeech: An ASR corpus based on public domain audio books
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur · 2015
Earlier work this paper cites.
Faster R-CNN: Towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun · 2015
Earlier work this paper cites.
ImageNet Large Scale Visual Recognition Challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li Fei-Fei · 2015
Earlier work this paper cites.
Learning to learn by gradient descent by gradient descent
Marcin Andrychowicz, Misha Denil, Sergio Gomez, Matthew W Hoffman, David Pfau, Tom Schaul, and Nando de Freitas · 2016
Earlier work this paper cites.
Learning to learn without gradient descent by gradient descent
Yutian Chen, Matthew W Hoffman, Sergio Gómez Colmenarejo, Misha Denil, Timothy P Lillicrap, Matt Botvinick, and Nando de Freitas · 2016
Earlier work this paper cites.
Learning step size controllers for robust neural network training
Christian Daniel, Jonathan Taylor, and Sebastian Nowozin · 2016
Earlier work this paper cites.
Incorporating Nesterov momentum into Adam
Timothy Dozat · 2016
Earlier work this paper cites.
David Ha, Andrew Dai, and Quoc V Le · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Gaussian error linear units (GELUs)
Dan Hendrycks and Kevin Gimpel · 2016
Earlier work this paper cites.
3d simulation for robot arm control with deep q-learning
Stephen James and Edward Johns · 2016
Earlier work this paper cites.
Optimization as a model for few-shot learning
Sachin Ravi and Hugo Larochelle · 2016
Earlier work this paper cites.
Cad2rl: Real single-image flight without a single real image
Fereshteh Sadeghi and Sergey Levine · 2016
Cited alongside, same era.
Neural optimizer search with reinforcement learning
Irwan Bello, Barret Zoph, Vijay Vasudevan, and Quoc Le · 2017
Cited alongside, same era.
Findings of the 2017 conference on machine translation (wmt17)
Ondrej Bojar, Rajen Chatterjee, Christian Federmann, Yvette Graham, Barry Haddow, Shujian Huang, Matthias Huck, Philipp Koehn, Qun Liu, Varvara Logacheva, Christof Monz, Matteo Negri, Matt Post, Raphael Rubino, Lucia Specia, and Marco Turchi · 2017
Cited alongside, same era.
Neural message passing for quantum chemistry
Justin Gilmer, Samuel S Schoenholz, Patrick F Riley, Oriol Vinyals, and George E Dahl · 2017
Cited alongside, same era.
Accurate, large minibatch SGD: Training Imagenet in 1 hour
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He · 2017
Disentangling adaptive gradient methods from learning rates
Naman Agarwal, Rohan Anil, Elad Hazan, Tomer Koren, and Cyril Zhang · 2020
Later among the works it cites.
Second order optimization made practical
Rohan Anil, Vineet Gupta, Tomer Koren, Kevin Regan, and Yoram Singer · 2020
Later among the works it cites.
The DeepMind JAX Ecosystem, 2020
Igor Babuschkin, Kate Baumli, Alison Bell, Surya Bhupatiraju, Jake Bruce, Peter Buchlovsky, David Budden, Trevor Cai, Aidan Clark, Ivo Danihelka, Claudio Fantacci, Jonathan Godwin, Chris Jones, Tom Hennigan, Matteo Hessel, Steven Kapturowski, Thomas Keck, Iurii Kemaev, Michael King, Lena Martens, Hamza Merzic, Vladimir Mikulik, Tamara Norman, John Quan, George Papamakarios, Roman Ring, Francisco Ruiz, Alvaro Sanchez, Rosalia Schneider, Eren Sezener, Stephen Spencer, Srivatsan Srinivasan, Wojciech Stokowiec, and Fabio Viola · 2020
Later among the works it cites.
On the distance between two neural networks and the stability of learning
Jeremy Bernstein, Arash Vahdat, Yisong Yue, and Ming-Yu Liu · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Self-normalizing neural networks
Günter Klambauer, Thomas Unterthiner, Andreas Mayr, and Sepp Hochreiter · 2017
Cited alongside, same era.
Generalizing Hamiltonian Monte Carlo with neural networks
Daniel Levy, Matthew D Hoffman, and Jascha Sohl-Dickstein · 2017
Cited alongside, same era.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2017
Cited alongside, same era.
Learning gradient descent: Better generalization and longer horizons
Kaifeng Lv, Shunhua Jiang, and Jian Li · 2017
Cited alongside, same era.
The effectiveness of data augmentation in image classification using deep learning
Luis Perez and Jason Wang · 2017
Cited alongside, same era.
Searching for activation functions
Prajit Ramachandran, Barret Zoph, and Quoc V Le · 2017
Cited alongside, same era.
Evolution strategies as a scalable alternative to reinforcement learning
Tim Salimans, Jonathan Ho, Xi Chen, Szymon Sidor, and Ilya Sutskever · 2017
Cited alongside, same era.
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Later among the works it cites.
Training stronger baselines for learning to optimize
Tianlong Chen, Weiyi Zhang, Zhou Jingyang, Shiyu Chang, Sijia Liu, Lisa Amini, and Zhangyang Wang · 2020
Later among the works it cites.
JaxNeRF: An efficient JAX implementation of NeRF, 2020
Boyang Deng, Jonathan T. Barron, and Pratul P. Srinivasan · 2020
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2020
Later among the works it cites.
Conformer: Convolution-augmented transformer for speech recognition
Anmol Gulati, James Qin, Chung-Cheng Chiu, Niki Parmar, Yu Zhang, Jiahui Yu, Wei Han, Shibo Wang, Zhengdong Zhang, Yonghui Wu, et al · 2020
Later among the works it cites.
Open graph benchmark: Datasets for machine learning on graphs
Weihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong, Hongyu Ren, Bowen Liu, Michele Catasta, and Jure Leskovec · 2020
Later among the works it cites.
A domain-specific supercomputer for training deep neural networks
Norman P Jouppi, Doe Hyun Yoon, George Kurian, Sheng Li, Nishant Patil, James Laudon, Cliff Young, and David Patterson · 2020
Later among the works it cites.
Reverse engineering learned optimizers reveals known and novel mechanisms
Niru Maheswaranathan, David Sussillo, Luke Metz, Ruoxi Sun, and Jascha Sohl-Dickstein · 2020
Later among the works it cites.
Mlperf training benchmark
Peter Mattson, Christine Cheng, Gregory Diamos, Cody Coleman, Paulius Micikevicius, David Patterson, Hanlin Tang, Gu-Yeon Wei, Peter Bailis, Victor Bittorf, et al · 2020
Later among the works it cites.
NeRF: Representing scenes as neural radiance fields for view synthesis
Ben Mildenhall, Pratul P Srinivasan, Matthew Tancik, Jonathan T Barron, Ravi Ramamoorthi, and Ren Ng · 2020
Later among the works it cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, Peter J Liu, et al · 2020
Later among the works it cites.
Zero: Memory optimizations toward training trillion parameter models
Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, and Yuxiong He · 2020
Later among the works it cites.
Automl-zero: Evolving machine learning algorithms from scratch
Esteban Real, Chen Liang, David So, and Quoc Le · 2020
Later among the works it cites.
Descending through a crowded valley–benchmarking deep learning optimizers
Robin M Schmidt, Frank Schneider, and Philipp Hennig · 2020
Later among the works it cites.
Improved adversarial training via learned optimizer
Yuanhao Xiong and Cho-Jui Hsieh · 2020
Later among the works it cites.
Adabelief optimizer: Adapting stepsizes by the belief in observed gradients
Juntang Zhuang, Tommy Tang, Yifan Ding, Sekhar C Tatikonda, Nicha Dvornek, Xenophon Papademetris, and James Duncan · 2020
Later among the works it cites.
A generalizable approach to learning optimizers
Diogo Almeida, Clemens Winter, Jie Tang, and Wojciech Zaremba · 2021
Later among the works it cites.
Brax–a differentiable physics engine for large scale rigid body simulation
C Daniel Freeman, Erik Frey, Anton Raichuk, Sertan Girgin, Igor Mordatch, and Olivier Bachem · 2021
Later among the works it cites.
init2winit: a JAX codebase for initialization, optimization, and tuning research
Justin M Gilmer, George E Dahl, and Zachary Nado · 2021
Later among the works it cites.
Learn2Hop: Learned optimization on rough landscapes
Amil Merchant, Luke Metz, Sam Schoenholz, and Ekin Dogus Cubuk · 2021
Later among the works it cites.
Periodic activation functions induce stationarity
Lassi Meronen, Martin Trapp, and Arno Solin · 2021
Later among the works it cites.
Training learned optimizers with randomly initialized learned optimizers
Luke Metz, C Daniel Freeman, Niru Maheswaranathan, and Jascha Sohl-Dickstein · 2021
Later among the works it cites.
Learning a minimax optimizer: A pilot study
Jiayi Shen, Xiaohan Chen, Howard Heaton, Tianlong Chen, Jialin Liu, Wotao Yin, and Zhangyang Wang · 2021
Later among the works it cites.
Searching for efficient transformers for language modeling
David So, Wojciech Mańke, Hanxiao Liu, Zihang Dai, Noam Shazeer, and Quoc V Le · 2021
Later among the works it cites.
Does knowledge distillation really work?
Samuel Stanton, Pavel Izmailov, Polina Kirichenko, Alexander A. Alemi, and Andrew Gordon Wilson · 2021
Later among the works it cites.
Roformer: Enhanced transformer with rotary position embedding
Jianlin Su, Yu Lu, Shengfeng Pan, Bo Wen, and Yunfeng Liu · 2021
Later among the works it cites.
MLP-mixer: An all-MLP architecture for vision
Ilya O Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer, Xiaohua Zhai, Thomas Unterthiner, Jessica Yung, Andreas Steiner, Daniel Keysers, Jakob Uszkoreit, et al · 2021
Later among the works it cites.
Unbiased gradient estimation in unrolled computation graphs with persistent evolution strategies
Paul Vicol, Luke Metz, and Jascha Sohl-Dickstein · 2021
Later among the works it cites.
Launchpad: A programming model for distributed machine learning research
Fan Yang, Gabriel Barth-Maron, Piotr Stańczyk, Matthew Hoffman, Siqi Liu, Manuel Kroiss, Aedan Pope, and Alban Rrustemi · 2021
Later among the works it cites.
Tutorial on amortized optimization for learning to optimize over continuous domains
Brandon Amos · 2022
Closest in time.
Pathways: Asynchronous distributed dataflow for ml
Paul Barham, Aakanksha Chowdhery, Jeff Dean, Sanjay Ghemawat, Steven Hand, Daniel Hurt, Michael Isard, Hyeontaek Lim, Ruoming Pang, Sudip Roy, et al · 2022
Closest in time.
PaLM: Scaling language modeling with pathways
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al · 2022
Closest in time.
Transformer-based learned optimization for physics-based reconstruction of articulated motion
Erik Gärtner, Luke Metz, C. Daniel Freeman, Misha Andriluka, and Cristian Sminchisescu · 2022
Closest in time.
A closer look at learned optimization: Stability, robustness, and inductive biases
James Harrison, Luke Metz, and Jascha Sohl-Dickstein · 2022
Closest in time.
Multi-game decision transformers
Kuang-Huei Lee, Ofir Nachum, Mengjiao Yang, Lisa Lee, Daniel Freeman, Winnie Xu, Sergio Guadarrama, Ian Fischer, Eric Jang, Henryk Michalewski, et al · 2022
Closest in time.
Practical tradeoffs between memory, compute, and performance in learned optimizers
Luke Metz, C Daniel Freeman, James Harrison, Niru Maheswaranathan, and Jascha Sohl-Dickstein · 2022
Closest in time.
A simple guard for learned optimizers
Isabeau Premont-Schwarz, Jaroslav Vitku, and Jan Feyereisl · 2022
Closest in time.
Scaling vision transformers
Xiaohua Zhai, Alexander Kolesnikov, Neil Houlsby, and Lucas Beyer · 2022
Closest in time.
Opt: Open pre-trained transformer language models
Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al · 2022
Closest in time.
Symbolic learning to optimize: Towards interpretability and scalability
Wenqing Zheng, Tianlong Chen, Ting-Kuei Hu, and Zhangyang Wang · 2022
Closest in time.