Fetching the paper…
Reading the bibliography…
Much as replacing hand-designed features with learned functions has revolutionized how we solve perceptual tasks, we believe learned algorithms will transform how we train models.
Using learned optimizers to make models robust to input noise
Luke Metz, Niru Maheswaranathan, Jonathon Shlens, Jascha Sohl-Dickstein, and Ekin D Cubuk · 1906
Earlier work this paper cites.
Ai memo 39-the new compiler
M Levin T Hart and Mike Levin · 1962
Earlier work this paper cites.
About convergence of random search method in extremal control of multi-parameter systems
LA Rastrigin · 1963
Earlier work this paper cites.
Evolutionsstrategie–optimierung technisher systeme nach prinzipien der biologischen evolution
Ingo Rechenberg · 1973
Earlier work this paper cites.
Updating quasi-newton matrices with limited storage
Jorge Nocedal · 1980
Earlier work this paper cites.
On the limited memory bfgs method for large scale optimization
Dong C Liu and Jorge Nocedal · 1989
Earlier work this paper cites.
Backpropagation through time: what it does and how to do it
Paul J Werbos · 1990
Earlier work this paper cites.
On the optimization of a synaptic learning rule
Samy Bengio, Yoshua Bengio, Jocelyn Cloutier, and Jan Gecsei · 1992
Earlier work this paper cites.
An investigation of the gradient descent process in neural networks
Barak Pearlmutter · 1996
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
The mnist database of handwritten digits
Yann LeCun · 1998
Earlier work this paper cites.
Evolution and design of distributed learning rules
Thomas Philip Runarsson and Magnus Thor Jonsson · 2000
Earlier work this paper cites.
Gradient-based optimization of hyperparameters
Yoshua Bengio · 2000
Earlier work this paper cites.
The google file system
Sanjay Ghemawat, Howard Gobioff, and Shun-Tak Leung · 2003
Earlier work this paper cites.
A guide to NumPy , volume 1
Travis E Oliphant · 2006
Earlier work this paper cites.
Matplotlib: A 2d graphics environment
J. D. Hunter · 2007
Earlier work this paper cites.
Cifar-10 and cifar-100 datasets
Alex Krizhevsky, Vinod Nair, and Geoffrey Hinton · 2009
Earlier work this paper cites.
Relative entropy policy search
Jan Peters, Katharina Mulling, and Yasemin Altun · 2010
Earlier work this paper cites.
The numpy array: a structure for efficient numerical computation
Stefan Van Der Walt, S Chris Colbert, and Gael Varoquaux · 2011
Earlier work this paper cites.
Random gradient-free minimization of convex functions
Yurii Nesterov and Vladimir Spokoiny · 2011
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
Tijmen Tieleman and Geoffrey Hinton · 2012
Earlier work this paper cites.
Auto-encoding variational bayes
Diederik P Kingma and Max Welling · 2013
Earlier work this paper cites.
Advances in optimizing recurrent networks
Yoshua Bengio, Nicolas Boulanger-Lewandowski, and Razvan Pascanu · 2013
Earlier work this paper cites.
Empirical evaluation of gated recurrent neural networks on sequence modeling
Junyoung Chung, Caglar Gulcehre, KyungHyun Cho, and Yoshua Bengio · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Cited alongside, same era.
Alex Graves, Greg Wayne, and Ivo Danihelka · 2014
Cited alongside, same era.
Automatic differentiation of algorithms for machine learning
Atilim Gunes Baydin and Barak A Pearlmutter · 2014
Cited alongside, same era.
Deep knowledge tracing
Chris Piech, Jonathan Bassen, Jonathan Huang, Surya Ganguli, Mehran Sahami, Leonidas J Guibas, and Jascha Sohl-Dickstein · 2015
Unbiasing truncated backpropagation through time
Corentin Tallec and Yann Ollivier · 2017
Later among the works it cites.
Model-agnostic meta-learning for fast adaptation of deep networks
Chelsea Finn, Pieter Abbeel, and Sergey Levine · 2017
Later among the works it cites.
Evolution strategies as a scalable alternative to reinforcement learning
Tim Salimans, Jonathan Ho, Xi Chen, Szymon Sidor, and Ilya Sutskever · 2017
Later among the works it cites.
Learning unsupervised learning rules
Luke Metz, Niru Maheswaranathan, Brian Cheung, and Jascha Sohl-Dickstein · 2018
Later among the works it cites.
Bilevel programming for hyperparameter optimization and meta-learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
ImageNet Large Scale Visual Recognition Challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li Fei-Fei · 2015
Cited alongside, same era.
Taking the human out of the loop: A review of bayesian optimization
Bobak Shahriari, Kevin Swersky, Ziyu Wang, Ryan P Adams, and Nando De Freitas · 2015
Cited alongside, same era.
Gradient-based hyperparameter optimization through reversible learning
Dougal Maclaurin, David Duvenaud, and Ryan Adams · 2015
Cited alongside, same era.
Learning step size controllers for robust neural network training
Christian Daniel, Jonathan Taylor, and Sebastian Nowozin · 2016
Cited alongside, same era.
Learning to learn by gradient descent by gradient descent
Marcin Andrychowicz, Misha Denil, Sergio Gomez, Matthew W Hoffman, David Pfau, Tom Schaul, and Nando de Freitas · 2016
Cited alongside, same era.
Understanding short-horizon bias in stochastic meta-optimization
Yuhuai Wu, Mengye Ren, Renjie Liao, and Roger B Grosse · 2016
Cited alongside, same era.
Incorporating nesterov momentum into adam
Timothy Dozat · 2016
Cited alongside, same era.
Luca Franceschi, Paolo Frasconi, Saverio Salzo, and Massimilano Pontil · 2018
Later among the works it cites.
Learning transferable architectures for scalable image recognition
Barret Zoph, Vijay Vasudevan, Jonathon Shlens, and Quoc V Le · 2018
Later among the works it cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Later among the works it cites.
Adafactor: Adaptive learning rates with sublinear memory cost
Noam Shazeer and Mitchell Stern · 2018
Later among the works it cites.
Quantifying generalization in reinforcement learning
Karl Cobbe, Oleg Klimov, Chris Hesse, Taehoon Kim, and John Schulman · 2018
Later among the works it cites.
On first-order meta-learning algorithms
Alex Nichol, Joshua Achiam, and John Schulman · 2018
Later among the works it cites.
Structured evolution with compact architectures for scalable policy optimization
Krzysztof Choromanski, Mark Rowland, Vikas Sindhwani, Richard E Turner, and Adrian Weller · 2018
Later among the works it cites.
Dota 2 with large scale deep reinforcement learning
Christopher Berner, Greg Brockman, Brooke Chan, Vicki Cheung, Przemysław Dębiak, Christy Dennison, David Farhi, Quirin Fischer, Shariq Hashme, Chris Hesse, et al · 2019
Later among the works it cites.
Alphastar: Mastering the real-time strategy game starcraft ii
Oriol Vinyals, Igor Babuschkin, Junyoung Chung, Michael Mathieu, Max Jaderberg, Wojciech M Czarnecki, Andrew Dudzik, Aja Huang, Petko Georgiev, Richard Powell, et al · 2019
Later among the works it cites.
On empirical comparisons of optimizers for deep learning
Dami Choi, Christopher J Shallue, Zachary Nado, Jaehoon Lee, Chris J Maddison, and George E Dahl · 2019
Later among the works it cites.
Learning an adaptive learning rate schedule
Zhen Xu, Andrew M Dai, Jonas Kemp, and Luke Metz · 2019
Later among the works it cites.
Meta-learning biologically plausible semi-supervised update rules
Keren Gu, Sam Greydanus, Luke Metz, Niru Maheswaranathan, and Jascha Sohl-Dickstein · 2019
Later among the works it cites.
On the tunability of optimizers in deep learning
Prabhu Teja Sivaprasad, Florian Mai, Thijs Vogels, Martin Jaggi, and François Fleuret · 2019
Later among the works it cites.
Energy and policy considerations for deep learning in nlp
Emma Strubell, Ananya Ganesh, and Andrew McCallum · 2019
Later among the works it cites.
Risks from learned optimization in advanced machine learning systems
Evan Hubinger, Chris van Merwijk, Vladimir Mikulik, Joar Skalse, and Scott Garrabrant · 2019
Later among the works it cites.
Guided evolutionary strategies: Augmenting random search with surrogate gradients
Niru Maheswaranathan, Luke Metz, George Tucker, Dami Choi, and Jascha Sohl-Dickstein · 2019
Later among the works it cites.
Using a thousand optimization tasks to learn hyperparameter search strategies
Luke Metz, Niru Maheswaranathan, Ruoxi Sun, C Daniel Freeman, Ben Poole, and Jascha Sohl-Dickstein · 2020
Closest in time.
Persisted evolutionary strategies
Paul Vicol and et. al · 2020
Closest in time.
Array programming with numpy
Charles R Harris, K Jarrod Millman, Stéfan J van der Walt, Ralf Gommers, Pauli Virtanen, David Cournapeau, Eric Wieser, Julian Taylor, Sebastian Berg, Nathaniel J Smith, et al · 2020
Closest in time.
mwaskom/seaborn, September 2020
Michael Waskom and the seaborn development team · 2020
Closest in time.