Fetching the paper…
Reading the bibliography…
Optimisers are an essential component for training machine learning models, and their design influences learning speed and generalisation.
A database for handwritten text recognition research
Jonathan J. Hull · 1994
Earlier work this paper cites.
Exponentiated gradient versus gradient descent for linear predictors
Jyrki Kivinen and Manfred K Warmuth · 1997
Earlier work this paper cites.
Rademacher and gaussian complexities: Risk bounds and structural results
Peter L Bartlett and Shahar Mendelson · 2002
Earlier work this paper cites.
MNIST handwritten digit database
Yann LeCun and Corinna Cortes · 2010
Earlier work this paper cites.
An analysis of single-layer networks in unsupervised feature learning
Adam Coates, Andrew Ng, and Honglak Lee · 2011
Earlier work this paper cites.
Reading digits in natural images with unsupervised feature learning
Yuval Netzer, Tao Wang, Adam Coates, Alessandro Bissacco, Bo Wu, and Andrew Y Ng · 2011
Earlier work this paper cites.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
Tijmen Tieleman and Geoffrey Hinton · 2012
Earlier work this paper cites.
Understanding machine learning: From theory to algorithms
Shai Shalev-Shwartz and Shai Ben-David · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Learning to learn by gradient descent by gradient descent
Marcin Andrychowicz, Misha Denil, Sergio Gomez, Matthew W Hoffman, David Pfau, Tom Schaul, Brendan Shillingford, and Nando De Freitas · 2016
Earlier work this paper cites.
Neural architecture search with reinforcement learning
Barret Zoph and Quoc V Le · 2016
Earlier work this paper cites.
Spectrally-normalized margin bounds for neural networks
Peter L Bartlett, Dylan J Foster, and Matus J Telgarsky · 2017
Earlier work this paper cites.
Neural optimizer search with reinforcement learning
Irwan Bello, Barret Zoph, Vijay Vasudevan, and Quoc V Le · 2017
Cited alongside, same era.
Learning to learn without gradient descent by gradient descent
Yutian Chen, Matthew W. Hoffman, Sergio Gomez Colmenarejo, Misha Denil, Timothy P. Lillicrap, Matt Botvinick, and Nando de Freitas · 2017
Cited alongside, same era.
Model-agnostic meta-learning for fast adaptation of deep networks
Chelsea Finn, Pieter Abbeel, and Sergey Levine · 2017
Cited alongside, same era.
Forward and reverse gradient-based hyperparameter optimization
Luca Franceschi, Michele Donini, Paolo Frasconi, and Massimiliano Pontil · 2017
Cited alongside, same era.
Learning to optimize
Ke Li and Jitendra Malik · 2017
Cited alongside, same era.
Meta-curvature
Eunbyung Park and Junier B Oliva · 2019
Later among the works it cites.
Regularized evolution for image classifier architecture search
Esteban Real, Alok Aggarwal, Yanping Huang, and Quoc V Le · 2019
Later among the works it cites.
Cold case: The lost mnist digits
Chhavi Yadav and Léon Bottou · 2019
Later among the works it cites.
BoTorch: A Framework for Efficient Monte-Carlo Bayesian Optimization
Maximilian Balandat, Brian Karrer, Daniel R. Jiang, Samuel Daulton, Benjamin Letham, Andrew Gordon Wilson, and Eytan Bakshy · 2020
Later among the works it cites.
Meta-learning with warped gradient descent
Sebastian Flennerhag, Andrei A Rusu, Razvan Pascanu, Francesco Visin, Hujun Yin, and Raia Hadsell · 2020
Later among the works it cites.
Meta-learning in neural networks: A survey
Timothy Hospedales, Antreas Antoniou, Paul Micaelli, and Amos Storkey · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zhenguo Li, Fengwei Zhou, Fei Chen, and Hang Li · 2017
Cited alongside, same era.
Learned optimizers that scale and generalize
Olga Wichrowska, Niru Maheswaranathan, Matthew W Hoffman, Sergio Gomez Colmenarejo, Misha Denil, Nando Freitas, and Jascha Sohl-Dickstein · 2017
Cited alongside, same era.
Fashion-mnist: a novel image dataset for benchmarking machine learning algorithms
Han Xiao, Kashif Rasul, and Roland Vollgraf · 2017
Cited alongside, same era.
How to train your MAML
Antreas Antoniou, Harrison Edwards, and Amos J. Storkey · 2018
Cited alongside, same era.
Metareg: Towards domain generalization using meta-regularization
Yogesh Balaji, Swami Sankaranarayanan, and Rama Chellappa · 2018
Cited alongside, same era.
Deep learning for classical japanese literature
Tarin Clanuwat, Mikel Bober-Irizar, Asanobu Kitamoto, Alex Lamb, Kazuaki Yamamoto, and David Ha · 2018
Cited alongside, same era.
Later among the works it cites.
Generalization bounds for deep convolutional neural networks
Philip M Long and Hanie Sedghi · 2020
Later among the works it cites.
Searching for robustness: Loss learning for noisy classification tasks
Boyan Gao, Henry Gouk, and Timothy M. Hospedales · 2021
Later among the works it cites.
Distance-based regularisation of deep networks for fine-tuning
Henry Gouk, Timothy M Hospedales, and Massimiliano Pontil · 2021
Later among the works it cites.
Gradient-based hyperparameter optimization over long horizons
Paul Micaelli and Amos J Storkey · 2021
Later among the works it cites.
Meta-learning bidirectional update rules
Mark Sandler, Max Vladymyrov, Andrey Zhmoginov, Nolan Miller, Tom Madams, Andrew Jackson, and Blaise Agüera Y Arcas · 2021
Later among the works it cites.