Fetching the paper…
Reading the bibliography…
If the trend of learned components eventually outperforming their hand-crafted version continues, learned optimizers will eventually outperform hand-crafted optimizers like SGD or Adam.
Learning an adaptive learning rate schedule
Xu, Z., Dai, A. M., Kemp, J., and Metz, L · 1909
Earlier work this paper cites.
BADGER: Learning to (Learn [Learning Algorithms] through Multi-Agent Communication)
Rosa, M., Afanasjeva, O., Andersson, S., Davidson, J., Guttenberg, N., Hlubuček, P., Poliak, M., Vítků, J., and Feyereisl, J · 1912
Earlier work this paper cites.
Long Short-Term Memory
Hochreiter, S. and Jurgen Schmidhuber · 1997
Earlier work this paper cites.
Learning to Learn Using Gradient Descent
Hochreiter, S., Younger, A. S., and Conwell, P. R · 2001
Earlier work this paper cites.
Adaptive behavior with fixed weights in RNN: an overview
Prokhorov, D., Feldkarnp, L., and Tyukin, I · 2002
Earlier work this paper cites.
Safeguarded Learned Convex Optimization
Heaton, H., Chen, X., Wang, Z., and Yin, W · 2003
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A · 2009
Earlier work this paper cites.
Metz, L., Maheswaranathan, N., Freeman, C. D., Poole, B., and Sohl-Dickstein, J · 2009
Earlier work this paper cites.
MNIST handwritten digit database
LeCun, Y. and Cortes, C · 2010
Earlier work this paper cites.
Reverse engineering learned optimizers reveals known and novel mechanisms
Maheswaranathan, N., Sussillo, D., Metz, L., Sun, R., and Sohl-Dickstein, J · 2011
Earlier work this paper cites.
Scikit-learn: Machine Learning in Python
Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., Blondel, M., Prettenhofer, P., Weiss, R., Dubourg, V., Vanderplas, J., Passos, A., Cournapeau, D., Brucher, M., Perrot, M., and Duchesnay, E · 2011
Cited alongside, same era.
ImageNet Large Scale Visual Recognition Challenge
Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M., Berg, A. C., and Fei-Fei, L · 2015
Cited alongside, same era.
Learning to Learn by Gradient Descent by Gradient Descent
Andrychowicz, M., Denil, M., Colmenarejo, S. G., Hoffman, M. W., Pfau, D., Schaul, T., Shillingford, B., and de Freitas, N · 2016
Cited alongside, same era.
Learning Gradient Descent: Better Generalization and Longer Horizons
Lv, K., Jiang, S., and Li, J · 2017
Cited alongside, same era.
Fashion-MNIST: a Novel Image Dataset for Benchmarking Machine Learning Algorithms
Xiao, H., Rasul, K., and Vollgraf, R · 2017
Meta learning backpropagation and improving it
Kirsch, L. and Schmidhuber, J · 2020
Later among the works it cites.
Improving generalization in meta reinforcement learning using learned objectives
Kirsch, L., van Steenkiste, S., and Schmidhuber, J · 2020
Later among the works it cites.
Stochastic gradient descent: An intuitive proof, 2020
Prazeres, M. O · 2020
Later among the works it cites.
Learning to Optimize: A Primer and A Benchmark
Chen, T., Chen, X., Chen, W., Heaton, H., Liu, J., Wang, Z., and Yin, W · 2021
Later among the works it cites.
Evolving reinforcement learning algorithms
Co-Reyes, J. D., Miao, Y., Peng, D., Real, E., Le, Q. V., Levine, S., Lee, H., and Faust, A · 2021
Later among the works it cites.
Meta-learning in neural networks: A survey
Hospedales, T. M., Antoniou, A., Micaelli, P., and Storkey, A. J · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Understanding Short-Horizon Bias in Stochastic Meta-Optimization
Wu, Y., Ren, M., Liao, R., and Grosse, R · 2018
Cited alongside, same era.
On the fenchel duality between strong convexity and lipschitz continuous gradient
Zhou, X · 2018
Cited alongside, same era.
Convergence theorems for gradient descent, 2019
Gower, R. M · 2019
Cited alongside, same era.
The bitter lesson, 2019
Sutton, R · 2019
Cited alongside, same era.
Meta-learning curiosity algorithms
Alet*, F., Schneider*, M. F., Lozano-Perez, T., and Kaelbling, L. P · 2020
Cited alongside, same era.
Later among the works it cites.
Overcoming barriers to the training of effective learned optimizers
Metz, L., Maheswaranathan, N., Freeman, C. D., Poole, B., and Sohl-Dickstein, J · 2021
Later among the works it cites.
Meta-learning bidirectional update rules
Sandler, M., Vladymyrov, M., Zhmoginov, A., Miller, N., Madams, T., Jackson, A., and Arcas, B. A. Y · 2021
Later among the works it cites.
Efficientnetv2: Smaller models and faster training
Tan, M. and Le, Q · 2021
Later among the works it cites.