Fetching the paper…
Reading the bibliography…
A core issue with learning to optimize neural networks has been the lack of generalization to real world problems.
Temporal difference learning and td-gammon
G. Tesauro · 1995
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
Learning to forget: Continual prediction with lstm
F. A. Gers, J. Schmidhuber, and F. Cummins · 1999
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping
A. Y. Ng, D. Harada, and S. Russell · 1999
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
X. Glorot and Y. Bengio · 2010
Earlier work this paper cites.
Practical recommendations for gradient-based training of deep architectures, 2012
Y. Bengio · 2012
Earlier work this paper cites.
Learning phrase representations using rnn encoder-decoder for statistical machine translation, 2014
K. Cho, B. van Merrienboer, C. Gulcehre, D. Bahdanau, F. Bougares, H. Schwenk, and Y. Bengio · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2014
Earlier work this paper cites.
Gradnets: Dynamic interpolation between neural architectures
D. Almeida and N. Sauder · 2015
Earlier work this paper cites.
Stop wasting my gradients: Practical svrg, 2015
R. Babanezhad, M. O. Ahmed, A. Virani, M. Schmidt, J. Konečný, and S. Sallinen · 2015
Earlier work this paper cites.
Deep residual learning for image recognition, 2015
K. He, X. Zhang, S. Ren, and J. Sun · 2015
Earlier work this paper cites.
Gradient-based hyperparameter optimization through reversible learning, 2015
D. Maclaurin, D. Duvenaud, and R. P. Adams · 2015
Earlier work this paper cites.
Deep speech 2: End-to-end speech recognition in english and mandarin
D. Amodei, S. Ananthanarayanan, R. Anubhai, J. Bai, E. Battenberg, C. Case, J. Casper, B. Catanzaro, Q. Cheng, G. Chen, et al · 2016
Earlier work this paper cites.
Learning to learn by gradient descent by gradient descent
M. Andrychowicz, M. Denil, S. Gomez, M. W. Hoffman, D. Pfau, T. Schaul, B. Shillingford, and N. De Freitas · 2016
Earlier work this paper cites.
Learning step size controllers for robust neural network training
C. Daniel, J. Taylor, and S. Nowozin · 2016
Earlier work this paper cites.
Sgdr: Stochastic gradient descent with warm restarts
I. Loshchilov and F. Hutter · 2016
Earlier work this paper cites.
Pointer sentinel mixture models, 2016
S. Merity, C. Xiong, J. Bradbury, and R. Socher · 2016
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. Van Den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, et al · 2016
Cited alongside, same era.
Neural collaborative filtering
X. He, L. Liao, H. Zhang, L. Nie, X. Hu, and T.-S. Chua · 2017
Cited alongside, same era.
Population based training of neural networks, 2017
M. Jaderberg, V. Dalibard, S. Osindero, W. M. Czarnecki, J. Donahue, A. Razavi, O. Vinyals, T. Green, I. Dunning, K. Simonyan, C. Fernando, and K. Kavukcuoglu · 2017
Cited alongside, same era.
Asymmetric actor critic for image-based robot learning, 2017
L. Pinto, M. Andrychowicz, P. Welinder, W. Zaremba, and P. Abbeel · 2017
Cited alongside, same era.
Proximal policy optimization algorithms, 2017
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Cited alongside, same era.
Cyclical learning rates for training neural networks, 2017
Single headed attention rnn: Stop thinking with your head, 2019
S. Merity · 2019
Later among the works it cites.
Understanding and correcting pathologies in the training of learned optimizers
L. Metz, N. Maheswaranathan, J. Nixon, D. Freeman, and J. Sohl-Dickstein · 2019
Later among the works it cites.
Dota 2 with large scale deep reinforcement learning, 2019
OpenAI, :, C. Berner, G. Brockman, B. Chan, V. Cheung, P. Dębiak, C. Dennison, D. Farhi, Q. Fischer, S. Hashme, C. Hesse, R. Józefowicz, S. Gray, C. Olsson, J. Pachocki, M. Petrov, H. P. d. O. Pinto, J. Raiman, T. Salimans, J. Schlatter, J. Schneider, S. Sidor, I. Sutskever, J. Tang, F. Wolski, and S. Zhang · 2019
Later among the works it cites.
Pytorch: An imperative style, high-performance deep learning library
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, et al · 2019
Later among the works it cites.
Language models are unsupervised multitask learners
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
L. N. Smith · 2017
Cited alongside, same era.
Attention is all you need, 2017
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin · 2017
Cited alongside, same era.
Learned optimizers that scale and generalize
O. Wichrowska, N. Maheswaranathan, M. W. Hoffman, S. G. Colmenarejo, M. Denil, N. Freitas, and J. Sohl-Dickstein · 2017
Cited alongside, same era.
Fashion-mnist: a novel image dataset for benchmarking machine learning algorithms, 2017
H. Xiao, K. Rasul, and R. Vollgraf · 2017
Cited alongside, same era.
Reinforcement learning for learning rate control, 2017
C. Xu, T. Qin, G. Wang, and T.-Y. Liu · 2017
Cited alongside, same era.
Large batch training of convolutional networks, 2017
Y. You, I. Gitman, and B. Ginsburg · 2017
Cited alongside, same era.
Online learning rate adaptation with hypergradient descent, 2018
A. G. Baydin, R. Cornish, D. M. Rubio, M. Schmidt, and F. Wood · 2018
Cited alongside, same era.
Huggingface’s transformers: State-of-the-art natural language processing
T. Wolf, L. Debut, V. Sanh, J. Chaumond, C. Delangue, A. Moi, P. Cistac, T. Rault, R. Louf, M. Funtowicz, et al · 2019
Later among the works it cites.
Learning an adaptive learning rate schedule
Z. Xu, A. M. Dai, J. Kemp, and L. Metz · 2019
Later among the works it cites.
Which algorithmic choices matter at which batch sizes? insights from a noisy quadratic model, 2019
G. Zhang, L. Li, Z. Nado, J. Martens, S. Sachdeva, G. E. Dahl, C. J. Shallue, and R. Grosse · 2019
Later among the works it cites.
Second order optimization made practical
R. Anil, V. Gupta, T. Koren, K. Regan, and Y. Singer · 2020
Later among the works it cites.
Language models are few-shot learners, 2020
T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei · 2020
Later among the works it cites.
On empirical comparisons of optimizers for deep learning, 2020
D. Choi, C. J. Shallue, Z. Nado, J. Lee, C. J. Maddison, and G. E. Dahl · 2020
Later among the works it cites.
On the iteration complexity of hypergradient computation, 2020
R. Grazzi, L. Franceschi, M. Pontil, and S. Salzo · 2020
Later among the works it cites.
Scaling laws for neural language models, 2020
J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei · 2020
Later among the works it cites.
The two regimes of deep network training, 2020
G. Leclerc and A. Madry · 2020
Later among the works it cites.
Optimizing neural networks with kronecker-factored approximate curvature, 2020
J. Martens and R. Grosse · 2020
Later among the works it cites.
Tasks, stability, architecture, and compute: Training more effective learned optimizers, and using them to train themselves, 2020
L. Metz, N. Maheswaranathan, C. D. Freeman, B. Poole, and J. Sohl-Dickstein · 2020
Later among the works it cites.
First-order preconditioning via hypergradient descent, 2020
T. Moskovitz, R. Wang, J. Lan, S. Kapoor, T. Miconi, J. Yosinski, and A. Rawal · 2020
Later among the works it cites.
Large batch optimization for deep learning: Training bert in 76 minutes, 2020
Y. You, J. Li, S. Reddi, J. Hseu, S. Kumar, S. Bhojanapalli, X. Song, J. Demmel, K. Keutzer, and C.-J. Hsieh · 2020
Later among the works it cites.