Fetching the paper…
Reading the bibliography…
Adaptive optimization algorithms such as Adam are widely used in deep learning.
RoBERTa: A Robustly Optimized BERT Pretraining Approach
Liu, Y.; Ott, M.; Goyal, N.; Du, J.; Joshi, M.; Chen, D.; Levy, O.; Lewis, M.; Zettlemoyer, L.; and Stoyanov, V. 2019b · 1907
Earlier work this paper cites.
Transformers without Tears: Improving the Normalization of Self-Attention
Nguyen, T. Q.; and Salazar, J. 2019 · 1910
Earlier work this paper cites.
Some methods of speeding up the convergence of iteration methods
Polyak, B. 1964 · 1964
Earlier work this paper cites.
A method for solving the convex programming problem with convergence rate 𝒪 ( 1 / k 2 ) \mathcal{O}(1/k^{2})
Nesterov, Y. E. 1983 · 1983
Earlier work this paper cites.
Language Models are Few-Shot Learners
Brown, T. B.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; Agarwal, S.; Herbert-Voss, A.; Krueger, G.; Henighan, T.; Child, R.; Ramesh, A.; Ziegler, D. M.; Wu, J.; Winter, C.; Hesse, C.; Chen, M.; Sigler, E.; Litwin, M.; Gray, S.; Chess, B.; Clark, J.; Berner, C.; McCandlish, S.; Radford, A.; Sutskever, I.; and Amodei, D. 2020 · 2005
Earlier work this paper cites.
Adaptive Subgradient Methods for Online Learning and Stochastic Optimization
Duchi, J. C.; Hazan, E.; and Singer, Y. 2010 · 2010
Earlier work this paper cites.
Neural networks for machine learning: Lecture 6a
Hinton, G.; Srivastava, N.; and Swersky, K. 2012 · 2012
Earlier work this paper cites.
ImageNet Classification with Deep Convolutional Neural Networks
Krizhevsky, A.; Sutskever, I.; and Hinton, G. E. 2012 · 2012
Earlier work this paper cites.
Adam: A Method for Stochastic Optimization
Kingma, D. P.; and Ba, J. 2014 · 2014
Earlier work this paper cites.
ImageNet Large Scale Visual Recognition Challenge
Russakovsky, O.; Deng, J.; Su, H.; Krause, J.; Satheesh, S.; Ma, S.; Huang, Z.; Karpathy, A.; Khosla, A.; Bernstein, M.; Berg, A. C.; and Fei-Fei, L. 2015 · 2015
Earlier work this paper cites.
Deep Residual Learning for Image Recognition
He, K.; Zhang, X.; Ren, S.; and Sun, J. 2016 · 2016
Earlier work this paper cites.
PyTorch Examples
Paszke, A.; Gross, S.; Chintala, S.; Chanan, G.; Yang, E.; DeVito, Z.; Lin, Z.; Desmaison, A.; Antiga, L.; and Lerer, A. 2016 · 2016
Earlier work this paper cites.
EMNIST: Extending MNIST to handwritten letters
Cohen, G.; Afshar, S.; Tapson, J.; and van Schaik, A. 2017 · 2017
Cited alongside, same era.
Convolutional Sequence to Sequence Learning
Gehring, J.; Auli, M.; Grangier, D.; Yarats, D.; and Dauphin, Y. N. 2017 · 2017
Cited alongside, same era.
Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour
Goyal, P.; Dollár, P.; Girshick, R. B.; Noordhuis, P.; Wesolowski, L.; Kyrola, A.; Tulloch, A.; Jia, Y.; and He, K. 2017 · 2017
Cited alongside, same era.
Improving Generalization Performance by Switching from Adam to SGD
Keskar, N. S.; and Socher, R. 2017 · 2017
Cited alongside, same era.
Stochastic Modified Equations and Adaptive Stochastic Gradient Algorithms
Li, Q.; Tai, C.; and E, W. 2017 · 2017
Cited alongside, same era.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Devlin, J.; Chang, M.; Lee, K.; and Toutanova, K. 2018 · 2018
Later among the works it cites.
Don’t Decay the Learning Rate, Increase the Batch Size
Smith, S. L.; Kindermans, P.; Ying, C.; and Le, Q. V. 2018 · 2018
Later among the works it cites.
A Closer Look at Deep Learning Heuristics: Learning rate restarts, Warmup and Distillation
Gotmare, A.; Keskar, N. S.; Xiong, C.; and Socher, R. 2019 · 2019
Closest in time.
Use What You Have: Video retrieval using representations from collaborative experts
Liu, Y.; Albanie, S.; Nagrani, A.; and Zisserman, A. 2019a · 2019
Closest in time.
Decoupled Weight Decay Regularization
Loshchilov, I.; and Hutter, F. 2019 · 2019
Closest in time.
Adaptive Gradient Methods with Dynamic Bound of Learning Rate
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Automatic differentiation in PyTorch
Paszke, A.; Gross, S.; Chintala, S.; Chanan, G.; Yang, E.; DeVito, Z.; Lin, Z.; Desmaison, A.; Antiga, L.; and Lerer, A. 2017 · 2017
Cited alongside, same era.
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, L.; and Polosukhin, I. 2017 · 2017
Cited alongside, same era.
YellowFin and the Art of Momentum Tuning
Zhang, J.; Mitliagkas, I.; and Ré, C. 2017 · 2017
Cited alongside, same era.
Adaptive Input Representations for Neural Language Modeling
Baevski, A.; and Auli, M. 2018 · 2018
Cited alongside, same era.
Online Learning Rate Adaptation with Hypergradient Descent
Baydin, A. G.; Cornish, R.; Martínez-Rubio, D.; Schmidt, M.; and Wood, F. 2018 · 2018
Cited alongside, same era.
Closing the Generalization Gap of Adaptive Gradient Methods in Training Deep Neural Networks
Chen, J.; and Gu, Q. 2018 · 2018
Cited alongside, same era.
Luo, L.; Xiong, Y.; Liu, Y.; and Sun, X. 2019 · 2019
Closest in time.
Quasi-hyperbolic momentum and Adam for deep learning
Ma, J.; and Yarats, D. 2019 · 2019
Closest in time.
fairseq: A Fast, Extensible Toolkit for Sequence Modeling
Ott, M.; Edunov, S.; Baevski, A.; Fan, A.; Gross, S.; Ng, N.; Grangier, D.; and Auli, M. 2019 · 2019
Closest in time.
Language models are unsupervised multitask learners
Radford, A.; Wu, J.; Child, R.; Luan, D.; Amodei, D.; and Sutskever, I. 2019 · 2019
Closest in time.
On the Variance of the Adaptive Learning Rate and Beyond
Liu, L.; Jiang, H.; He, P.; Chen, W.; Liu, X.; Gao, J.; and Han, J. 2020 · 2020
Closest in time.
Parallel WaveGAN: A fast waveform generation model based on generative adversarial networks with multi-resolution spectrogram
Yamamoto, R.; Song, E.; and Kim, J.-M. 2020 · 2020
Closest in time.