Attention is all you need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017 · 2017
Later among the works it cites.
Deep & cross network for ad click predictions
Wang, R.; Fu, B.; Fu, G.; and Wang, M. 2017 · 2017
Later among the works it cites.
Second-order optimization for non-convex machine learning: An empirical study
Original
Xu, P.; Roosta-Khorasan, F.; and Mahoney, M. W. 2017 · 2017
Later among the works it cites.
A progressive batching L-BFGS method for machine learning
Original
Bollapragada, R.; Mudigere, D.; Nocedal, J.; Shi, H.-J. M.; and Tang, P. T. P. 2018 · 2018
Later among the works it cites.
Optimization methods for large-scale machine learning
Bottou, L.; Curtis, F. E.; and Nocedal, J. 2018 · 2018
Later among the works it cites.
Accelerated methods for nonconvex optimization
Carmon, Y.; Duchi, J. C.; Hinder, O.; and Sidford, A. 2018 · 2018
Later among the works it cites.
Shampoo: Preconditioned stochastic tensor optimization
Original
Gupta, V.; Koren, T.; and Singer, Y. 2018 · 2018
Later among the works it cites.
Scaling Neural Machine Translation
Ott, M.; Edunov, S.; Grangier, D.; and Auli, M. 2018 · 2018
Later among the works it cites.
Sentence encoders on stilts: Supplementary training on intermediate labeled-data tasks
Original
Phang, J.; Févry, T.; and Bowman, S. R. 2018 · 2018
Later among the works it cites.
How does batch normalization help optimization?
Santurkar, S.; Tsipras, D.; Ilyas, A.; and Madry, A. 2018 · 2018
Later among the works it cites.
Adafactor: Adaptive Learning Rates with Sublinear Memory Cost
Shazeer, N.; and Stern, M. 2018 · 2018
Later among the works it cites.
Exact and inexact subsampled Newton methods for optimization
Bollapragada, R.; Byrd, R. H.; and Nocedal, J. 2019 · 2019
Later among the works it cites.
Entropy-sgd: Biasing gradient descent into wide valleys
Chaudhari, P.; Choromanska, A.; Soatto, S.; LeCun, Y.; Baldassi, C.; Borgs, C.; Chayes, J.; Sagun, L.; and Zecchina, R. 2019 · 2019
Later among the works it cites.
Transformer-XL: Attentive Language Models beyond a Fixed-Length Context
Dai, Z.; Yang, Z.; Yang, Y.; Carbonell, J. G.; Le, Q.; and Salakhutdinov, R. 2019 · 2019
Later among the works it cites.
Decoupled Weight Decay Regularization
Loshchilov, I.; and Hutter, F. 2019 · 2019
Later among the works it cites.
A tensorized transformer for language modeling
Ma, X.; Zhang, P.; Zhang, S.; Duan, N.; Hou, Y.; Zhou, M.; and Song, D. 2019 · 2019
Later among the works it cites.
fairseq: A Fast, Extensible Toolkit for Sequence Modeling
Ott, M.; Edunov, S.; Baevski, A.; Fan, A.; Gross, S.; Ng, N.; Grangier, D.; and Auli, M. 2019 · 2019
Later among the works it cites.
PyTorch: An Imperative Style, High-Performance Deep Learning Library
Paszke, A.; Gross, S.; Massa, F.; Lerer, A.; Bradbury, J.; Chanan, G.; Killeen, T.; Lin, Z.; Gimelshein, N.; Antiga, L.; Desmaison, A.; Kopf, A.; Yang, E.; DeVito, Z.; Raison, M.; Tejani, A.; Chilamkurthy, S.; Steiner, B.; Fang, L.; Bai, J.; and Chintala, S. 2019 · 2019
Later among the works it cites.
Learning Deep Transformer Models for Machine Translation
Wang, Q.; Li, B.; Xiao, T.; Zhu, J.; Li, C.; Wong, D. F.; and Chao, L. S. 2019 · 2019
Later among the works it cites.
Large batch optimization for deep learning: Training bert in 76 minutes
You, Y.; Li, J.; Reddi, S.; Hseu, J.; Kumar, S.; Bhojanapalli, S.; Song, X.; Demmel, J.; Keutzer, K.; and Hsieh, C.-J. 2019 · 2019
Later among the works it cites.
https://github.com/amirgholami/ADAHESSIAN.git
2020 · 2020
Closest in time.