Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, L. Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Modularized morphing of neural networks
Original
Tao Wei, Changhu Wang, and Chang Wen Chen. 2017 · 2017
Cited alongside, same era.
Multi-level residual networks from dynamical systems view
Bo Chang, Lili Meng, Eldad Haber, Frederick Tung, and David Begert. 2018 · 2018
Cited alongside, same era.
Empower sequence labeling with task-aware neural language model
Liyuan Liu, Jingbo Shang, Xiang Ren, Frank Fangzheng Xu, Huan Gui, Jian Peng, and Jiawei Han. 2018 · 2018
Cited alongside, same era.
Deep contextualized word representations
Matthew E Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018 · 2018
Cited alongside, same era.
Improving language understanding by generative pre-training
A. Radford. 2018 · 2018
Cited alongside, same era.
Know what you don’t know: Unanswerable questions for squad
Pranav Rajpurkar, Robin Jia, and Percy Liang. 2018 · 2018
Cited alongside, same era.
Glue: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman. 2018 · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
On the variance of the adaptive learning rate and beyond
Liyuan Liu, Haoming Jiang, Pengcheng He, W. Chen, Xiaodong Liu, Jianfeng Gao, and J. Han. 2020a
Cited in the paper.
Understanding the difficulty of training transformers
Liyuan Liu, X. Liu, Jianfeng Gao, Weizhu Chen, and J. Han. 2020b
Cited in the paper.