Music transformer: Generating music with long-term structure
Cheng-Zhi Anna Huang, Ashish Vaswani, Jakob Uszkoreit, Ian Simon, Curtis Hawthorne, Noam Shazeer, Andrew M Dai, Matthew D Hoffman, Monica Dinculescu, and Douglas Eck. 2018 · 2018
Later among the works it cites.
Generating wikipedia by summarizing long sequences
Original
Peter J Liu, Mohammad Saleh, Etienne Pot, Ben Goodrich, Ryan Sepassi, Lukasz Kaiser, and Noam Shazeer. 2018 · 2018
Later among the works it cites.
Document-level neural machine translation with hierarchical attention networks
Original
Lesly Miculicich, Dhananjay Ram, Nikolaos Pappas, and James Henderson. 2018 · 2018
Later among the works it cites.
Representation learning with contrastive predictive coding
Original
Aaron van den Oord, Yazhe Li, and Oriol Vinyals. 2018 · 2018
Later among the works it cites.
Self-attention with relative position representations
Original
Peter Shaw, Jakob Uszkoreit, and Ashish Vaswani. 2018 · 2018
Later among the works it cites.
Adafactor: Adaptive learning rates with sublinear memory cost
Original
Noam Shazeer and Mitchell Stern. 2018 · 2018
Later among the works it cites.
Bi-directional block self-attention for fast and memory-efficient sequence modeling
Original
Tao Shen, Tianyi Zhou, Guodong Long, Jing Jiang, and Chengqi Zhang. 2018 · 2018
Later among the works it cites.
Glue: A multi-task benchmark and analysis platform for natural language understanding
Original
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman. 2018 · 2018
Later among the works it cites.
Constructing datasets for multi-hop reading comprehension across documents
Johannes Welbl, Pontus Stenetorp, and Sebastian Riedel. 2018 · 2018
Later among the works it cites.
Hotpotqa: A dataset for diverse, explainable multi-hop question answering
Original
Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William W Cohen, Ruslan Salakhutdinov, and Christopher D Manning. 2018 · 2018
Later among the works it cites.
Natural questions: a benchmark for question answering research
Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, et al. 2019 · 2019
Later among the works it cites.
Xlnet: Generalized autoregressive pretraining for language understanding
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Russ R Salakhutdinov, and Quoc V Le. 2019 · 2019
Later among the works it cites.
Large batch optimization for deep learning: Training bert in 76 minutes
Yang You, Jing Li, Sashank Reddi, Jonathan Hseu, Sanjiv Kumar, Srinadh Bhojanapalli, Xiaodan Song, James Demmel, Kurt Keutzer, and Cho-Jui Hsieh. 2019 · 2019
Later among the works it cites.
Leveraging pre-trained checkpoints for sequence generation tasks
Sascha Rothe, Shashi Narayan, and Aliaksei Severyn. 2020 · 2020
Closest in time.