ListOps: A Diagnostic Dataset for Latent Tree Learning
Nangia, N.; and Bowman, S. R. 2018 · 2018
Later among the works it cites.
Deep contextualized word representations
Peters, M. E.; Neumann, M.; Iyyer, M.; Gardner, M.; Clark, C.; Lee, K.; and Zettlemoyer, L. 2018 · 2018
Later among the works it cites.
Improving language understanding with unsupervised learning
Radford, A.; Narasimhan, K.; Salimans, T.; and Sutskever, I. 2018 · 2018
Later among the works it cites.
GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding
Wang, A.; Singh, A.; Michael, J.; Hill, F.; Levy, O.; and Bowman, S. R. 2018 · 2018
Later among the works it cites.
A Broad-Coverage Challenge Corpus for Sentence Understanding through Inference
Williams, A.; Nangia, N.; and Bowman, S. R. 2018 · 2018
Later among the works it cites.
ELECTRA: Pre-training Text Encoders as Discriminators Rather Than Generators
Clark, K.; Luong, M.-T.; Le, Q. V.; and Manning, C. D. 2019 · 2019
Later among the works it cites.
Transformer-XL: Attentive Language Models beyond a Fixed-Length Context
Dai, Z.; Yang, Z.; Yang, Y.; Carbonell, J. G.; Le, Q.; and Salakhutdinov, R. 2019 · 2019
Later among the works it cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Devlin, J.; Chang, M.-W.; Lee, K.; and Toutanova, K. 2019 · 2019
Later among the works it cites.
Gpipe: Efficient training of giant neural networks using pipeline parallelism
Huang, Y.; Cheng, Y.; Bapna, A.; Firat, O.; Chen, D.; Chen, M.; Lee, H.; Ngiam, J.; Le, Q. V.; Wu, Y.; et al. 2019 · 2019
Later among the works it cites.
Reformer: The Efficient Transformer
Kitaev, N.; Kaiser, L.; and Levskaya, A. 2019 · 2019
Later among the works it cites.
Large memory layers with product keys
Lample, G.; Sablayrolles, A.; Ranzato, M.; Denoyer, L.; and Jégou, H. 2019 · 2019
Later among the works it cites.
Set transformer: A framework for attention-based permutation-invariant neural networks
Lee, J.; Lee, Y.; Kim, J.; Kosiorek, A.; Choi, S.; and Teh, Y. W. 2019 · 2019
Later among the works it cites.
Are sixteen heads really better than one?
Michel, P.; Levy, O.; and Neubig, G. 2019 · 2019
Later among the works it cites.
Sampled softmax with random fourier features
Rawat, A. S.; Chen, J.; Yu, F. X. X.; Suresh, A. T.; and Kumar, S. 2019 · 2019
Later among the works it cites.
GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding
Wang, A.; Singh, A.; Michael, J.; Hill, F.; Levy, O.; and Bowman, S. R. 2019 · 2019
Later among the works it cites.
Q8BERT: Quantized 8bit BERT
Zafrir, O.; Boudoukh, G.; Izsak, P.; and Wasserblat, M. 2019 · 2019
Later among the works it cites.
Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention
Katharopoulos, A.; Vyas, A.; Pappas, N.; and Fleuret, F. 2020 · 2020
Later among the works it cites.
ALBERT: A lite BERT for self-supervised learning of language representations
Lan, Z.; Chen, M.; Goodman, S.; Gimpel, K.; Sharma, P.; and Soricut, R. 2020 · 2020
Later among the works it cites.
BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension
Lewis, M.; Liu, Y.; Goyal, N.; Ghazvininejad, M.; Mohamed, A.; Levy, O.; Stoyanov, V.; and Zettlemoyer, L. 2020 · 2020
Later among the works it cites.
Fast transformers with clustered attention
Vyas, A.; Katharopoulos, A.; and Fleuret, F. 2020 · 2020
Later among the works it cites.
LambdaNetworks: Modeling long-range Interactions without Attention
Bello, I. 2021 · 2021
Closest in time.