Fetching the paper…
Reading the bibliography…
Transformer models achieve state-of-the-art performance on a wide range of NLP tasks.
Learning word vectors for sentiment analysis
Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, Andrew Y. Ng, and Christopher Potts · 2011
Earlier work this paper cites.
Adam: A method for stochastic optimization, 2014
Diederik P. Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Practical and optimal LSH for angular distance
Alexandr Andoni, Piotr Indyk, Thijs Laarhoven, Ilya P. Razenshteyn, and Ludwig Schmidt · 2015
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Efficient attention using a fixed-size memory representation
Denny Britz, Melody Y Guan, and Minh-Thang Luong · 2017
Earlier work this paper cites.
Chung-Cheng Chiu and Colin Raffel · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
Self-attention with relative position representations
Peter Shaw, Jakob Uszkoreit, and Ashish Vaswani · 2018
Earlier work this paper cites.
Bi-directional block self-attention for fast and memory-efficient sequence modeling
Tao Shen, Tianyi Zhou, Guodong Long, Jing Jiang, and Chengqi Zhang · 2018
Earlier work this paper cites.
Factorized attention: Self-attention with linear complexities
Zhuoran Shen, Mingyuan Zhang, Shuai Yi, Junjie Yan, and Haiyu Zhao · 2018
Earlier work this paper cites.
A discourse-aware attention model for abstractive summarization of long documents
Arman Cohan, Franck Dernoncourt, Doo Soon Kim, Trung Bui, Seokhwan Kim, Walter Chang, and Nazli Goharian · 2018
Earlier work this paper cites.
Transformer-xl: Attentive language models beyond a fixed-length context
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc V Le, and Ruslan Salakhutdinov · 2019
Earlier work this paper cites.
Compressive transformers for long-range sequence modelling
Jack W Rae, Anna Potapenko, Siddhant M Jayakumar, and Timothy P Lillicrap · 2019
Earlier work this paper cites.
Generating long sequences with sparse transformers
Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever · 2019
Earlier work this paper cites.
What does bert look at? an analysis of bert’s attention
Kevin Clark, Urvashi Khandelwal, Omer Levy, and Christopher D Manning · 2019
Cited alongside, same era.
Pay less attention with lightweight and dynamic convolutions
Felix Wu, Angela Fan, Alexei Baevski, Yann N Dauphin, and Michael Auli · 2019
Cited alongside, same era.
Xingxing Zhang, Furu Wei, and Ming Zhou · 2019
Cited alongside, same era.
Qipeng Guo, Xipeng Qiu, Pengfei Liu, Yunfan Shao, Xiangyang Xue, and Zheng Zhang · 2019
Cited alongside, same era.
Fixed encoder self-attention patterns in transformer-based machine translation
Alessandro Raganato, Yves Scherrer, and Jörg Tiedemann · 2020
Later among the works it cites.
Etc: Encoding long and structured inputs in transformers
Joshua Ainslie, Santiago Ontanon, Chris Alberti, Vaclav Cvicek, Zachary Fisher, Philip Pham, Anirudh Ravula, Sumit Sanghai, Qifan Wang, and Li Yang · 2020
Later among the works it cites.
Transformers are rnns: Fast autoregressive transformers with linear attention
Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, and François Fleuret · 2020
Later among the works it cites.
Legal-bert: The muppets straight out of law school
Ilias Chalkidis, Manos Fergadiotis, Prodromos Malakasiotis, Nikolaos Aletras, and Ion Androutsopoulos · 2020
Later among the works it cites.
Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 2019
Cited alongside, same era.
Long document classification from local word glimpses via recurrent attention learning
Jun He, Liqun Wang, Liu Liu, Jiao Feng, and Hao Wu · 2019
Cited alongside, same era.
Bigpatent: A large-scale dataset for abstractive and coherent summarization
Eva Sharma, Chen Li, and Lu Wang · 2019
Cited alongside, same era.
Multi-news: a large-scale multi-document summarization dataset and abstractive hierarchical model, 2019
Alexander R. Fabbri, Irene Li, Tianwei She, Suyi Li, and Dragomir R. Radev · 2019
Cited alongside, same era.
Pegasus: Pre-training with extracted gap-sentences for abstractive summarization, 2019
Jingqing Zhang, Yao Zhao, Mohammad Saleh, and Peter J. Liu · 2019
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, Peter J Liu, et al · 2020
Cited alongside, same era.
Longformer: The long-document transformer
Iz Beltagy, Matthew E. Peters, and Arman Cohan · 2020
Cited alongside, same era.
Linformer: Self-attention with linear complexity
Sinong Wang, Belinda Z. Li, Madian Khabsa, Han Fang, and Hao Ma · 2020
Cited alongside, same era.
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer · 2020
Later among the works it cites.
Train short, test long: Attention with linear biases enables input length extrapolation
Ofir Press, Noah A. Smith, and Mike Lewis · 2021
Later among the works it cites.
Rethinking attention with performers
Krzysztof Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Andreea Gane, Tamas Sarlos, Peter Hawkins, Jared Davis, Afroz Mohiuddin, Lukasz Kaiser, David Belanger, Lucy Colwell, and Adrian Weller · 2021
Later among the works it cites.
Efficient attentions for long document summarization
Luyang Huang, Shuyang Cao, Nikolaus Parulian, Heng Ji, and Lu Wang · 2021
Later among the works it cites.
Mediasum: A large-scale media interview dataset for dialogue summarization
Chenguang Zhu, Yang Liu, Jie Mei, and Michael Zeng · 2021
Later among the works it cites.
Hierarchical learning for generation with long source sequences
Tobias Rohde, Xiaoxia Wu, and Yinhan Liu · 2021
Later among the works it cites.
Longt5: Efficient text-to-text transformer for long sequences
Mandy Guo, Joshua Ainslie, David C. Uthus, Santiago Ontañón, Jianmo Ni, Yun-Hsuan Sung, and Yinfei Yang · 2021
Later among the works it cites.
Lexglue: A benchmark dataset for legal language understanding in english
Ilias Chalkidis, Abhik Jana, Dirk Hartung, Michael Bommarito, Ion Androutsopoulos, Daniel Martin Katz, and Nikolaos Aletras · 2022
Closest in time.
Primera: Pyramid-based masked sentence pre-training for multi-document summarization
Wen Xiao, Iz Beltagy, Giuseppe Carenini, and Arman Cohan · 2022
Closest in time.