Adaptive attention span in transformers
Sainbayar Sukhbaatar, Edouard Grave, Piotr Bojanowski, and Armand Joulin. 2019 · 2019
Later among the works it cites.
Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned
Elena Voita, David Talbot, Fedor Moiseev, Rico Sennrich, and Ivan Titov. 2019 · 2019
Later among the works it cites.
Pubtator central: automated concept annotation for biomedical full text articles
Chih-Hsuan Wei, Alexis Allot, Robert Leaman, and Zhiyong Lu. 2019 · 2019
Later among the works it cites.
Extractive summarization of long documents by combining global and local context
Wen Xiao and Giuseppe Carenini. 2019 · 2019
Later among the works it cites.
Longformer: The long-document transformer
Original
Iz Beltagy, Matthew E. Peters, and Arman Cohan. 2020 · 2020
Later among the works it cites.
Rethinking attention with performers
Original
Krzysztof Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Andreea Gane, Tamas Sarlos, Peter Hawkins, Jared Davis, Afroz Mohiuddin, Lukasz Kaiser, David Belanger, Lucy Colwell, and Adrian Weller. 2020 · 2020
Later among the works it cites.
FEQA: A question answering evaluation framework for faithfulness assessment in abstractive summarization
Esin Durmus, He He, and Mona Diab. 2020 · 2020
Later among the works it cites.
A divide-and-conquer approach to the summarization of long documents
Alexios Gidiotis and Grigorios Tsoumakas. 2020 · 2020
Later among the works it cites.
Transformers are rnns: Fast autoregressive transformers with linear attention
Original
Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, and François Fleuret. 2020 · 2020
Later among the works it cites.
Reformer: The efficient transformer
Nikita Kitaev, Lukasz Kaiser, and Anselm Levskaya. 2020 · 2020
Later among the works it cites.
Evaluating the factual consistency of abstractive text summarization
Wojciech Kryscinski, Bryan McCann, Caiming Xiong, and Richard Socher. 2020 · 2020
Later among the works it cites.
Biobert: a pre-trained biomedical language representation model for biomedical text mining
Jinhyuk Lee, Wonjin Yoon, Sungdong Kim, Donghyeon Kim, Sunkyu Kim, Chan Ho So, and Jaewoo Kang. 2020 · 2020
Later among the works it cites.
BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020 · 2020
Later among the works it cites.
A joint neural model for information extraction with global features
Ying Lin, Heng Ji, Fei Huang, and Lingfei Wu. 2020 · 2020
Later among the works it cites.
On extractive and abstractive neural document summarization with transformer language models
Jonathan Pilault, Raymond Li, Sandeep Subramanian, and Chris Pal. 2020 · 2020
Later among the works it cites.
Seal: Segment-wise extractive-abstractive long-form text summarization
Original
Yao Zhao, Mohammad Saleh, and Peter J. Liu. 2020 · 2020
Later among the works it cites.