Fetching the paper…
Reading the bibliography…
We propose a novel method to sparsify attention in the Transformer model by learning to select the most-informative token representations during the training process, thus focusing on the task-specific parts of an input.
Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime G. Carbonell, Quoc V. Le, and Ruslan Salakhutdinov. 2019 · 1901
Earlier work this paper cites.
Types of document knowledge: From structures to strategies
Peter B. Mosenthal and Irwin S. Kirsch. 1992 · 1992
Earlier work this paper cites.
Understanding the strategies of document literacy and their conditions of use
Peter B Mosenthal. 1996 · 1996
Earlier work this paper cites.
Reformer: The Efficient Transformer
Nikita Kitaev, Lukasz Kaiser, and Anselm Levskaya. 2020 · 2001
Earlier work this paper cites.
Longformer: The Long-Document Transformer
Iz Beltagy, Matthew E. Peters, and Arman Cohan. 2020 · 2004
Earlier work this paper cites.
Making do with what we have: use your bootstraps
Guillaume Calmettes, Gordon B. Drummond, and Sarah L. Vowler. 2012 · 2012
Earlier work this paper cites.
Coarse-to-fine question answering for long documents
Eunsol Choi, Daniel Hewlett, Jakob Uszkoreit, Illia Polosukhin, Alexandre Lacoste, and Jonathan Berant. 2017 · 2017
Earlier work this paper cites.
Differentiable scheduled sampling for credit assignment
Kartik Goyal, Chris Dyer, and Taylor Berg-Kirkpatrick. 2017 · 2017
Earlier work this paper cites.
Temporal dynamics of eye-tracking and eeg during reading and relevance decisions
Jacek Gwizdka, Rahilsadat Hosseini, Michael Cole, and Shouyi Wang. 2017 · 2017
Earlier work this paper cites.
The narrativeqa reading comprehension challenge
Tomás Kociský, Jonathan Schwarz, Phil Blunsom, Chris Dyer, Karl Moritz Hermann, Gábor Melis, and Edward Grefenstette. 2017 · 2017
Earlier work this paper cites.
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Fast abstractive summarization with reinforce-selected sentence rewriting
Yen-Chun Chen and Mohit Bansal. 2018 · 2018
Earlier work this paper cites.
A discourse-aware attention model for abstractive summarization of long documents
Arman Cohan, Franck Dernoncourt, Doo Soon Kim, Trung Bui, Seokhwan Kim, Walter Chang, and Nazli Goharian. 2018 · 2018
Earlier work this paper cites.
Bottom-up abstractive summarization
Sebastian Gehrmann, Yuntian Deng, and Alexander M. Rush. 2018 · 2018
Earlier work this paper cites.
A continuous relaxation of beam search for end-to-end training of neural sequence models
Kartik Goyal, Graham Neubig, Chris Dyer, and Taylor Berg-Kirkpatrick. 2018 · 2018
Cited alongside, same era.
A unified model for extractive and abstractive summarization using inconsistency loss
Wan-Ting Hsu, Chieh-Kai Lin, Ming-Ying Lee, Kerui Min, Jing Tang, and Min Sun. 2018 · 2018
Cited alongside, same era.
Neural nearest neighbors networks
Tobias Plötz and Stefan Roth. 2018 · 2018
Cited alongside, same era.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. 2018 · 2018
Cited alongside, same era.
Generating long sequences with sparse transformers
Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever. 2019 · 2019
PoWER-BERT: Accelerating BERT inference via progressive word-vector elimination
Saurabh Goyal, Anamitra Roy Choudhury, Saurabh Raje, Venkatesan Chakaravarthy, Yogish Sabharwal, and Ashish Verma. 2020 · 2020
Closest in time.
Successive Halving Top-k Operator
Michał Pietruszka, Łukasz Borchmann, and Filip Graliǹski. 2020 · 2020
Closest in time.
Do transformers need deep long-range memory?
Jack Rae and Ali Razavi. 2020 · 2020
Closest in time.
Efficient content-based sparse attention with routing transformers
Aurko Roy, Mohammad Saffar, Ashish Vaswani, and David Grangier. 2020 · 2020
Closest in time.
Sparse Sinkhorn Attention
Yi Tay, Dara Bahri, Liu Yang, Donald Metzler, and Da-Cheng Juan. 2020 · 2020
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
RoBERTa: A Robustly Optimized BERT Pretraining Approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 2019
Cited alongside, same era.
fairseq: A fast, extensible toolkit for sequence modeling
Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli. 2019 · 2019
Cited alongside, same era.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. 2019 · 2019
Cited alongside, same era.
On extractive and abstractive neural document summarization with transformer language models
Sandeep Subramanian, Raymond Li, Jonathan Pilault, and Christopher Pal. 2019 · 2019
Cited alongside, same era.
Reparameterizable subset sampling via continuous relaxations
Sang Michael Xie and Stefano Ermon. 2019 · 2019
Cited alongside, same era.
Funnel-transformer: Filtering out sequential redundancy for efficient language processing
Zihang Dai, Guokun Lai, Yiming Yang, and Quoc V. Le. 2020 · 2020
Cited alongside, same era.
Sinong Wang, Belinda Z. Li, Madian Khabsa, Han Fang, and Hao Ma. 2020 · 2020
Closest in time.
Differentiable top-k operator with optimal transport
Yujia Xie, Hanjun Dai, Minshuo Chen, Bo Dai, Tuo Zhao, Hongyuan Zha, Wei Wei, and Tomas Pfister. 2020 · 2020
Closest in time.
Rethinking attention with performers
Krzysztof Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Andreea Gane, Tamas Sarlos, Peter Hawkins, Jared Davis, Afroz Mohiuddin, Lukasz Kaiser, David Belanger, Lucy Colwell, and Adrian Weller. 2021 · 2021
Closest in time.
Multiscale vision transformers
Haoqi Fan, Bo Xiong, Karttikeya Mangalam, Yanghao Li, Zhicheng Yan, Jitendra Malik, and Christoph Feichtenhofer. 2021 · 2021
Closest in time.
Hierarchical learning for generation with long source sequences
Tobias Rohde, Xiaoxia Wu, and Yinhan Liu. 2021 · 2021
Closest in time.
Efficient attention: Attention with linear complexities
Zhuoran Shen, Mingyuan Zhang, Haiyu Zhao, Shuai Yi, and Hongsheng Li. 2021 · 2021
Closest in time.
Doc2dict: Information extraction as text generation
Benjamin Townsend, Eamon Ito-Fisher, Lily Zhang, and Madison May. 2021 · 2021
Closest in time.
Poolingformer: Long document modeling with pooling attention
Hang Zhang, Yeyun Gong, Yelong Shen, Weisheng Li, Jiancheng Lv, Nan Duan, and Weizhu Chen. 2021 · 2021
Closest in time.