Fetching the paper…
Reading the bibliography…
Transformer-based models are unable to process long sequences due to their self-attention operation, which scales quadratically with the sequence length.
Pay less attention with lightweight and dynamic convolutions
Felix Wu, Angela Fan, Alexei Baevski, Yann Dauphin, and Michael Auli. 2019 · 1901
Earlier work this paper cites.
Generating long sequences with sparse transformers
Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever. 2019 · 1904
Earlier work this paper cites.
Unsupervised data augmentation for consistency training
Qizhe Xie, Zihang Dai, Eduard H. Hovy, Minh-Thang Luong, and Quoc V. Le. 2019 · 1904
Earlier work this paper cites.
What does bert look at? an analysis of bert’s attention
Kevin Clark, Urvashi Khandelwal, Omer Levy, and Christopher D. Manning. 2019 · 1906
Earlier work this paper cites.
RoBERTa: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Span selection pre-training for question answering
Michael Glaß, Alfio Massimiliano Gliozzo, Rishav Chakravarti, Anthony Ferritto, Lin Pan, Gaudani Bhargav, Dinesh Garg, and Avirup Sil. 2019 · 1909
Earlier work this paper cites.
Multi-hop question answering via reasoning chains
Jifan Chen, Shih-Ting Lin, and Greg Durrett. 2019 · 1910
Earlier work this paper cites.
Blockwise self-attention for long document understanding
Jiezhong Qiu, Hao Ma, Omer Levy, Scott Yih, Sinong Wang, and Jie Tang. 2019 · 1911
Earlier work this paper cites.
Select, answer and explain: Interpretable multi-hop reading comprehension over multiple documents
Ming Tu, Kevin Huang, Guangtao Wang, Jing Huang, Xiaodong He, and Bufang Zhou. 2019 · 1911
Earlier work this paper cites.
BP-Transformer: Modelling long-range context via binary partitioning
Zihao Ye, Qipeng Guo, Quan Gan, Xipeng Qiu, and Zheng Zhang. 2019 · 1911
Earlier work this paper cites.
On layer normalization in the transformer architecture
Ruibin Xiong, Yunchang Yang, Di He, Kai Zheng, Shu xin Zheng, Chen Xing, Huishuai Zhang, Yanyan Lan, Li-Wei Wang, and Tie-Yan Liu. 2020 · 2002
Earlier work this paper cites.
Efficient content-based sparse attention with routing transformers
Aurko Roy, Mohammad Saffar, Ashish Vaswani, and David Grangier. 2020 · 2003
Earlier work this paper cites.
A divide-and-conquer approach to the summarization of academic articles
Alexios Gidiotis and Grigorios Tsoumakas. 2020 · 2004
Earlier work this paper cites.
A simple yet strong pipeline for HotpotQA
Dirk Groeneveld, Tushar Khot, Mausam, and Ashish Sabhwaral. 2020 · 2004
Earlier work this paper cites.
Is graph structure necessary for multi-hop reasoning?
Nan Shao, Yiming Cui, Ting Liu, Shijin Wang, and Guoping Hu. 2020 · 2004
Earlier work this paper cites.
Gmat: Global memory augmentation for transformers
Ankit Gupta and Jonathan Berant. 2020 · 2006
Earlier work this paper cites.
Big bird: Transformers for longer sequences
Manzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie, C. Alberti, S. Ontañón, Philip Pham, Anirudh Ravula, Qifan Wang, L. Yang, and A. Ahmed. 2020 · 2007
Earlier work this paper cites.
Large text compression benchmark
Matt Mahoney. 2009 · 2009
Earlier work this paper cites.
Learning word vectors for sentiment analysis
Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, Andrew Y. Ng, and Christopher Potts. 2011 · 2011
Earlier work this paper cites.
CoNLL-2012 shared task: Modeling multilingual unrestricted coreference in OntoNotes
Sameer Pradhan, Alessandro Moschitti, Nianwen Xue, Olga Uryupina, and Yuchen Zhang. 2012 · 2012
Earlier work this paper cites.
Sequence to sequence learning with neural networks
Ilya Sutskever, Oriol Vinyals, and Quoc V. Le. 2014 · 2014
Cited alongside, same era.
Semi-supervised sequence learning
Andrew M Dai and Quoc V Le. 2015 · 2015
Cited alongside, same era.
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books
Yukun Zhu, Ryan Kiros, Richard S. Zemel, Ruslan Salakhutdinov, Raquel Urtasun, Antonio Torralba, and Sanja Fidler. 2015 · 2015
Cited alongside, same era.
Training deep nets with sublinear memory cost
Tianqi Chen, Bing Xu, Chiyuan Zhang, and Carlos Guestrin. 2016 · 2016
Cited alongside, same era.
Wavenet: A generative model for raw audio
Aäron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew W. Senior, and Koray Kavukcuoglu. 2016 · 2016
Cited alongside, same era.
Transformer-XL: Attentive language models beyond a fixed-length context
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime G. Carbonell, Quoc V. Le, and Ruslan Salakhutdinov. 2019 · 2019
Later among the works it cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Later among the works it cites.
MRQA 2019 shared task: Evaluating generalization in reading comprehension
Adam Fisch, Alon Talmor, Robin Jia, Minjoon Seo, Eunsol Choi, and Danqi Chen. 2019 · 2019
Later among the works it cites.
BERT for coreference resolution: Baselines and analysis
Mandar Joshi, Omer Levy, Luke Zettlemoyer, and Daniel Weld. 2019 · 2019
Later among the works it cites.
SemEval-2019 task 4: Hyperpartisan news detection
Johannes Kiesel, Maria Mestre, Rishabh Shukla, Emmanuel Vincent, Payam Adineh, David Corney, Benno Stein, and Martin Potthast. 2019 · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Danqi Chen, Adam Fisch, Jason Weston, and Antoine Bordes. 2017 · 2017
Cited alongside, same era.
Simple and effective multi-paragraph reading comprehension
Christopher Clark and Matt Gardner. 2017 · 2017
Cited alongside, same era.
Gpu kernels for block-sparse weights
Scott Gray, Alec Radford, and Diederik P. Kingma. 2017 · 2017
Cited alongside, same era.
TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension
Mandar Joshi, Eunsol Choi, Daniel S. Weld, and Luke Zettlemoyer. 2017 · 2017
Cited alongside, same era.
Semi-supervised classification with graph convolutional networks
Thomas N Kipf and Max Welling. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Character-level language modeling with deeper self-attention
Rami Al-Rfou, Dokook Choe, Noah Constant, Mandy Guo, and Llion Jones. 2018 · 2018
Cited alongside, same era.
Olga V. Kovaleva, Alexey Romanov, Anna Rogers, and Anna Rumshisky. 2019 · 2019
Later among the works it cites.
fairseq: A fast, extensible toolkit for sequence modeling
Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli. 2019 · 2019
Later among the works it cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Later among the works it cites.
Adaptive attention span in transformers
Sainbayar Sukhbaatar, Edouard Grave, Piotr Bojanowski, and Armand Joulin. 2019 · 2019
Later among the works it cites.
Defending against neural fake news
Rowan Zellers, Ari Holtzman, Hannah Rashkin, Yonatan Bisk, Ali Farhadi, Franziska Roesner, and Yejin Choi. 2019 · 2019
Later among the works it cites.
ETC: Encoding long and structured inputs in transformers
Joshua Ainslie, Santiago Ontanon, Chris Alberti, Vaclav Cvicek, Zachary Fisher, Philip Pham, Anirudh Ravula, Sumit Sanghai, Qifan Wang, and Li Yang. 2020 · 2020
Closest in time.
Hierarchical graph network for multi-hop question answering
Yuwei Fang, Siqi Sun, Zhe Gan, Rohit Pillai, Shuohang Wang, and Jingjing Liu. 2020 · 2020
Closest in time.
Reformer: The efficient transformer
Nikita Kitaev, Lukasz Kaiser, and Anselm Levskaya. 2020 · 2020
Closest in time.
BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020 · 2020
Closest in time.
Compressive transformers for long-range sequence modelling
Jack W. Rae, Anna Potapenko, Siddhant M. Jayakumar, and Timothy P. Lillicrap. 2020 · 2020
Closest in time.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, W. Li, and Peter J. Liu. 2020 · 2020
Closest in time.
On extractive and abstractive neural document summarization with transformer language models
Sandeep Subramanian, Raymond Li, Jonathan Pilault, and C. Pal. 2020 · 2020
Closest in time.
Graph sequential network for reasoning over sequences
Ming Tu, Jinke Huang, Xiaodong He, and Bowen Zhou. 2020 · 2020
Closest in time.
Pegasus: Pre-training with extracted gap-sentences for abstractive summarization
Jingqing Zhang, Yao Zhao, Mohammad Saleh, and Peter J Liu. 2020 · 2020
Closest in time.