Fetching the paper…
Reading the bibliography…
Following the success of dot-product attention in Transformers, numerous approximations have been recently proposed to address its quadratic complexity with respect to the input length.
Learning word vectors for sentiment analysis
A. L. Maas, R. E. Daly, P. T. Pham, D. Huang, A. Y. Ng, and C. Potts · 2011
Earlier work this paper cites.
The ACL anthology network corpus
D. R. Radev, P. Muthukrishnan, V. Qazvinian, and A. Abu-Jbara · 2013
Earlier work this paper cites.
MCTest: A challenge dataset for the open-domain machine comprehension of text
M. Richardson, C. J. Burges, and E. Renshaw · 2013
Earlier work this paper cites.
Elementary school science and math tests as a driver for AI: take the aristo challenge!
P. Clark · 2015
Earlier work this paper cites.
Training deep nets with sublinear memory cost
T. Chen, B. Xu, C. Zhang, and C. Guestrin · 2016
Earlier work this paper cites.
SQuAD: 100,000+ questions for machine comprehension of text
P. Rajpurkar, J. Zhang, K. Lopyrev, and P. Liang · 2016
Earlier work this paper cites.
The reversible residual network: Backpropagation without storing activations
A. N. Gomez, M. Ren, R. Urtasun, and R. B. Grosse · 2017
Earlier work this paper cites.
RACE: Large-scale ReAding comprehension dataset from examinations
G. Lai, Q. Xie, H. Liu, Y. Yang, and E. Hovy · 2017
Earlier work this paper cites.
Pointer sentinel mixture models
S. Merity, C. Xiong, J. Bradbury, and R. Socher · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Think you have solved question answering? try ARC, the AI2 reasoning challenge
P. Clark, I. Cowhey, O. Etzioni, T. Khot, A. Sabharwal, C. Schoenick, and O. Tafjord · 2018
Earlier work this paper cites.
Understanding back-translation at scale
S. Edunov, M. Ott, M. Auli, and D. Grangier · 2018
Earlier work this paper cites.
The NarrativeQA reading comprehension challenge
T. Kočiský, J. Schwarz, P. Blunsom, C. Dyer, K. M. Hermann, G. Melis, and E. Grefenstette · 2018
Earlier work this paper cites.
Can a suit of armor conduct electricity? a new dataset for open book question answering
T. Mihaylov, P. Clark, T. Khot, and A. Sabharwal · 2018
Earlier work this paper cites.
ListOps: A diagnostic dataset for latent tree learning
N. Nangia and S. Bowman · 2018
Earlier work this paper cites.
Deep contextualized word representations
M. Peters, M. Neumann, M. Iyyer, M. Gardner, C. Clark, K. Lee, and L. Zettlemoyer · 2018
Earlier work this paper cites.
Know what you don’t know: Unanswerable questions for SQuAD
P. Rajpurkar, R. Jia, and P. Liang · 2018
Earlier work this paper cites.
Generating long sequences with sparse transformers
R. Child, S. Gray, A. Radford, and I. Sutskever · 2019
Earlier work this paper cites.
BoolQ: Exploring the surprising difficulty of natural yes/no questions
C. Clark, K. Lee, M.-W. Chang, T. Kwiatkowski, M. Collins, and K. Toutanova · 2019
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2019
Cited alongside, same era.
A structural probe for finding syntax in word representations
J. Hewitt and C. D. Manning · 2019
Cited alongside, same era.
Reasoning over paragraph effects in situations
K. Lin, O. Tafjord, P. Clark, and M. Gardner · 2019
Cited alongside, same era.
Language models as knowledge bases?
F. Petroni, T. Rocktäschel, S. Riedel, P. Lewis, A. Bakhtin, Y. Wu, and A. Miller · 2019
Cited alongside, same era.
Augmenting self-attention with persistent memory
S. Sukhbaatar, E. Grave, G. Lample, H. Jégou, and A. Joulin · 2019
Cited alongside, same era.
CommonsenseQA: A question answering challenge targeting commonsense knowledge
Reformer: The efficient transformer
N. Kitaev, L. Kaiser, and A. Levskaya · 2020
Later among the works it cites.
Blockwise self-attention for long document understanding
J. Qiu, H. Ma, O. Levy, W.-t. Yih, S. Wang, and J. Tang · 2020
Later among the works it cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu · 2020
Later among the works it cites.
Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters
J. Rasley, S. Rajbhandari, O. Ruwase, and Y. He · 2020
Later among the works it cites.
How much knowledge can you pack into the parameters of a language model?
A. Roberts, C. Raffel, and N. Shazeer · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Talmor, J. Herzig, N. Lourie, and J. Berant · 2019
Cited alongside, same era.
Triton: an intermediate language and compiler for tiled neural network computations
P. Tillet, H. Kung, and D. Cox · 2019
Cited alongside, same era.
Huggingface’s transformers: State-of-the-art natural language processing
T. Wolf, L. Debut, V. Sanh, J. Chaumond, C. Delangue, A. Moi, P. Cistac, T. Rault, R. Louf, M. Funtowicz, and J. Brew · 2019
Cited alongside, same era.
BP-Transformer: Modelling long-range context via binary partitioning
Z. Ye, Q. Guo, Q. Gan, X. Qiu, and Z. Zhang · 2019
Cited alongside, same era.
On the power and limitations of random features for understanding neural networks
G. Yehudai and O. Shamir · 2019
Cited alongside, same era.
Explicit sparse transformer: Concentrated attention through explicit selection
G. Zhao, J. Lin, Z. Zhang, X. Ren, Q. Su, and X. Sun · 2019
Cited alongside, same era.
Longformer: The long-document transformer
I. Beltagy, M. E. Peters, and A. Cohan · 2020
Cited alongside, same era.
N. M. Shazeer · 2020
Later among the works it cites.
Sparse sinkhorn attention
Y. Tay, D. Bahri, L. Yang, D. Metzler, and D. Juan · 2020
Later among the works it cites.
Efficient transformers: A survey
Y. Tay, M. Dehghani, D. Bahri, and D. Metzler · 2020
Later among the works it cites.
Fast transformers with clustered attention
A. Vyas, A. Katharopoulos, and F. Fleuret · 2020
Later among the works it cites.
Linformer: Self-attention with linear complexity
S. Wang, B. Z. Li, M. Khabsa, H. Fang, and H. Ma · 2020
Later among the works it cites.
Big Bird: Transformers for longer sequences
M. Zaheer, G. Guruganesh, K. A. Dubey, J. Ainslie, C. Alberti, S. Ontanon, P. Pham, A. Ravula, Q. Wang, L. Yang, et al · 2020
Later among the works it cites.
Lambdanetworks: Modeling long-range interactions without attention
I. Bello · 2021
Closest in time.
Rethinking attention with performers
K. M. Choromanski, V. Likhosherstov, D. Dohan, X. Song, A. Gane, T. Sarlos, P. Hawkins, J. Q. Davis, A. Mohiuddin, L. Kaiser, D. B. Belanger, L. J. Colwell, and A. Weller · 2021
Closest in time.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
W. Fedus, B. Zoph, and N. M. Shazeer · 2021
Closest in time.
Value-aware approximate attention
A. Gupta and J. Berant · 2021
Closest in time.
Random feature attention
H. Peng, N. Pappas, D. Yogatama, R. Schwartz, N. Smith, and L. Kong · 2021
Closest in time.
Efficient content-based sparse attention with routing transformers
A. Roy, M. Saffar, A. Vaswani, and D. Grangier · 2021
Closest in time.
Long Range Arena : A benchmark for efficient transformers
Y. Tay, M. Dehghani, S. Abnar, Y. Shen, D. Bahri, P. Pham, J. Rao, L. Yang, S. Ruder, and D. Metzler · 2021
Closest in time.