Sparse sinkhorn attention
Original
Yi Tay, Dara Bahri, Liu Yang, Donald Metzler, and Da-Cheng Juan · 2002
Earlier work this paper cites.
Beyond 512 tokens: Siamese multi-depth transformer-based hierarchical encoder for document matching
Original
Liu Yang, Mingyang Zhang, Cheng Li, Michael Bendersky, and Marc Najork · 2004
Earlier work this paper cites.
Synthesizer: Rethinking self-attention in transformer models
Original
Yi Tay, Dara Bahri, Donald Metzler, Da-Cheng Juan, Zhe Zhao, and Che Zheng · 2005
Earlier work this paper cites.
Masked language modeling for proteins via linearly scalable long-context transformers
Original
Krzysztof Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Jared Davis, Tamas Sarlos, David Belanger, Lucy Colwell, and Adrian Weller · 2006
Earlier work this paper cites.
Rethinking attention with performers
Original
Krzysztof Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Andreea Gane, Tamas Sarlos, Peter Hawkins, Jared Davis, Afroz Mohiuddin, Lukasz Kaiser, et al · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky · 2009
Earlier work this paper cites.
Efficient transformers: A survey
Original
Yi Tay, Mostafa Dehghani, Dara Bahri, and Donald Metzler · 2009
Earlier work this paper cites.
Parallel and serial grouping of image elements in visual perception
R. Houtkamp and P. R. Roelfsema · 2010
Earlier work this paper cites.
Learning word vectors for sentiment analysis
Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, Andrew Y. Ng, and Christopher Potts · 2011
Earlier work this paper cites.
The acl anthology network corpus
Dragomir R Radev, Pradeep Muthukrishnan, Vahed Qazvinian, and Amjad Abu-Jbara · 2013
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Y Wu, Jason Chuang, Christopher D Manning, Andrew Y Ng, and Christopher Potts · 2013
Earlier work this paper cites.
A deep relevance matching model for ad-hoc retrieval
Jiafeng Guo, Yixing Fan, Qingyao Ai, and W Bruce Croft · 2016
Earlier work this paper cites.
The LAMBADA dataset: Word prediction requiring a broad discourse context
Denis Paperno, Angeliki Lazaridou, Quan Ngoc Pham, Raffaella Bernardi, Sandro Pezzelle, Marco Baroni, Gemma Boleda, and Raquel Fern · 2016
Earlier work this paper cites.
Frustratingly Short Attention Spans in Neural Language Modeling
Michał Daniluk, Tim Rockt, Johannes Welbl, and Sebastian Riedel · 2017
Earlier work this paper cites.
TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension
Mandar Joshi, Eunsol Choi, Daniel S Weld, Luke Zettlemoyer, and Paul G Allen · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Original
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.