Fetching the paper…
Reading the bibliography…
Pairwise dot product-based attention allows Transformers to exchange information between tokens in an input-dependent way, and is key to their success across diverse applications in language and vision.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
SQuAD: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang · 2016
Earlier work this paper cites.
Revisiting unreasonable effectiveness of data in deep learning era
Chen Sun, Abhinav Shrivastava, Saurabh Singh, and Abhinav Gupta · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Findings of the 2018 conference on machine translation (wmt18)
Ondřej Bojar, Christian Federmann, Mark Fishel, Yvette Graham, Barry Haddow, Matthias Huck, Philipp Koehn, and Christof Monz · 2018
Earlier work this paper cites.
Mostafa Dehghani, Stephan Gouws, Oriol Vinyals, Jakob Uszkoreit, and Łukasz Kaiser · 2018
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional Transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
A call for clarity in reporting bleu scores
Matt Post · 2018
Earlier work this paper cites.
Tensor2tensor for neural machine translation
Ashish Vaswani, Samy Bengio, Eugene Brevdo, Francois Chollet, Aidan N Gomez, Stephan Gouws, Llion Jones, Łukasz Kaiser, Nal Kalchbrenner, Niki Parmar, et al · 2018
Earlier work this paper cites.
A broad-coverage challenge corpus for sentence understanding through inference
Adina Williams, Nikita Nangia, and Samuel Bowman · 2018
Earlier work this paper cites.
Generating long sequences with sparse Transformers
Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever · 2019
Earlier work this paper cites.
What does BERT look at? an analysis of BERT’s attention
Kevin Clark, Urvashi Khandelwal, Omer Levy, and Christopher D Manning · 2019
Cited alongside, same era.
A structural probe for finding syntax in word representations
John Hewitt and Christopher D Manning · 2019
Cited alongside, same era.
Attention is not explanation
Sarthak Jain and Byron C Wallace · 2019
Cited alongside, same era.
Albert: A lite bert for self-supervised learning of language representations
Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut · 2019
Cited alongside, same era.
Are sixteen heads really better than one?
Paul Michel, Omer Levy, and Graham Neubig · 2019
Cited alongside, same era.
Is attention interpretable?
Sofia Serrano and Noah A Smith · 2019
Limits to depth efficiencies of self-attention
Yoav Levine, Noam Wies, Or Sharir, Hofit Bata, and Amnon Shashua · 2020
Later among the works it cites.
Random feature attention
Hao Peng, Nikolaos Pappas, Dani Yogatama, Roy Schwartz, Noah Smith, and Lingpeng Kong · 2020
Later among the works it cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu · 2020
Later among the works it cites.
Fixed encoder self-attention patterns in transformer-based machine translation
Alessandro Raganato, Yves Scherrer, and Jörg Tiedemann · 2020
Later among the works it cites.
A primer in bertology: What we know about how bert works
Anna Rogers, Olga Kovaleva, and Anna Rumshisky · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned
Elena Voita, David Talbot, Fedor Moiseev, Rico Sennrich, and Ivan Titov · 2019
Cited alongside, same era.
Glue: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman · 2019
Cited alongside, same era.
Attention is not not explanation
Sarah Wiegreffe and Yuval Pinter · 2019
Cited alongside, same era.
Rethinking attention with performers
Krzysztof Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Andreea Gane, Tamas Sarlos, Peter Hawkins, Jared Davis, Afroz Mohiuddin, Lukasz Kaiser, et al · 2020
Cited alongside, same era.
Reformer: The efficient Transformer
Nikita Kitaev, Łukasz Kaiser, and Anselm Levskaya · 2020
Cited alongside, same era.
Superglue: a stickier benchmark for general-purpose language understanding systems
Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman
Cited in the paper.
Yi Tay, Mostafa Dehghani, Dara Bahri, and Donald Metzler · 2020
Later among the works it cites.
O(n) connections are expressive enough: Universal approximability of sparse transformers
Chulhee Yun, Yin-Wen Chang, Srinadh Bhojanapalli, Ankit Singh Rawat, Sashank Reddi, and Sanjiv Kumar · 2020
Later among the works it cites.
Big bird: Transformers for longer sequences
Manzil Zaheer, Guru Guruganesh, Avinava Dubey, Joshua Ainslie, Chris Alberti, Santiago Ontanon, Philip Pham, Anirudh Ravula, Qifan Wang, Li Yang, et al · 2020
Later among the works it cites.
Eigen analysis of self-attention and its reconstruction from partial computation
Srinadh Bhojanapalli, Ayan Chakrabarti, Himanshu Jain, Sanjiv Kumar, Michal Lukasik, and Andreas Veit · 2021
Closest in time.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2021
Closest in time.
Long range arena : A benchmark for efficient transformers
Yi Tay, Mostafa Dehghani, Samira Abnar, Yikang Shen, Dara Bahri, Philip Pham, Jinfeng Rao, Liu Yang, Sebastian Ruder, and Donald Metzler · 2021
Closest in time.