Fetching the paper…
Reading the bibliography…
The dot product self-attention is known to be central and indispensable to state-of-the-art Transformer models.
Learning word vectors for sentiment analysis
Maas, A. L., Daly, R. E., Pham, P. T., Huang, D., Ng, A. Y., and Potts, C · 2011
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Bahdanau, D., Cho, K., and Bengio, Y · 2014
Earlier work this paper cites.
Graves, A., Wayne, G., and Danihelka, I · 2014
Earlier work this paper cites.
Weston, J., Chopra, S., and Bordes, A · 2014
Earlier work this paper cites.
Effective approaches to attention-based neural machine translation
Luong, M.-T., Pham, H., and Manning, C. D · 2015
Earlier work this paper cites.
Character-level convolutional networks for text classification
Zhang, X., Zhao, J., and LeCun, Y · 2015
Earlier work this paper cites.
Long short-term memory-networks for machine reading
Cheng, J., Dong, L., and Lapata, M · 2016
Earlier work this paper cites.
A decomposable attention model for natural language inference
Parikh, A. P., Täckström, O., Das, D., and Uszkoreit, J · 2016
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Cited alongside, same era.
Gated self-matching networks for reading comprehension and question answering
Wang, W., Yang, N., Wei, F., Chang, B., and Zhou, M · 2017
Cited alongside, same era.
Dehghani, M., Gouws, S., Vinyals, O., Uszkoreit, J., and Kaiser, Ł · 2018
Cited alongside, same era.
Huang, C.-Z. A., Vaswani, A., Uszkoreit, J., Shazeer, N., Simon, I., Hawthorne, C., Dai, A. M., Hoffman, M. D., Dinculescu, M., and Eck, D · 2018
Cited alongside, same era.
Phrase-indexed question answering: A new challenge for scalable document comprehension
Seo, M., Kwiatkowski, T., Parikh, A. P., Farhadi, A., and Hajishirzi, H · 2018
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2019
Later among the works it cites.
Pay less attention with lightweight and dynamic convolutions
Wu, F., Fan, A., Baevski, A., Dauphin, Y. N., and Auli, M · 2019
Later among the works it cites.
Longformer: The long-document transformer
Beltagy, I., Peters, M. E., and Cohan, A · 2020
Closest in time.
Reformer: The efficient transformer
Kitaev, N., Kaiser, Ł., and Levskaya, A · 2020
Closest in time.
Fixed encoder self-attention patterns in transformer-based machine translation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Mesh-tensorflow: Deep learning for supercomputers
Shazeer, N., Cheng, Y., Parmar, N., Tran, D., Vaswani, A., Koanantakool, P., Hawkins, P., Lee, H., Hong, M., Young, C., et al · 2018
Cited alongside, same era.
Generating long sequences with sparse transformers
Child, R., Gray, S., Radford, A., and Sutskever, I · 2019
Cited alongside, same era.
On the relationship between self-attention and convolutional layers
Cordonnier, J.-B., Loukas, A., and Jaggi, M · 2019
Cited alongside, same era.
Raganato, A., Scherrer, Y., and Tiedemann, J · 2020
Closest in time.
Tay, Y., Bahri, D., Yang, L., Metzler, D., and Juan, D.-C · 2020
Closest in time.
Linformer: Self-attention with linear complexity
Wang, S., Li, B., Khabsa, M., Fang, H., and Ma, H · 2020
Closest in time.
Mlp-mixer: An all-mlp architecture for vision
Tolstikhin, I., Houlsby, N., Kolesnikov, A., Beyer, L., Zhai, X., Unterthiner, T., Yung, J., Keysers, D., Uszkoreit, J., Lucic, M., et al · 2021
Closest in time.