Fetching the paper…
Reading the bibliography…
Attention is an important mechanism that can be employed for a variety of deep learning models across many different domains and tasks.
R. J. Williams, “Simple statistical gradient-following algorithms for connectionist reinforcement learning,” Machine Learning , vol. 8, no. 3, pp. 229–256, 1992
1992
Earlier work this paper cites.
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu, “BLEU: a method for automatic evaluation of machine translation,” in 40th Annual Meeting of the Association for Computational Linguistics (ACL 2002) . ACL, 2002, pp. 311–318
2002
Earlier work this paper cites.
P. Schwarz, P. Matějka, and J. Černockỳ, “Towards lower error rates in phoneme recognition,” in 7th International Conference on Text, Speech and Dialogue (TSD 2004) , ser. LNCS, vol. 3206. Springer, 2004, pp. 465–472
2004
Earlier work this paper cites.
D. S. Turaga, Y. Chen, and J. Caviedes, “No reference PSNR estimation for compressed pictures,” Signal Processing: Image Communication , vol. 19, no. 2, pp. 173–184, 2004
2004
Earlier work this paper cites.
S. Banerjee and A. Lavie, “METEOR: An automatic metric for MT evaluation with improved correlation with human judgments,” in 2005 Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and/or Summarization . ACL, 2005, pp. 65–72
2005
Earlier work this paper cites.
M. Popović and H. Ney, “Word error rates: Decomposition over POS classes and applications for error analysis,” in 2nd Workshop on Statistical Machine Translation (WMT 2007) . ACL, 2007, pp. 48–55
2007
Earlier work this paper cites.
H. Larochelle and G. E. Hinton, “Learning to combine foveal glimpses with a third-order Boltzmann machine,” in 24th Annual Conference in Neural Information Processing Systems (NIPS 2010) . Curran Associates, Inc., 2010, pp. 1243–1251
2010
Earlier work this paper cites.
P. Ndajah, H. Kikuchi, M. Yukawa, H. Watanabe, and S. Muramatsu, “SSIM image quality metric for denoised images,” in 3rd WSEAS International Conference on Visualization, Imaging and Simulation (VIS 2010) . WSEAS, 2010, pp. 53–58
2010
Earlier work this paper cites.
R. Sennrich, “Perplexity minimization for translation model domain adaptation in statistical machine translation,” in 13th Conference of the European Chapter of the Association for Computational Linguistics (EACL 2012) . ACL, 2012, pp. 539–549
2012
Earlier work this paper cites.
V. Mnih, N. Heess, A. Graves, and k. kavukcuoglu, “Recurrent models of visual attention,” in 27th Annual Conference on Neural Information Processing Systems (NIPS 2014) . Curran Associates, Inc., 2014, pp. 2204–2212
2014
Earlier work this paper cites.
K. Cho, B. van Merriënboer, D. Bahdanau, and Y. Bengio, “On the properties of neural machine translation: Encoder–decoder approaches,” in 8th Workshop on Syntax, Semantics and Structure in Statistical Translation (SSST 2014) . ACL, 2014, pp. 103–111
2014
Earlier work this paper cites.
A. Graves, G. Wayne, and I. Danihelka, “Neural Turing machines,” arXiv:1410.5401 , 2014
2014
Earlier work this paper cites.
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, “Generative adversarial nets,” in 27th Annual Conference on Neural Information Processing Systems (NIPS 2014) . Curran Associates, Inc., 2014, pp. 2672–2680
2014
Earlier work this paper cites.
D. Bahdanau, K. Cho, and Y. Bengio, “Neural machine translation by jointly learning to align and translate,” in 3rd International Conference on Learning Representation (ICLR 2015) , 2015
2015
Earlier work this paper cites.
T. Luong, H. Pham, and C. D. Manning, “Effective approaches to attention-based neural machine translation,” in 2015 Conference on Empirical Methods in Natural Language Processing (EMNLP 2015) . ACL, 2015, pp. 1412–1421
2015
Earlier work this paper cites.
J. K. Chorowski, D. Bahdanau, D. Serdyuk, K. Cho, and Y. Bengio, “Attention-based models for speech recognition,” in 28th Annual Conference on Neural Information Processing Systems (NIPS 2015) . Curran Associates, Inc., 2015, pp. 577–585
2015
Earlier work this paper cites.
Z. Yang, D. Yang, C. Dyer, X. He, A. Smola, and E. Hovy, “Hierarchical attention networks for document classification,” in 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT 2016) . ACL, 2016, pp. 1480–1489
2016
Earlier work this paper cites.
Y. Wang, M. Huang, X. Zhu, and L. Zhao, “Attention-based LSTM for aspect-level sentiment classification,” in 2016 Conference on Empirical Methods in Natural Language Processing (EMNLP 2016) . ACL, 2016, pp. 606–615
2016
Earlier work this paper cites.
D. Bahdanau, J. Chorowski, D. Serdyuk, P. Brakel, and Y. Bengio, “End-to-end attention-based large vocabulary speech recognition,” in 2016 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP 2016) . IEEE Signal Processing Society, 2016, pp. 4945–4949
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
S. Sharma, R. Kiros, and R. Salakhutdinov, “Action recognition using visual attention,” in Proceedings of the 4th International Conference on Learning Representations Workshop (ICLR 2016) , 2016
2016
Earlier work this paper cites.
J. Lu, J. Yang, D. Batra, and D. Parikh, “Hierarchical question-image co-attention for visual question answering,” in 30th Annual Conference on Neural Information Processing Systems (NIPS 2016) . Curran Associates, Inc., 2016, pp. 289–297
2016
Earlier work this paper cites.
M. Seo, A. Kembhavi, A. Farhadi, and H. Hajishirzi, “Bidirectional attention flow for machine comprehension,” in 4th International Conference on Learning Representations (ICLR 2016) , 2016
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
J. L. Ba, J. R. Kiros, and G. E. Hinton, “Layer normalization,” arXiv:1607.06450 , 2016
2016
Earlier work this paper cites.
Z. Yang, X. He, J. Gao, L. Deng, and A. Smola, “Stacked attention networks for image question answering,” in 2016 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2016) , 2016, pp. 21–29
2016
Earlier work this paper cites.
M. A. Rahman and Y. Wang, “Optimizing intersection-over-union in deep neural networks for image segmentation,” in 12th International Symposium on Visual Computing (ISVC 2016) , ser. LNCS, vol. 10072. Springer, 2016, pp. 234–244
2016
Earlier work this paper cites.
S. Kim, T. Hori, and S. Watanabe, “Joint CTC-attention based end-to-end speech recognition using multi-task learning,” in 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP 2017) . IEEE Signal Processing Society, 2017, pp. 4835–4839
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. u. Kaiser, and I. Polosukhin, “Attention is all you need,” in 31st Annual Conference on Neural Information Processing Systems (NIPS 2017) . Curran Associates, Inc., 2017, pp. 5998–6008
2017
Earlier work this paper cites.
M. Daniluk, T. Rocktäschel, J. Welbl, and S. Riedel, “Frustratingly short attention spans in neural language modeling,” in 5th International Conference on Learning Representations (ICLR 2017) , 2017
2017
Earlier work this paper cites.
Y. Xu, Q. Kong, Q. Huang, W. Wang, and M. D. Plumbley, “Attention and localization based on a deep convolutional recurrent model for weakly supervised audio tagging,” in Proceedings of the 18th Annual Conference of the International Speech Communication Association (Interspeech 2017) . ISCA, 2017, pp. 3083–3087
2017
Earlier work this paper cites.
L. Gao, Z. Guo, H. Zhang, X. Xu, and H. T. Shen, “Video captioning with attention-based LSTM and semantic consistency,” IEEE Transactions on Multimedia , vol. 19, no. 9, pp. 2045–2055, 2017
2017
Earlier work this paper cites.
D. Ma, S. Li, X. Zhang, and H. Wang, “Interactive attention networks for aspect-level sentiment classification,” in 26th International Joint Conference on Artificial Intelligence (IJCAI 2017) . IJCAI, 2017, pp. 4068–4074
2017
Earlier work this paper cites.
D. Britz, A. Goldie, M.-T. Luong, and Q. Le, “Massive exploration of neural machine translation architectures,” in 2017 Conference on Empirical Methods in Natural Language Processing (EMNLP 2017) . ACL, 2017, pp. 1442–1451
2017
Earlier work this paper cites.
S. Seo, J. Huang, H. Yang, and Y. Liu, “Interpretable convolutional neural networks with dual local and global attention for review rating prediction,” in 11th ACM Conference on Recommender Systems (RecSys 2017) . ACM, 2017, pp. 297–305
2017
Earlier work this paper cites.
Z. Lin, M. Feng, C. N. d. Santos, M. Yu, B. Xiang, B. Zhou, and Y. Bengio, “A structured self-attentive sentence embedding,” in 5th International Conference on Learning Representations (ICLR 2017) , 2017
2017
Earlier work this paper cites.
P. Veličković, G. Cucurull, A. Casanova, A. Romero, P. Liò, and Y. Bengio, “Graph attention networks,” in 5th International Conference on Learning Representations (ICLR 2017) , 2017
2017
Earlier work this paper cites.
Y. Gong and S. R. Bowman, “Ruminating reader: Reasoning with gated multi-hop attention,” in 5th International Conference on Learning Representation (ICLR 2017) , 2017
2017
Earlier work this paper cites.
S. Sabour, N. Frosst, and G. E. Hinton, “Dynamic routing between capsules,” in 31st Annual Conference on Neural Information Processing Systems (NIPS 2017) . Curran Associates, Inc., 2017, p. 3859–3869
2017
Earlier work this paper cites.
S. Liu, Y. Chen, K. Liu, and J. Zhao, “Exploiting argument information to improve event detection via supervised attention mechanisms,” in 55th Annual Meeting of the Association for Computational Linguistics (ACL 2017) . ACL, 2017, pp. 1789–1798
2017
Earlier work this paper cites.
C. Liu, J. Mao, F. Sha, and A. Yuille, “Attention correctness in neural image captioning,” in 31st AAAI Conference on Artificial Intelligence (AAAI 2017) . AAAI Press, 2017, pp. 4176–4182
2017
Earlier work this paper cites.
A. Das, H. Agrawal, L. Zitnick, D. Parikh, and D. Batra, “Human attention in visual question answering: Do humans and deep networks look at the same regions?” Computer Vision and Image Understanding , vol. 163, pp. 90 – 100, 2017
2017
Cited alongside, same era.
Y. Yu, J. Choi, Y. Kim, K. Yoo, S.-H. Lee, and G. Kim, “Supervising neural attention models for video captioning by human gaze data,” in 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2017) . IEEE Computer Society, 2017
2017
Cited alongside, same era.
P. Anderson, X. He, C. Buehler, D. Teney, M. Johnson, S. Gould, and L. Zhang, “Bottom-up and top-down attention for image captioning and visual question answering,” in 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2018) , 2018, pp. 6077–6086
2018
Cited alongside, same era.
Y. Ma, H. Peng, and E. Cambria, “Targeted aspect-based sentiment analysis via embedding commonsense knowledge into an attentive LSTM,” in 32nd AAAI Conference on Artificial Intelligence (AAAI 2018) . AAAI Press, 2018, pp. 5876–5883
L. Wu, L. Chen, R. Hong, Y. Fu, X. Xie, and M. Wang, “A hierarchical attention model for social contextual image recommendation,” IEEE Transactions on Knowledge and Data Engineering , 2019
2019
Later among the works it cites.
V. A. Sindagi and V. M. Patel, “HA-CCN: Hierarchical attention-based crowd counting network,” IEEE Transactions on Image Processing , vol. 29, pp. 323–335, 2019
2019
Later among the works it cites.
G. I. Winata, Z. Lin, and P. Fung, “Learning multilingual meta-embeddings for code-switching named entity recognition,” in 4th Workshop on Representation Learning for NLP (RepL4NLP 2019) . ACL, 2019, pp. 181–186
2019
Later among the works it cites.
R. Jin, L. Lu, J. Lee, and A. Usman, “Multi-representational convolutional neural networks for text classification,” Computational Intelligence , vol. 35, no. 3, pp. 599–609, 2019
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2018
Cited alongside, same era.
N. Parmar, A. Vaswani, J. Uszkoreit, L. Kaiser, N. Shazeer, A. Ku, and D. Tran, “Image Transformer,” in 35th International Conference on Machine Learning (ICML 2018) , vol. 80. PMLR, 2018, pp. 4055–4064
2018
Cited alongside, same era.
L. Zhou, Y. Zhou, J. J. Corso, R. Socher, and C. Xiong, “End-to-end dense video captioning with masked transformer,” in 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2018) . IEEE Computer Society, 2018, pp. 8739–8748
2018
Cited alongside, same era.
C. Yu, K. S. Barsim, Q. Kong, and B. Yang, “Multi-level attention model for weakly supervised audio classification,” in Proceedings of the Detection and Classification of Acoustic Scenes and Events 2018 Workshop (DCASE 2018) , 2018, pp. 188–192
2018
Cited alongside, same era.
H. Ying, F. Zhuang, F. Zhang, Y. Liu, G. Xu, X. Xie, H. Xiong, and J. Wu, “Sequential recommender system based on hierarchical attention networks,” in 27th International Joint Conference on Artificial Intelligence (IJCAI 2018) . IJCAI, 2018, pp. 3926–3932
2018
Cited alongside, same era.
H. Song, D. Rajan, J. Thiagarajan, and A. Spanias, “Attend and diagnose: Clinical time series analysis using attention models,” in 32nd AAAI Conference on Artificial Intelligence (AAAI 2018) . AAAI Press, 2018, pp. 4091–4098
2018
Cited alongside, same era.
P. Veličković, G. Cucurull, A. Casanova, A. Romero, P. Lio, and Y. Bengio, “Graph attention networks,” in 6th International Conference on Learning Representations (ICLR 2018) , 2018
2018
Cited alongside, same era.
F. Fan, Y. Feng, and D. Zhao, “Multi-grained attention network for aspect-level sentiment classification,” in 2018 Conference on Empirical Methods in Natural Language Processing (EMNLP 2018) . ACL, 2018, pp. 3433–3442
2018
Cited alongside, same era.
2018
Cited alongside, same era.
J. Wang, C. Sun, S. Li, X. Liu, L. Si, M. Zhang, and G. Zhou, “Aspect sentiment classification towards question-answering with reinforced bidirectional attention network,” in 57th Annual Meeting of the Association for Computational Linguistics (ACL 2019) . ACL, 2019, pp. 3548–3557
2019
Later among the works it cites.
O. Arshad, I. Gallo, S. Nawaz, and A. Calefati, “Aiding intra-text representations with visual context for multimodal named entity recognition,” in 2019 International Conference on Document Analysis and Recognition (ICDAR 2019) . IEEE, 2019, pp. 337–342
2019
Later among the works it cites.
R. Tan, J. Sun, B. Su, and G. Liu, “Extending the transformer with context and multi-dimensional mechanism for dialogue response generation,” in 8th International Conference on Natural Language Processing and Chinese Computing (NLPCC 2019) , ser. LNCS, J. Tang, M.-Y. Kan, D. Zhao, S. Li, and H. Zan, Eds., vol. 11839. Springer, 2019, pp. 189–199
2019
Later among the works it cites.
H. Wang, G. Liu, A. Liu, Z. Li, and K. Zheng, “Dmran: A hierarchical fine-grained attention-based network for recommendation,” in 28th International Joint Conference on Artificial Intelligence (IJCAI 2019)
2019
Later among the works it cites.
H. Zhang, I. Goodfellow, D. Metaxas, and A. Odena, “Self-attention generative adversarial networks,” in 36th International Conference on Machine Learning (ICML 2019) , ser. Proceedings of Machine Learning Research, K. Chaudhuri and R. Salakhutdinov, Eds., vol. 97. PMLR, 2019, pp. 7354–7363
2019
Later among the works it cites.
J. Salazar, K. Kirchhoff, and Z. Huang, “Self-attention networks for connectionist temporal classification in speech recognition,” in 2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP 2019) . IEEE, 2019, pp. 7115–7119
2019
Later among the works it cites.
S. Iida, R. Kimura, H. Cui, P.-H. Hung, T. Utsuro, and M. Nagata, “Attention over heads: A multi-hop attention for neural machine translation,” in 57th Annual Meeting of the Association for Computational Linguistics: Student Research Workshop (ACL-SRW 2019) . ACL, 2019, pp. 217–222
2019
Later among the works it cites.
S. Yoon, S. Byun, S. Dey, and K. Jung, “Speech emotion recognition using multi-hop attention mechanism,” in 2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP 2019) . IEEE, 2019, pp. 2822–2826
2019
Later among the works it cites.
M. India, P. Safari, and J. Hernando, “Self multi-head attention for speaker recognition,” in Proceedings of the 20th Annual Conference of the International Speech Communication Association (Interspeech 2019) . ISCA, 2019, pp. 2822–2826
2019
Later among the works it cites.
C. Wu, F. Wu, S. Ge, T. Qi, Y. Huang, and X. Xie, “Neural news recommendation with multi-head self-attention,” in 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP 2019) . ACL, 2019, pp. 6389–6394
2019
Later among the works it cites.
P. Zhong, D. Wang, and C. Miao, “Knowledge-enriched transformer for emotion detection in textual conversations,” in 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP 2019) . ACL, 2019, pp. 165–176
2019
Later among the works it cites.
Y. Zhou, R. Ji, J. Su, X. Sun, and W. Chen, “Dynamic capsule attention for visual question answering,” in 33rd AAAI Conference on Artificial Intelligence (AAAI 2019) , vol. 33, no. 01. AAAI Press, 2019, pp. 9324–9331
2019
Later among the works it cites.
Z. Dai, Z. Yang, Y. Yang, J. Carbonell, Q. Le, and R. Salakhutdinov, “Transformer-XL: Attentive language models beyond a fixed-length context,” in 57th Annual Meeting of the Association for Computational Linguistics (ACL 2019) . ACL, 2019, pp. 2978–2988
2019
Later among the works it cites.
X. Li, J. Song, L. Gao, X. Liu, W. Huang, X. He, and C. Gan, “Beyond RNNs: Positional self-attention with co-attention for video question answering,” in 33rd AAAI Conference on Artificial Intelligence (AAAI 2019) , vol. 33. AAAI Press, 2019, pp. 8658–8665
2019
Later among the works it cites.
S. Jain and B. C. Wallace, “Attention is not explanation,” in 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT 2019) . ACL, 2019, pp. 3543–3556
2019
Later among the works it cites.
S. Wiegreffe and Y. Pinter, “Attention is not not explanation,” in 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP 2019) . ACL, 2019, pp. 11–20
2019
Later among the works it cites.
R. He, W. S. Lee, H. T. Ng, and D. Dahlmeier, “An unsupervised neural attention model for aspect extraction,” in 55th Annual Meeting of the Association for Computational Linguistics (ACL 2017) . ACL, 2017, pp. 388–397
2019
Later among the works it cites.
D. Hu, “An introductory survey on attention mechanisms in NLP problems,” in Proceedings of the 2019 Intelligent Systems Conference (IntelliSys 2019) , ser. AISC, vol. 1038. Springer, 2020, pp. 432–448
2020
Later among the works it cites.
Y.-J. Lu and C.-T. Li, “GCAN: Graph-aware co-attention networks for explainable fake news detection on social media,” in 58th Annual Meeting of the Association for Computational Linguistics (ACL 2020) . ACL, 2020, pp. 505–514
2020
Later among the works it cites.
Y. Liu, W. Wang, Y. Hu, J. Hao, X. Chen, and Y. Gao, “Multi-agent game abstraction via graph attention neural network,” in 34th AAAI Conference on Artificial Intelligence (AAAI 2020) , vol. 34, no. 05. AAAI Press, 2020, pp. 7211–7218
2020
Later among the works it cites.
M. Jiang, C. Li, J. Kong, Z. Teng, and D. Zhuang, “Cross-level reinforced attention network for person re-identification,” Journal of Visual Communication and Image Representation , vol. 69, p. 102775, 2020
2020
Later among the works it cites.
L. Chen, B. Lv, C. Wang, S. Zhu, B. Tan, and K. Yu, “Schema-guided multi-domain dialogue state tracking with graph attention neural networks,” in 34th AAAI Conference on Artificial Intelligence (AAAI 2020) , vol. 34, no. 05. AAAI Press, 2020, pp. 7521–7528
2020
Later among the works it cites.
H. Zhao, J. Jia, and V. Koltun, “Exploring self-attention for image recognition,” in 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2020) , 2020, pp. 10 076–10 085
2020
Later among the works it cites.
A. Sankar, Y. Wu, L. Gou, W. Zhang, and H. Yang, “Dysat: Deep neural representation learning on dynamic graphs via self-attention networks,” in 13th International Conference on Web Search and Data Mining (WSDM 2020) , 2020, pp. 519–527
2020
Later among the works it cites.
M. Cornia, M. Stefanini, L. Baraldi, and R. Cucchiara, “Meshed-memory transformer for image captioning,” in 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2020) , 2020, pp. 10 578–10 587
2020
Later among the works it cites.
N. Kitaev, Ł. Kaiser, and A. Levskaya, “Reformer: The efficient Transformer,” in 8th International Conference on Learning Representations (ICLR 2020) , 2020
2020
Later among the works it cites.
2020
Later among the works it cites.
Z. Wu, Z. Liu, J. Lin, Y. Lin, and S. Han, “Lite transformer with long-short range attention,” in 8th International Conference on Learning Representations (ICLR 2020) , 2020
2020
Later among the works it cites.
2020
Later among the works it cites.
2020
Later among the works it cites.
A. K. Mohankumar, P. Nema, S. Narasimhan, M. M. Khapra, B. V. Srinivasan, and B. Ravindran, “Towards transparent and explainable attention models,” in 58th Annual Meeting of the Association for Computational Linguistics (ACL 2020) . ACL, 2020, pp. 4206–4216
2020
Later among the works it cites.
S. Chaudhari, V. Mithal, G. Polatkan, and R. Ramanath, “An attentive survey of attention models,” ACM Transactions on Intelligent Systems and Technology , vol. 12, no. 5, pp. 1–32, 2021
2021
Later among the works it cites.
A. Sinha and J. Dolz, “Multi-scale self-guided attention for medical image segmentation,” IEEE Journal of Biomedical and Health Informatics , vol. 25, no. 1, pp. 121–130, 2021
2021
Later among the works it cites.
Y. Wang, W. Chen, D. Pi, and L. Yue, “Adversarially regularized medication recommendation model with multi-hop memory network,” Knowledge and Information Systems , vol. 63, no. 1, pp. 125–142, 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
Y. Tay, D. Bahri, D. Metzler, D.-C. Juan, Z. Zhao, and C. Zheng, “Synthesizer: Rethinking self-attention for transformer models,” in Proceedings of the 38th International Conference on Machine Learning (ICML 2021) , vol. 139. PMLR, 2021, pp. 10 183–10 192
2021
Later among the works it cites.
Y. Wang, A. Sun, M. Huang, and X. Zhu, “Aspect-level sentiment analysis using AS-capsules,” in The World Wide Web Conference , 2019, pp. 2033–2044
2044
Closest in time.
K. Xu, J. Ba, R. Kiros, K. Cho, A. Courville, R. Salakhudinov, R. Zemel, and Y. Bengio, “Show, attend and tell: Neural image caption generation with visual attention,” in 32nd International Conference on Machine Learning (ICML 2015) , vol. 37. PMLR, 2015, pp. 2048–2057
2057
Closest in time.