Fetching the paper…
Reading the bibliography…
Attention Model has now become an important concept in neural networks that has been researched within diverse application domains.
Language Models are Few-Shot Learners. In Advances in Neural Information Processing Systems , Vol. 33. 1877–1901
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020 · 1901
Earlier work this paper cites.
Dipole: Diagnosis Prediction in Healthcare via Attention-Based Bidirectional Recurrent Neural Networks. In Proceedings of the 23rd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining . Association for Computing Machinery, 1903–1911
Fenglong Ma, Radha Chitta, Jing Zhou, Quanzeng You, Tong Sun, and Jing Gao. 2017a · 1911
Earlier work this paper cites.
On estimating regression
Elizbar A Nadaraya. 1964 · 1964
Earlier work this paper cites.
Smooth regression analysis
Geoffrey S Watson. 1964 · 1964
Earlier work this paper cites.
Multiple Object Recognition with Visual Attention
Lei Jimmy Ba, Volodymyr Mnih, and Koray Kavukcuoglu. 2014 · 2014
Earlier work this paper cites.
Learning Phrase Representations using RNN Encoder–Decoder for Statistical Machine Translation. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing . Association for Computational Linguistics, 1724–1734
Kyunghyun Cho, Bart van Merriënboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio. 2014b · 2014
Earlier work this paper cites.
Alex Graves, Greg Wayne, and Ivo Danihelka. 2014a · 2014
Earlier work this paper cites.
Alex Graves, Greg Wayne, and Ivo Danihelka. 2014b · 2014
Earlier work this paper cites.
Recurrent Models of Visual Attention. In Proceedings of the 27th International Conference on Neural Information Processing Systems - Volume 2 . MIT Press, 2204–2212
Volodymyr Mnih, Nicolas Heess, Alex Graves, and Koray Kavukcuoglu. 2014 · 2014
Earlier work this paper cites.
Jason Weston, Sumit Chopra, and Antoine Bordes. 2014 · 2014
Earlier work this paper cites.
Neural Machine Translation by Jointly Learning to Align and Translate. In 3rd International Conference on Learning Representations
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2015 · 2015
Earlier work this paper cites.
Describing multimedia content using attention-based encoder-decoder networks
Kyunghyun Cho, Aaron Courville, and Yoshua Bengio. 2015 · 2015
Earlier work this paper cites.
Attention-based Models for Speech Recognition. In Advances in Neural Information Processing Systems . MIT Press, 577–585
Jan Chorowski, Dzmitry Bahdanau, Dmitriy Serdyuk, Kyunghyun Cho, and Yoshua Bengio. 2015 · 2015
Earlier work this paper cites.
DRAW: A Recurrent Neural Network For Image Generation (Proceedings of Machine Learning Research, Vol. 37) . PMLR, 1462–1471
Karol Gregor, Ivo Danihelka, Alex Graves, Danilo Rezende, and Daan Wierstra. 2015 · 2015
Earlier work this paper cites.
Teaching Machines to Read and Comprehend
Karl Moritz Hermann, Tomas Kocisky, Edward Grefenstette, Lasse Espeholt, Will Kay, Mustafa Suleyman, and Phil Blunsom. 2015 · 2015
Earlier work this paper cites.
Effective approaches to attention-based neural machine translation
Minh-Thang Luong, Hieu Pham, and Christopher D Manning. 2015a · 2015
Earlier work this paper cites.
Effective Approaches to Attention-based Neural Machine Translation. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing . Association for Computational Linguistics, 1412–1421
Thang Luong, Hieu Pham, and Christopher D. Manning. 2015b · 2015
Earlier work this paper cites.
A Neural Attention Model for Abstractive Sentence Summarization. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing . Association for Computational Linguistics, 379–389
Alexander M. Rush, Sumit Chopra, and Jason Weston. 2015 · 2015
Earlier work this paper cites.
End-To-End Memory Networks
Sainbayar Sukhbaatar, arthur szlam, Jason Weston, and Rob Fergus. 2015 · 2015
Earlier work this paper cites.
Pointer Networks. In Advances in Neural Information Processing Systems 28 . MIT Press, 2692–2700
Oriol Vinyals, Meire Fortunato, and Navdeep Jaitly. 2015 · 2015
Earlier work this paper cites.
Describing Videos by Exploiting Temporal Structure. In Proceedings of the 2015 IEEE International Conference on Computer Vision . IEEE Computer Society, 4507–4515
Li Yao, Atousa Torabi, Kyunghyun Cho, Nicolas Ballas, Christopher Pal, Hugo Larochelle, and Aaron Courville. 2015 · 2015
Earlier work this paper cites.
Listen, attend and spell: A neural network for large vocabulary conversational speech recognition. In International Conference on Acoustics, Speech, and Signal Processing . IEEE, 4960–4964
William Chan, Navdeep Jaitly, Quoc Le, and Oriol Vinyals. 2016 · 2016
Earlier work this paper cites.
Abstractive Sentence Summarization with Attentive Recurrent Neural Networks. In Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies . Association for Computational Linguistics, 93–98
Sumit Chopra, Michael Auli, and Alexander M. Rush. 2016 · 2016
Earlier work this paper cites.
GAKE: Graph Aware Knowledge Embedding. In Proceedings of the 26th International Conference on Computational Linguistics . The COLING 2016 Organizing Committee, 641–651
Jun Feng, Minlie Huang, Yang Yang, and Xiaoyan Zhu. 2016 · 2016
Earlier work this paper cites.
Ask Me Anything: Dynamic Memory Networks for Natural Language Processing. In Proceedings of The 33rd International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 48) . PMLR, 1378–1387
Ankit Kumar, Ozan Irsoy, Peter Ondruska, Mohit Iyyer, James Bradbury, Ishaan Gulrajani, Victor Zhong, Romain Paulus, and Richard Socher. 2016 · 2016
Earlier work this paper cites.
Understanding Neural Networks through Representation Erasure
Jiwei Li, Will Monroe, and Dan Jurafsky. 2016 · 2016
Earlier work this paper cites.
Hierarchical Question-Image Co-Attention for Visual Question Answering
Jiasen Lu, Jianwei Yang, Dhruv Batra, and Devi Parikh. 2016 · 2016
Earlier work this paper cites.
From softmax to sparsemax: A sparse model of attention and multi-label classification. In International Conference on Machine Learning . PMLR, 1614–1623
Andre Martins and Ramon Astudillo. 2016 · 2016
Earlier work this paper cites.
Abstractive Text Summarization using Sequence-to-sequence RNNs and Beyond. In Proceedings of The 20th SIGNLL Conference on Computational Natural Language Learning . Association for Computational Linguistics, 280–290
Ramesh Nallapati, Bowen Zhou, Cicero dos Santos, Çağlar Gu̇lçehre, and Bing Xiang. 2016 · 2016
Earlier work this paper cites.
Iterative alternating neural attention for machine reading
Alessandro Sordoni, Philip Bachman, Adam Trischler, and Yoshua Bengio. 2016 · 2016
Earlier work this paper cites.
Aspect Level Sentiment Classification with Deep Memory Network. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing . Association for Computational Linguistics, 214–224
Duyu Tang, Bing Qin, and Ting Liu. 2016 · 2016
Earlier work this paper cites.
Survey on the attention based RNN model and its applications in computer vision
Feng Wang and David MJ Tax. 2016 · 2016
Earlier work this paper cites.
Attention-based LSTM for Aspect-level Sentiment Classification. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing . Association for Computational Linguistics, 606–615
Yequan Wang, Minlie Huang, Xiaoyan Zhu, and Li Zhao. 2016 · 2016
Earlier work this paper cites.
Dynamic Memory Networks for Visual and Textual Question Answering. In Proceedings of the 33rd International Conference on International Conference on Machine Learning - Volume 48 . JMLR.org, 2397–2406
Caiming Xiong, Stephen Merity, and Richard Socher. 2016 · 2016
Earlier work this paper cites.
Hierarchical Attention Networks for Document Classification. In Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies . 1480–1489
Zichao Yang, Diyi Yang, Chris Dyer, Xiaodong He, Alex Smola, and Eduard Hovy. 2016 · 2016
Earlier work this paper cites.
Massive Exploration of Neural Machine Translation Architectures. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing . Association for Computational Linguistics, 1442–1451
Denny Britz, Anna Goldie, Minh-Thang Luong, and Quoc Le. 2017 · 2017
Earlier work this paper cites.
Chung-Cheng Chiu and Colin Raffel. 2017 · 2017
Earlier work this paper cites.
Interactive Visualization and Manipulation of Attention-based Neural Machine Translation. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing: System Demonstrations . Association for Computational Linguistics, 121–126
Jaesong Lee, Joong-Hwi Shin, and Jun-Seok Kim. 2017 · 2017
Earlier work this paper cites.
A Structured Self-attentive Sentence Embedding
Zhouhan Lin, Minwei Feng, Cícero Nogueira dos Santos, Mo Yu, Bing Xiang, Bowen Zhou, and Yoshua Bengio. 2017 · 2017
Earlier work this paper cites.
Interactive attention networks for aspect-level sentiment classification
Dehong Ma, Sujian Li, Xiaodong Zhang, and Houfeng Wang. 2017b · 2017
Earlier work this paper cites.
Deeper attention to abusive user content moderation. In Proceedings of the 2017 conference on empirical methods in natural language processing . 1125–1135
John Pavlopoulos, Prodromos Malakasiotis, and Ion Androutsopoulos. 2017 · 2017
Earlier work this paper cites.
Attention is All you Need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Coupled Multi-Layer Attentions for Co-Extraction of Aspect and Opinion Terms. In Proceedings of the Thirty-First AAAI Conference on Artificial Intelligence . AAAI Press, 3316–3322
Wenya Wang, Sinno Jialin Pan, Daniel Dahlmeier, and Xiaokui Xiao. 2017 · 2017
Earlier work this paper cites.
Link Prediction via Ranking Metric Dual-Level Attention Network Learning. In Proceedings of the 26th International Joint Conference on Artificial Intelligence . AAAI Press, 3525–3531
Zhou Zhao, Ben Gao, Vicent W. Zheng, Deng Cai, Xiaofei He, and Yueting Zhuang. 2017 · 2017
Earlier work this paper cites.
Watch Your Step: Learning Node Embeddings via Graph Attention
Sami Abu-El-Haija, Bryan Perozzi, Rami Al-Rfou, and Alexander A Alemi. 2018 · 2018
Cited alongside, same era.
Self-Attention: A Better Building Block for Sentiment Analysis Neural Network Classifiers. In Proceedings of the 9th Workshop on Computational Approaches to Subjectivity, Sentiment and Social Media Analysis . Association for Computational Linguistics, 130–139
Artaches Ambartsoumian and Fred Popowich. 2018 · 2018
Cited alongside, same era.
Bottom-Up and Top-Down Attention for Image Captioning and Visual Question Answering. In The IEEE Conference on Computer Vision and Pattern Recognition . IEEE Computer Society, 6077–6086
Peter Anderson, Xiaodong He, Chris Buehler, Damien Teney, Mark Johnson, Stephen Gould, and Lei Zhang. 2018 · 2018
Cited alongside, same era.
Graph-to-Sequence Learning using Gated Graph Neural Networks. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics . Association for Computational Linguistics, 273–283
Daniel Beck, Gholamreza Haffari, and Trevor Cohn. 2018 · 2018
Cited alongside, same era.
LXMERT: Learning Cross-Modality Encoder Representations from Transformers. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, EMNLP-IJCNLP 2019, Hong Kong, China, November 3-7, 2019 . Association for Computational Linguistics, 5099–5110
Hao Tan and Mohit Bansal. 2019 · 2019
Closest in time.
Compositional De-Attention Networks
Yi Tay, Anh Tuan Luu, Aston Zhang, Shuohang Wang, and Siu Cheung Hui. 2019 · 2019
Closest in time.
Analyzing Multi-Head Self-Attention: Specialized Heads Do the Heavy Lifting, the Rest Can Be Pruned. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics . Association for Computational Linguistics, 5797–5808
Elena Voita, David Talbot, Fedor Moiseev, Rico Sennrich, and Ivan Titov. 2019 · 2019
Closest in time.
Hierarchical User and Item Representation with Three-Tier Attention for Recommendation. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies . Association for Computational Linguistics, 1818–1826
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A Survey of Methods for Explaining Black Box Models
Riccardo Guidotti, Anna Monreale, Salvatore Ruggieri, Franco Turini, Fosca Giannotti, and Dino Pedreschi. 2018 · 2018
Cited alongside, same era.
NAIS: Neural Attentive Item Similarity Model for Recommendation
Xiangnan He, Zhankui He, Jingkuan Song, Zhenguang Liu, Yu-Gang Jiang, and Tat-Seng Chua. 2018 · 2018
Cited alongside, same era.
Learn to Pay Attention. In International Conference on Learning Representations
Saumya Jetley, Nicholas A. Lord, Namhoon Lee, and Philip Torr. 2018 · 2018
Cited alongside, same era.
Self-Attentive Sequential Recommendation. In IEEE International Conference on Data Mining . IEEE Computer Society, 197–206
Wang-Cheng Kang and Julian J. McAuley. 2018 · 2018
Cited alongside, same era.
Dynamic Meta-Embeddings for Improved Sentence Representations. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing . Association for Computational Linguistics, 1466–1477
Douwe Kiela, Changhan Wang, and Kyunghyun Cho. 2018 · 2018
Cited alongside, same era.
Graph Classification Using Structural Attention. In Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining . Association for Computing Machinery, 1666–1674
John Boaz Lee, Ryan Rossi, and Xiangnan Kong. 2018 · 2018
Cited alongside, same era.
Visual Interrogation of Attention-Based Models for Natural Language Inference and Machine Comprehension. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing: System Demonstrations . Association for Computational Linguistics, 36–41
Shusen Liu, Tao Li, Zhimin Li, Vivek Srikumar, Valerio Pascucci, and Peer-Timo Bremer. 2018 · 2018
Cited alongside, same era.
Targeted Aspect-Based Sentiment Analysis via Embedding Commonsense Knowledge into an Attentive LSTM. In Proceedings of the Thirty-Second AAAI Conference on Artificial Intelligence . AAAI Press, 5876–5883
Yukun Ma, Haiyun Peng, and Erik Cambria. 2018 · 2018
Cited alongside, same era.
Chuhan Wu, Fangzhao Wu, Junxin Liu, and Yongfeng Huang. 2019 · 2019
Closest in time.
XLNet: Generalized Autoregressive Pretraining for Language Understanding. In Advances in Neural Information Processing Systems , Vol. 32
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Russ R Salakhutdinov, and Quoc V Le. 2019 · 2019
Closest in time.
Deep modular co-attention networks for visual question answering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 6281–6290
Zhou Yu, Jun Yu, Yuhao Cui, Dacheng Tao, and Qi Tian. 2019 · 2019
Closest in time.
Deep Learning Based Recommender System: A Survey and New Perspectives
Shuai Zhang, Lina Yao, Aixin Sun, and Yi Tay. 2019b · 2019
Closest in time.
End-to-End Object Detection with Transformers. In European Conference on Computer Vision (Lecture Notes in Computer Science, Vol. 12346) . Springer, 213–229
Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. 2020 · 2020
Closest in time.
Generative Pretraining From Pixels. In Proceedings of the 37th International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 119) . PMLR, 1691–1703
Mark Chen, Alec Radford, Rewon Child, Jeffrey Wu, Heewoo Jun, David Luan, and Ilya Sutskever. 2020 · 2020
Closest in time.
DETERRENT: Knowledge Guided Graph Attention Network for Detecting Healthcare Misinformation. In Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining . Association for Computing Machinery, 492–502
Limeng Cui, Haeseung Seo, Maryam Tabar, Fenglong Ma, Suhang Wang, and Dongwon Lee. 2020 · 2020
Closest in time.
Hierarchical Bi-Directional Self-Attention Networks for Paper Review Rating Recommendation. In Proceedings of the 28th International Conference on Computational Linguistics . International Committee on Computational Linguistics, 6302–6314
Zhongfen Deng, Hao Peng, Congying Xia, Jianxin Li, Lifang He, and Philip Yu. 2020 · 2020
Closest in time.
Policy learning with partial observation and mechanical constraints for multi-person modeling
Keisuke Fujii, Naoya Takeishi, Yoshinobu Kawahara, and Kazuya Takeda. 2020 · 2020
Closest in time.
Attention in Natural Language Processing
Andrea Galassi, Marco Lippi, and Paolo Torroni. 2020 · 2020
Closest in time.
Reformer: The Efficient Transformer. In 8th International Conference on Learning Representations
Nikita Kitaev, Lukasz Kaiser, and Anselm Levskaya. 2020 · 2020
Closest in time.
Attention is Not Only a Weight: Analyzing Transformers with Vector Norms. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) . Association for Computational Linguistics, 7057–7075
Goro Kobayashi, Tatsuki Kuribayashi, Sho Yokoi, and Kentaro Inui. 2020 · 2020
Closest in time.
ALBERT: A Lite BERT for Self-supervised Learning of Language Representations. In International Conference on Learning Representations
Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. 2020 · 2020
Closest in time.
BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics . Association for Computational Linguistics, 7871–7880
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020 · 2020
Closest in time.
Sac: Accelerating and structuring self-attention via sparse adaptive connection
Xiaoya Li, Yuxian Meng, Mingxin Zhou, Qinghong Han, Fei Wu, and Jiwei Li. 2020b · 2020
Closest in time.
SPACE: Unsupervised Object-Oriented Scene Representation via Spatial Attention and Decomposition. In 8th International Conference on Learning Representations
Zhixuan Lin, Yi-Fu Wu, Skand Vishwanath Peri, Weihao Sun, Gautam Singh, Fei Deng, Jindong Jiang, and Sungjin Ahn. 2020 · 2020
Closest in time.
Auto Learning Attention
Benteng Ma, Jing Zhang, Yong Xia, and Dacheng Tao. 2020 · 2020
Closest in time.
Sparse and Continuous Attention Mechanisms
André FT Martins, Marcos Treviso, António Farinhas, Vlad Niculae, Mário AT Figueiredo, and Pedro MQ Aguiar. 2020 · 2020
Closest in time.
Improved knowledge distillation via teacher assistant. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 34. 5191–5198
Seyed Iman Mirzadeh, Mehrdad Farajtabar, Ang Li, Nir Levine, Akihiro Matsukawa, and Hassan Ghasemzadeh. 2020 · 2020
Closest in time.
Towards transparent and explainable attention models
Akash Kumar Mohankumar, Preksha Nema, Sharan Narasimhan, Mitesh M Khapra, Balaji Vasan Srinivasan, and Balaraman Ravindran. 2020 · 2020
Closest in time.
Towards Understanding Attention-Based Speech Recognition Models
C. Qin and D. Qu. 2020 · 2020
Closest in time.
Training data-efficient image transformers and distillation through attention
Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Hervé Jégou. 2020 · 2020
Closest in time.
Self-Attention with Linear Complexity
Sinong Wang, Belinda Li, Madian Khabsa, Han Fang, and H Linformer Ma. 2020a · 2020
Closest in time.
Minilm: Deep self-attention distillation for task-agnostic compression of pre-trained transformers
Wenhui Wang, Furu Wei, Li Dong, Hangbo Bao, Nan Yang, and Ming Zhou. 2020d · 2020
Closest in time.
Cross-Media Keyphrase Prediction: A Unified Framework with Multi-Modality Multi-Head Attention and Image Wordings. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing . Association for Computational Linguistics, 3311–3324
Yue Wang, Jing Li, Michael Lyu, and Irwin King. 2020b · 2020
Closest in time.
A Hierarchical Attention Model for Social Contextual Image Recommendation
Le Wu, Lei Chen, Richang Hong, Yanjie Fu, Xing Xie, and Meng Wang. 2020 · 2020
Closest in time.
Self-Attention Guided Copy Mechanism for Abstractive Summarization. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics . Association for Computational Linguistics, 1355–1362
Song Xu, Haoran Li, Peng Yuan, Youzheng Wu, Xiaodong He, and Bowen Zhou. 2020 · 2020
Closest in time.
Neural Machine Translation With GRU-Gated Attention Model
B. Zhang, D. Xiong, J. Xie, and J. Su. 2020 · 2020
Closest in time.
Hyper-SAGNN: a self-attention based graph neural network for hypergraphs. In International Conference on Learning Representations
Ruochi Zhang, Yuesong Zou, and Jian Ma. 2020 · 2020
Closest in time.
Exploring Self-Attention for Image Recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
Hengshuang Zhao, Jiaya Jia, and Vladlen Koltun. 2020 · 2020
Closest in time.
GMAN: A Graph Multi-Attention Network for Traffic Prediction. In The Thirty-Fourth AAAI Conference on Artificial Intelligence . AAAI Press, 1234–1241
Chuanpan Zheng, Xiaoliang Fan, Cheng Wang, and Jianzhong Qi. 2020 · 2020
Closest in time.
A Graphical and Attentional Framework for Dual-Target Cross-Domain Recommendation. In Proceedings of the Twenty-Ninth International Joint Conference on Artificial Intelligence . 3001–3008
Feng Zhu, Yan Wang, Chaochao Chen, Guanfeng Liu, and Xiaolin Zheng. 2020 · 2020
Closest in time.
Rethinking Attention with Performers. In International Conference on Learning Representations
Krzysztof Marcin Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Andreea Gane, Tamas Sarlos, Peter Hawkins, Jared Quincy Davis, Afroz Mohiuddin, Lukasz Kaiser, David Benjamin Belanger, Lucy J Colwell, and Adrian Weller. 2021 · 2021
Closest in time.
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. In International Conference on Learning Representations
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. 2021 · 2021
Closest in time.
Transformers in Vision: A Survey
Salman Khan, Muzammal Naseer, Munawar Hayat, Syed Waqas Zamir, Fahad Shahbaz Khan, and Mubarak Shah. 2021 · 2021
Closest in time.
Training data-efficient image transformers and distillation through attention
Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Hervé Jégou. 2021 · 2021
Closest in time.
Social explorative attention based recommendation for content distribution platforms
Wenyi Xiao, Huan Zhao, Haojie Pan, Yangqiu Song, Vincent W. Zheng, and Qiang Yang. 2021 · 2021
Closest in time.
Deformable DETR: Deformable Transformers for End-to-End Object Detection. In International Conference on Learning Representations
Xizhou Zhu, Weijie Su, Lewei Lu, Bin Li, Xiaogang Wang, and Jifeng Dai. 2021 · 2021
Closest in time.
Heterogeneous graph attention network. In The World Wide Web Conference . 2022–2032
Xiao Wang, Houye Ji, Chuan Shi, Bai Wang, Yanfang Ye, Peng Cui, and Philip S Yu. 2019b · 2032
Closest in time.
web (Proceedings of Machine Learning Research, Vol. 37) . 2048–2057
Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron Courville, Ruslan Salakhudinov, Rich Zemel, and Yoshua Bengio. 2015 · 2057
Closest in time.