Fetching the paper…
Reading the bibliography…
While the self-attention mechanism has been widely used in a wide variety of tasks, it has the unfortunate property of a quadratic cost with respect to the input length, which makes it difficult to deal with long inputs.
Eca-net: Efficient channel attention for deep convolutional neural networks
Qilong Wang, Banggu Wu, Pengfei Zhu, Peihua Li, Wangmeng Zuo, and Qinghua Hu. 2019 · 1910
Earlier work this paper cites.
Collective classification in network data
Prithviraj Sen, Galileo Namata, Mustafa Bilgic, Lise Getoor, Brian Galligher, and Tina Eliassi-Rad. 2008 · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. 2009 · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al. 2009 · 2009
Earlier work this paper cites.
Molecular signatures database (msigdb) 3.0
Arthur Liberzon, Aravind Subramanian, Reid Pinchback, Helga Thorvaldsdóttir, Pablo Tamayo, and Jill P Mesirov. 2011 · 2011
Earlier work this paper cites.
Large text compression benchmark
Matt Mahoney. 2011 · 2011
Earlier work this paper cites.
Deep Learning via Semi-supervised Embedding , pages 639–655. Springer Berlin Heidelberg, Berlin, Heidelberg
Jason Weston, Frédéric Ratle, Hossein Mobahi, and Ronan Collobert. 2012 · 2012
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
A fast and accurate dependency parser using neural networks
Danqi Chen and Christopher Manning. 2014 · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba. 2014 · 2014
Earlier work this paper cites.
Deepwalk: Online learning of social representations
Bryan Perozzi, Rami Al-Rfou, and Steven Skiena. 2014 · 2014
Earlier work this paper cites.
Convolutional neural networks on graphs with fast localized spectral filtering
Michaël Defferrard, Xavier Bresson, and Pierre Vandergheynst. 2016 · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Earlier work this paper cites.
Semi-supervised classification with graph convolutional networks
Thomas N. Kipf and Max Welling. 2016 · 2016
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Earlier work this paper cites.
Rethinking the inception architecture for computer vision
C. Szegedy, V. Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna. 2016 · 2016
Earlier work this paper cites.
Wide residual networks
Sergey Zagoruyko and Nikos Komodakis. 2016 · 2016
Earlier work this paper cites.
Geometric deep learning on graphs and manifolds using mixture model cnns
F. Monti, D. Boscaini, J. Masci, E. Rodola, J. Svoboda, and M. M. Bronstein. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Petar Veličković, Guillem Cucurull, Arantxa Casanova, Adriana Romero, Pietro Lio, and Yoshua Bengio. 2017 · 2017
Cited alongside, same era.
Predicting multicellular function through multi-layer tissue networks
Marinka Zitnik and Jure Leskovec. 2017 · 2017
Cited alongside, same era.
Watch your step: Learning node embeddings via graph attention
Sami Abu-El-Haija, Bryan Perozzi, Rami Al-Rfou, and Alexander A Alemi. 2018 · 2018
Cited alongside, same era.
Graph classification using structural attention
John Boaz Lee, Ryan Rossi, and Xiangnan Kong. 2018 · 2018
Cited alongside, same era.
Star-transformer
Qipeng Guo, Xipeng Qiu, Pengfei Liu, Yunfan Shao, Xiangyang Xue, and Zheng Zhang. 2019 · 2019
Later among the works it cites.
Local relation networks for image recognition
Han Hu, Zheng Zhang, Zhenda Xie, and Stephen Lin. 2019 · 2019
Later among the works it cites.
Self-attention graph pooling
Junhyun Lee, Inyeop Lee, and Jaewoo Kang. 2019 · 2019
Later among the works it cites.
Lanczosnet: Multi-scale deep graph convolutional networks
Renjie Liao, Zhizhen Zhao, Raquel Urtasun, and Richard Zemel. 2019 · 2019
Later among the works it cites.
Stand-alone self-attention in vision models
Niki Parmar, Prajit Ramachandran, Ashish Vaswani, Irwan Bello, Anselm Levskaya, and Jon Shlens. 2019 · 2019
Later among the works it cites.
Adaptive attention span in transformers
Sainbayar Sukhbaatar, Edouard Grave, Piotr Bojanowski, and Armand Joulin. 2019 · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yanbin Liu, Juho Lee, Minseop Park, Saehoon Kim, Eunho Yang, Sung Ju Hwang, and Yi Yang. 2018 · 2018
Cited alongside, same era.
Scaling neural machine translation
Myle Ott, Sergey Edunov, David Grangier, and Michael Auli. 2018 · 2018
Cited alongside, same era.
Self-attention with relative position representations
Peter Shaw, Jakob Uszkoreit, and Ashish Vaswani. 2018 · 2018
Cited alongside, same era.
Graph attention networks
Petar Veličković, Guillem Cucurull, Arantxa Casanova, Adriana Romero, Pietro Liò, and Yoshua Bengio. 2018 · 2018
Cited alongside, same era.
Non-local neural networks
Xiaolong Wang, Ross Girshick, Abhinav Gupta, and Kaiming He. 2018 · 2018
Cited alongside, same era.
Representation learning on graphs with jumping knowledge networks
Keyulu Xu, Chengtao Li, Yonglong Tian, Tomohiro Sonobe, Ken-ichi Kawarabayashi, and Stefanie Jegelka. 2018 · 2018
Cited alongside, same era.
Glomo: Unsupervisedly learned relational graphs as transferable representations
Zhilin Yang, Jake Zhao, Bhuwan Dhingra, Kaiming He, William W. Cohen, Ruslan Salakhutdinov, and Yann LeCun. 2018 · 2018
Cited alongside, same era.
Deep graph infomax
Petar Veličković, William Fedus, William L. Hamilton, Pietro Liò, Yoshua Bengio, and R Devon Hjelm. 2019 · 2019
Later among the works it cites.
Simplifying graph convolutional networks
Felix Wu, Amauri Souza, Tianyi Zhang, Christopher Fifty, Tao Yu, and Kilian Weinberger. 2019 · 2019
Later among the works it cites.
Sparse graph attention networks
Yang Ye and Shihao Ji. 2019 · 2019
Later among the works it cites.
Bp-transformer: Modelling long-range context via binary partitioning
Zihao Ye, Qipeng Guo, Quan Gan, Xipeng Qiu, and Zheng Zhang. 2019 · 2019
Later among the works it cites.
Reformer: The efficient transformer
Nikita Kitaev, Lukasz Kaiser, and Anselm Levskaya. 2020 · 2020
Closest in time.
Geom-gcn: Geometric graph convolutional networks
Hongbin Pei, Bingzhe Wei, Kevin Chen-Chuan Chang, Yu Lei, and Bo Yang. 2020 · 2020
Closest in time.
Unifying graph convolutional neural networks and label propagation
Hongwei Wang and Jure Leskovec. 2020 · 2020
Closest in time.
Curvature graph network
Ze Ye, Kin Sum Liu, Tengfei Ma, Jie Gao, and Chao Chen. 2020 · 2020
Closest in time.
Adaptive structural fingerprints for graph attention networks
Kai Zhang, Yaokang Zhu, Jun Wang, and Jie Zhang. 2020 · 2020
Closest in time.
Spatial transformer networks
Max Jaderberg, Karen Simonyan, Andrew Zisserman, and koray kavukcuoglu. 2015 · 2025
Closest in time.
Show, attend and tell: Neural image caption generation with visual attention
Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron Courville, Ruslan Salakhudinov, Rich Zemel, and Yoshua Bengio. 2015 · 2057
Closest in time.