Fetching the paper…
Reading the bibliography…
While originally designed for natural language processing tasks, the self-attention mechanism has recently taken various computer vision areas by storm.
F. Rosenblatt, “The perceptron: a probabilistic model for information storage and organization in the brain.” Psychological review , vol. 65, no. 6, p. 386, 1958
1958
Earlier work this paper cites.
A. M. Treisman and G. Gelade, “A feature-integration theory of attention,” Cognitive psychology , vol. 12, no. 1, pp. 97–136, 1980
1980
Earlier work this paper cites.
D. E. Rumelhart, G. E. Hinton, and R. J. Williams, “Learning internal representations by error propagation,” California Univ San Diego La Jolla Inst for Cognitive Science, Tech. Rep., 1985
1985
Earlier work this paper cites.
Y. LeCun, B. Boser, J. S. Denker, D. Henderson, R. E. Howard, W. Hubbard, and L. D. Jackel, “Backpropagation applied to handwritten zip code recognition,” Neural computation , vol. 1, no. 4, pp. 541–551, 1989
1989
Earlier work this paper cites.
B. T. Polyak and A. B. Juditsky, “Acceleration of stochastic approximation by averaging,” SIAM journal on control and optimization , vol. 30, no. 4, pp. 838–855, 1992
1992
Earlier work this paper cites.
J. K. Tsotsos, S. M. Culhane, W. Y. K. Wai, Y. Lai, N. Davis, and F. Nuflo, “Modeling visual attention via selective tuning,” Artificial intelligence , vol. 78, no. 1-2, pp. 507–545, 1995
1995
Earlier work this paper cites.
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner, “Gradient-based learning applied to document recognition,” Proceedings of the IEEE , vol. 86, no. 11, pp. 2278–2324, 1998
1998
Earlier work this paper cites.
J. P. Gottlieb, M. Kusunoki, and M. E. Goldberg, “The representation of visual salience in monkey parietal cortex,” Nature , vol. 391, no. 6666, pp. 481–484, 1998
1998
Earlier work this paper cites.
J. M. Wolfe and T. S. Horowitz, “What attributes guide the deployment of visual attention and how do they do it?” Nature reviews neuroscience , vol. 5, no. 6, pp. 495–501, 2004
2004
Earlier work this paper cites.
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” in 2009 IEEE conference on computer vision and pattern recognition . Ieee, 2009, pp. 248–255
2009
Earlier work this paper cites.
P. Welinder, S. Branson, T. Mita, C. Wah, F. Schroff, S. Belongie, and P. Perona, “Caltech-ucsd birds 200,” 2010
2010
Earlier work this paper cites.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,” Adv. Neural Inform. Process. Syst. , vol. 25, pp. 1097–1105, 2012
2012
Earlier work this paper cites.
2013
Earlier work this paper cites.
C. Yang, L. Zhang, H. Lu, X. Ruan, and M.-H. Yang, “Saliency detection via graph-based manifold ranking,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2013, pp. 3166–3173
2013
Earlier work this paper cites.
2014
Earlier work this paper cites.
M. Lin, Q. Chen, and S. Yan, “Network in network,” in Int. Conf. Learn. Represent. , 2014
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
V. Mnih, N. Heess, A. Graves et al. , “Recurrent models of visual attention,” in Adv. Neural Inform. Process. Syst. , 2014, pp. 2204–2212
2014
Earlier work this paper cites.
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick, “Microsoft coco: Common objects in context,” in Eur. Conf. Comput. Vis. Springer, 2014, pp. 740–755
2014
Earlier work this paper cites.
Y. Li, X. Hou, C. Koch, J. M. Rehg, and A. L. Yuille, “The secrets of salient object segmentation,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2014, pp. 280–287
2014
Earlier work this paper cites.
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich, “Going deeper with convolutions,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2015, pp. 1–9
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
S. Ioffe and C. Szegedy, “Batch normalization: Accelerating deep network training by reducing internal covariate shift,” in Int. Conf. Mach. Learn. PMLR, 2015, pp. 448–456
2015
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2016, pp. 770–778
2016
Earlier work this paper cites.
B. Zhou, A. Khosla, A. Lapedriza, A. Oliva, and A. Torralba, “Learning deep features for discriminative localization,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2016, pp. 2921–2929
2016
Earlier work this paper cites.
——, “Sgdr: Stochastic gradient descent with warm restarts,” arXiv preprint arXiv:1608.03983 , 2016
2016
Earlier work this paper cites.
2017
Earlier work this paper cites.
S. Xie, R. Girshick, P. Dollár, Z. Tu, and K. He, “Aggregated residual transformations for deep neural networks,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2017, pp. 1492–1500
2017
Earlier work this paper cites.
G. Huang, Z. Liu, L. Van Der Maaten, and K. Q. Weinberger, “Densely connected convolutional networks,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2017, pp. 4700–4708
2017
Earlier work this paper cites.
J. Dai, H. Qi, Y. Xiong, Y. Li, G. Zhang, H. Hu, and Y. Wei, “Deformable convolutional networks,” in Int. Conf. Comput. Vis. , 2017, pp. 764–773
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in Adv. Neural Inform. Process. Syst. , 2017, pp. 5998–6008
2017
Earlier work this paper cites.
L. Chen, H. Zhang, J. Xiao, L. Nie, J. Shao, W. Liu, and T.-S. Chua, “Sca-cnn: Spatial and channel-wise attention in convolutional networks for image captioning,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2017, pp. 5659–5667
2017
Earlier work this paper cites.
F. Wang, M. Jiang, C. Qian, S. Yang, C. Li, H. Zhang, X. Wang, and X. Tang, “Residual attention network for image classification,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2017, pp. 3156–3164
2017
Earlier work this paper cites.
R. R. Selvaraju, M. Cogswell, A. Das, R. Vedantam, D. Parikh, and D. Batra, “Grad-cam: Visual explanations from deep networks via gradient-based localization,” in Int. Conf. Comput. Vis. , 2017, pp. 618–626
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” 2017
2017
Earlier work this paper cites.
T.-Y. Lin, P. Goyal, R. Girshick, K. He, and P. Dollár, “Focal loss for dense object detection,” in Int. Conf. Comput. Vis. , 2017, pp. 2980–2988
2017
Earlier work this paper cites.
K. He, G. Gkioxari, P. Dollar, and R. Girshick, “Mask r-cnn,” in Int. Conf. Comput. Vis. , Oct 2017
2017
Earlier work this paper cites.
L. Wang, H. Lu, Y. Wang, M. Feng, D. Wang, B. Yin, and X. Ruan, “Learning to detect salient objects with image-level supervision,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2017, pp. 136–145
2017
Earlier work this paper cites.
X. Zhang, X. Zhou, M. Lin, and J. Sun, “Shufflenet: An extremely efficient convolutional neural network for mobile devices,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2018, pp. 6848–6856
2018
Earlier work this paper cites.
J. Hu, L. Shen, and G. Sun, “Squeeze-and-excitation networks,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2018, pp. 7132–7141
2018
Earlier work this paper cites.
2018
Cited alongside, same era.
S. Woo, J. Park, J.-Y. Lee, and I. S. Kweon, “Cbam: Convolutional block attention module,” in Eur. Conf. Comput. Vis. , 2018, pp. 3–19
2018
Cited alongside, same era.
M. Sandler, A. Howard, M. Zhu, A. Zhmoginov, and L.-C. Chen, “Mobilenetv2: Inverted residuals and linear bottlenecks,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2018, pp. 4510–4520
2018
Cited alongside, same era.
H. Hu, J. Gu, Z. Zhang, J. Dai, and Y. Wei, “Relation networks for object detection,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2018, pp. 3588–3597
2018
Cited alongside, same era.
X. Wang, R. Girshick, A. Gupta, and K. He, “Non-local neural networks,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2018, pp. 7794–7803
E. Xie, W. Wang, Z. Yu, A. Anandkumar, J. M. Alvarez, and P. Luo, “Segformer: Simple and efficient design for semantic segmentation with transformers,” Adv. Neural Inform. Process. Syst. , vol. 34, 2021
2021
Later among the works it cites.
H. Touvron, M. Cord, M. Douze, F. Massa, A. Sablayrolles, and H. Jégou, “Training data-efficient image transformers & distillation through attention,” in Int. Conf. Mach. Learn. PMLR, 2021, pp. 10 347–10 357
2021
Later among the works it cites.
W. Wang, E. Xie, X. Li, D.-P. Fan, K. Song, D. Liang, T. Lu, P. Luo, and L. Shao, “Pyramid vision transformer: A versatile backbone for dense prediction without convolutions,” in Int. Conf. Comput. Vis. , 2021
2021
Later among the works it cites.
J. Yang, C. Li, P. Zhang, X. Dai, B. Xiao, L. Yuan, and J. Gao, “Focal self-attention for local-global interactions in vision transformers,” 2021
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2018
Cited alongside, same era.
2018
Cited alongside, same era.
S. Xie, S. Liu, Z. Chen, and Z. Tu, “Attentional shapecontextnet for point cloud recognition,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2018, pp. 4606–4615
2018
Cited alongside, same era.
J. Hu, L. Shen, S. Albanie, G. Sun, and A. Vedaldi, “Gather-excite: Exploiting feature context in convolutional neural networks,” Adv. Neural Inform. Process. Syst. , vol. 31, 2018
2018
Cited alongside, same era.
2018
Cited alongside, same era.
T. Xiao, Y. Liu, B. Zhou, Y. Jiang, and J. Sun, “Unified perceptual parsing for scene understanding,” in Eur. Conf. Comput. Vis. , 2018, pp. 418–434
2018
Cited alongside, same era.
B. Xiao, H. Wu, and Y. Wei, “Simple baselines for human pose estimation and tracking,” in Proceedings of the European conference on computer vision (ECCV) , 2018, pp. 466–481
2018
Cited alongside, same era.
2018
Cited alongside, same era.
A. Ali, H. Touvron, M. Caron, P. Bojanowski, M. Douze, A. Joulin, I. Laptev, N. Neverova, G. Synnaeve, J. Verbeek et al. , “Xcit: Cross-covariance image transformers,” Adv. Neural Inform. Process. Syst. , vol. 34, 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
I. Bello, W. Fedus, X. Du, E. D. Cubuk, A. Srinivas, T.-Y. Lin, J. Shlens, and B. Zoph, “Revisiting resnets: Improved training and scaling strategies,” Advances in Neural Information Processing Systems , vol. 34, 2021
2021
Later among the works it cites.
Z. Geng, M.-H. Guo, H. Chen, X. Li, K. Wei, and Z. Lin, “Is attention better than matrix decomposition?” in Int. Conf. Learn. Represent. , 2021
2021
Later among the works it cites.
M.-H. Guo, J.-X. Cai, Z.-N. Liu, T.-J. Mu, R. R. Martin, and S.-M. Hu, “Pct: Point cloud transformer,” Computational Visual Media , vol. 7, no. 2, pp. 187–199, 2021
2021
Later among the works it cites.
A. Srinivas, T.-Y. Lin, N. Parmar, J. Shlens, P. Abbeel, and A. Vaswani, “Bottleneck transformers for visual recognition,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2021, pp. 16 519–16 529
2021
Later among the works it cites.
L. Yuan, Y. Chen, T. Wang, W. Yu, Y. Shi, Z.-H. Jiang, F. E. Tay, J. Feng, and S. Yan, “Tokens-to-token vit: Training vision transformers from scratch on imagenet,” in Int. Conf. Comput. Vis. , October 2021, pp. 558–567
2021
Later among the works it cites.
2021
Later among the works it cites.
I. Bello, “Lambdanetworks: Modeling long-range interactions without attention,” in International Conference on Learning Representations , 2021
2021
Later among the works it cites.
Y. Xu, Q. Zhang, J. Zhang, and D. Tao, “Vitae: Vision transformer advanced by exploring intrinsic inductive bias,” Advances in Neural Information Processing Systems , vol. 34, pp. 28 522–28 535, 2021
2021
Later among the works it cites.
R. Liu, H. Deng, Y. Huang, X. Shi, L. Lu, W. Sun, X. Wang, J. Dai, and H. Li, “Fuseformer: Fusing fine-grained information in transformers for video inpainting,” in Int. Conf. Comput. Vis. , 2021, pp. 14 040–14 049
2021
Later among the works it cites.
H. Wu, B. Xiao, N. Codella, M. Liu, X. Dai, L. Yuan, and L. Zhang, “Cvt: Introducing convolutions to vision transformers,” in Int. Conf. Comput. Vis. , 2021, pp. 22–31
2021
Later among the works it cites.
S. Liu, L. Zhang, X. Yang, H. Su, and J. Zhu, “Query2label: A simple transformer way to multi-label classification,” 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
K. He, X. Chen, S. Xie, Y. Li, P. Dollár, and R. Girshick, “Masked autoencoders are scalable vision learners,” 2021
2021
Later among the works it cites.
Z. Qin, P. Zhang, F. Wu, and X. Li, “Fcanet: Frequency channel attention networks,” in Int. Conf. Comput. Vis. , 2021, pp. 783–792
2021
Later among the works it cites.
I. O. Tolstikhin, N. Houlsby, A. Kolesnikov, L. Beyer, X. Zhai, T. Unterthiner, J. Yung, A. Steiner, D. Keysers, J. Uszkoreit et al. , “Mlp-mixer: An all-mlp architecture for vision,” Adv. Neural Inform. Process. Syst. , vol. 34, 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
H. Liu, Z. Dai, D. So, and Q. V. Le, “Pay attention to MLPs,” in Adv. Neural Inform. Process. Syst. , A. Beygelzimer, Y. Dauphin, P. Liang, and J. W. Vaughan, Eds., 2021
2021
Later among the works it cites.
M.-H. Guo, Z.-N. Liu, T.-J. Mu, D. Liang, R. R. Martin, and S.-M. Hu, “Can attention enable mlps to catch up with cnns?” Computational Visual Media , vol. 7, no. 3, pp. 283–288, 2021
2021
Later among the works it cites.
R. Liu, Y. Li, L. Tao, D. Liang, S.-M. Hu, and H.-T. Zheng, “Are we ready for a new paradigm shift? a survey on visual deep mlp,” 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
H. Touvron, M. Cord, A. Sablayrolles, G. Synnaeve, and H. Jégou, “Going deeper with image transformers,” in Int. Conf. Comput. Vis. , 2021, pp. 32–42
2021
Later among the works it cites.
K. Han, A. Xiao, E. Wu, J. Guo, C. Xu, and Y. Wang, “Transformer in transformer,” Adv. Neural Inform. Process. Syst. , vol. 34, 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
X. Chu, Z. Tian, Y. Wang, B. Zhang, H. Ren, X. Wei, H. Xia, and C. Shen, “Twins: Revisiting the design of spatial attention in vision transformers,” Adv. Neural Inform. Process. Syst. , vol. 34, 2021
2021
Later among the works it cites.
M. Tan and Q. Le, “Efficientnetv2: Smaller models and faster training,” in International Conference on Machine Learning . PMLR, 2021, pp. 10 096–10 106
2021
Later among the works it cites.
Z. Dai, H. Liu, Q. Le, and M. Tan, “Coatnet: Marrying convolution and attention for all data sizes,” Adv. Neural Inform. Process. Syst. , vol. 34, 2021
2021
Later among the works it cites.
P. Sun, R. Zhang, Y. Jiang, T. Kong, C. Xu, W. Zhan, M. Tomizuka, L. Li, Z. Yuan, C. Wang et al. , “Sparse r-cnn: End-to-end object detection with learnable proposals,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2021, pp. 14 454–14 463
2021
Later among the works it cites.
Z. Liu, H. Mao, C.-Y. Wu, C. Feichtenhofer, T. Darrell, and S. Xie, “A convnet for the 2020s,” 2022
2022
Closest in time.
M.-H. Guo, T.-X. Xu, J.-J. Liu, Z.-N. Liu, P.-T. Jiang, T.-J. Mu, S.-H. Zhang, R. R. Martin, M.-M. Cheng, and S.-M. Hu, “Attention mechanisms in computer vision: A survey,” Computational Visual Media , pp. 1–38, 2022
2022
Closest in time.
H. Bao, L. Dong, S. Piao, and F. Wei, “BEit: BERT pre-training of image transformers,” in Int. Conf. Learn. Represent. , 2022
2022
Closest in time.
S. Liu, F. Li, H. Zhang, X. Yang, X. Qi, H. Su, J. Zhu, and L. Zhang, “DAB-DETR: Dynamic anchor boxes are better queries for DETR,” in Int. Conf. Learn. Represent. , 2022
2022
Closest in time.
B. Cheng, I. Misra, A. G. Schwing, A. Kirillov, and R. Girdhar, “Masked-attention mask transformer for universal image segmentation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 1290–1299
2022
Closest in time.
Y.-H. Wu, Y. Liu, L. Zhang, M.-M. Cheng, and B. Ren, “Edn: Salient object detection via extremely-downsampled network,” IEEE Transactions on Image Processing , vol. 31, pp. 3125–3136, 2022
2022
Closest in time.