Fetching the paper…
Reading the bibliography…
Attention mechanisms, especially self-attention, have played an increasingly important role in deep feature representation for visual tasks.
Z. Wu, S. Song, A. Khosla, F. Yu, L. Zhang, X. Tang, and J. Xiao, “3d shapenets: A deep representation for volumetric shapes,” in CVPR , 2015, pp. 1912–1920
1920
Earlier work this paper cites.
Y. Cao, J. Xu, S. Lin, F. Wei, and H. Hu, “Gcnet: Non-local networks meet squeeze-excitation networks and beyond,” in ICCV Workshops , 2019, pp. 1971–1980
1980
Earlier work this paper cites.
B. A. Olshausen and D. J. Field, “Emergence of simple-cell receptive field properties by learning a sparse code for natural images,” Nature , vol. 381, pp. 607–609, 1996
1996
Earlier work this paper cites.
M. Aharon, M. Elad, and A. M. Bruckstein, “K-SVD: an algorithm for designing overcomplete dictionaries for sparse representation,” IEEE Trans. Signal Process. , vol. 54, no. 11, pp. 4311–4322, 2006
2006
Earlier work this paper cites.
J. Deng, W. Dong, R. Socher, L. Li, K. Li, and F. Li, “Imagenet: A large-scale hierarchical image database,” in CVPR , 2009, pp. 248–255
2009
Earlier work this paper cites.
M. Everingham, L. V. Gool, C. K. I. Williams, J. M. Winn, and A. Zisserman, “The pascal visual object classes (VOC) challenge,” Int. J. Comput. Vis. , vol. 88, no. 2, pp. 303–338, 2010
2010
Earlier work this paper cites.
V. Mnih, N. Heess, A. Graves, and K. Kavukcuoglu, “Recurrent models of visual attention,” in NIPS , 2014, pp. 2204–2212
2014
Earlier work this paper cites.
T. Lin, M. Maire, S. J. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick, “Microsoft COCO: common objects in context,” in ECCV , vol. 8693, 2014, pp. 740–755
2014
Earlier work this paper cites.
D. Bahdanau, K. Cho, and Y. Bengio, “Neural machine translation by jointly learning to align and translate,” in ICLR , 2015
2015
Earlier work this paper cites.
S. Ioffe and C. Szegedy, “Batch normalization: Accelerating deep network training by reducing internal covariate shift,” in ICML , vol. 37, 2015, pp. 448–456
2015
Earlier work this paper cites.
A. Radford, L. Metz, and S. Chintala, “Unsupervised representation learning with deep convolutional generative adversarial networks,” in ICLR (Poster) , 2016
2016
Earlier work this paper cites.
J. L. Ba, J. R. Kiros, and G. E. Hinton, “Layer normalization,” 2016
2016
Earlier work this paper cites.
M. Cordts, M. Omran, S. Ramos, T. Rehfeld, M. Enzweiler, R. Benenson, U. Franke, S. Roth, and B. Schiele, “The cityscapes dataset for semantic urban scene understanding,” in CVPR , 2016, pp. 3213–3223
2016
Earlier work this paper cites.
T. Salimans, I. J. Goodfellow, W. Zaremba, V. Cheung, A. Radford, and X. Chen, “Improved techniques for training gans,” in NIPS , 2016, pp. 2226–2234
2016
Earlier work this paper cites.
L. Yi, V. G. Kim, D. Ceylan, I. Shen, M. Yan, H. Su, C. Lu, Q. Huang, A. Sheffer, and L. J. Guibas, “A scalable active framework for region annotation in 3d shape collections,” ACM Trans. Graph. , vol. 35, no. 6, pp. 210:1–210:12, 2016
2016
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” in NIPS , 2017, pp. 5998–6008
2017
Earlier work this paper cites.
Z. Lin, M. Feng, C. N. dos Santos, M. Yu, B. Xiang, B. Zhou, and Y. Bengio, “A structured self-attentive sentence embedding,” in ICLR (Poster) , 2017
2017
Earlier work this paper cites.
E. Shelhamer, J. Long, and T. Darrell, “Fully convolutional networks for semantic segmentation,” IEEE Trans. Pattern Anal. Mach. Intell. , vol. 39, no. 4, pp. 640–651, 2017
2017
Earlier work this paper cites.
S. Ren, K. He, R. B. Girshick, and J. Sun, “Faster R-CNN: towards real-time object detection with region proposal networks,” IEEE Trans. Pattern Anal. Mach. Intell. , vol. 39, no. 6, pp. 1137–1149, 2017
2017
Earlier work this paper cites.
H. Zhao, J. Shi, X. Qi, X. Wang, and J. Jia, “Pyramid scene parsing network,” in CVPR , 2017, pp. 6230–6239
2017
Earlier work this paper cites.
X. Mao, Q. Li, H. Xie, R. Y. K. Lau, Z. Wang, and S. P. Smolley, “Least squares generative adversarial networks,” in ICCV , 2017, pp. 2813–2821
2017
Earlier work this paper cites.
C. R. Qi, H. Su, K. Mo, and L. J. Guibas, “Pointnet: Deep learning on point sets for 3d classification and segmentation,” in CVPR , 2017, pp. 77–85
2017
Earlier work this paper cites.
R. Klokov and V. S. Lempitsky, “Escape from cells: Deep kd-networks for the recognition of 3d point cloud models,” in ICCV , 2017, pp. 863–872
2017
Earlier work this paper cites.
C. R. Qi, L. Yi, H. Su, and L. J. Guibas, “Pointnet++: Deep hierarchical feature learning on point sets in a metric space,” in NIPS , 2017, pp. 5099–5108
2017
Earlier work this paper cites.
M. Heusel, H. Ramsauer, T. Unterthiner, B. Nessler, and S. Hochreiter, “Gans trained by a two time-scale update rule converge to a local nash equilibrium,” in NIPS , 2017, pp. 6626–6637
2017
Earlier work this paper cites.
X. Wang, R. B. Girshick, A. Gupta, and K. He, “Non-local neural networks,” in CVPR , 2018, pp. 7794–7803
2018
Earlier work this paper cites.
H. Hu, J. Gu, Z. Zhang, J. Dai, and Y. Wei, “Relation networks for object detection,” in CVPR , 2018, pp. 3588–3597
2018
Earlier work this paper cites.
S. Xie, S. Liu, Z. Chen, and Z. Tu, “Attentional shapecontextnet for point cloud recognition,” in CVPR , 2018, pp. 4606–4615
2018
Earlier work this paper cites.
C. Yu, J. Wang, C. Peng, C. Gao, G. Yu, and N. Sang, “Learning a discriminative feature network for semantic segmentation,” in CVPR , 2018, pp. 1857–1866
2018
Earlier work this paper cites.
H. Zhang, K. J. Dana, J. Shi, Z. Zhang, X. Wang, A. Tyagi, and A. Agrawal, “Context encoding for semantic segmentation,” in CVPR , 2018, pp. 7151–7160
2018
Earlier work this paper cites.
H. Zhao, Y. Zhang, S. Liu, J. Shi, C. C. Loy, D. Lin, and J. Jia, “Psanet: Point-wise spatial attention network for scene parsing,” in ECCV , vol. 11213, 2018, pp. 270–286
2018
Earlier work this paper cites.
X. Wei, B. Gong, Z. Liu, W. Lu, and L. Wang, “Improving the improved training of wasserstein gans: A consistency term and its dual effect,” in ICLR (Poster) , 2018
2018
Earlier work this paper cites.
T. Miyato and M. Koyama, “cgans with projection discriminator,” in ICLR (Poster) , 2018
2018
Earlier work this paper cites.
J. Li, B. M. Chen, and G. H. Lee, “So-net: Self-organizing network for point cloud analysis,” in CVPR , 2018, pp. 9397–9406
2018
Earlier work this paper cites.
M. Atzmon, H. Maron, and Y. Lipman, “Point convolutional neural networks by extension operators,” ACM Trans. Graph. , vol. 37, no. 4, pp. 71:1–71:12, 2018
2018
Cited alongside, same era.
Y. Li, R. Bu, M. Sun, W. Wu, X. Di, and B. Chen, “Pointcnn: Convolution on x-transformed points,” in NeurIPS , 2018, pp. 828–838
2018
Cited alongside, same era.
T. Le and Y. Duan, “Pointgrid: A deep network for 3d shape understanding,” in CVPR , 2018, pp. 9204–9214
2018
Cited alongside, same era.
Y. Chen, Y. Kalantidis, J. Li, S. Yan, and J. Feng, “A 2 {}^{\mbox{2}} -nets: Double attention networks,” 2018
2018
Cited alongside, same era.
J. Devlin, M. Chang, K. Lee, and K. Toutanova, “BERT: pre-training of deep bidirectional transformers for language understanding,” in NAACL-HLT , 2019, pp. 4171–4186
2019
Cited alongside, same era.
J. Lee, W. Yoon, S. Kim, D. Kim, S. Kim, C. H. So, and J. Kang, “Biobert: a pre-trained biomedical language representation model for biomedical text mining,” Bioinform. , vol. 36, no. 4, pp. 1234–1240, 2020
2020
Later among the works it cites.
N. Carion, F. Massa, G. Synnaeve, N. Usunier, A. Kirillov, and S. Zagoruyko, “End-to-end object detection with transformers,” in ECCV , vol. 12346, 2020, pp. 213–229
2020
Later among the works it cites.
M. Chen, A. Radford, R. Child, J. Wu, H. Jun, D. Luan, and I. Sutskever, “Generative pretraining from pixels,” in ICML , vol. 119, 2020, pp. 1691–1703
2020
Later among the works it cites.
H. Chen, Y. Wang, T. Guo, C. Xu, Y. Deng, Z. Liu, S. Ma, C. Xu, C. Xu, and W. Gao, “Pre-trained image processing transformer,” 2020
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Fu, J. Liu, H. Tian, Y. Li, Y. Bao, Z. Fang, and H. Lu, “Dual attention network for scene segmentation,” in CVPR , 2019, pp. 3146–3154
2019
Cited alongside, same era.
Z. Huang, X. Wang, L. Huang, C. Huang, Y. Wei, and W. Liu, “Ccnet: Criss-cross attention for semantic segmentation,” in ICCV , 2019, pp. 603–612
2019
Cited alongside, same era.
X. Li, Z. Zhong, J. Wu, Y. Yang, Z. Lin, and H. Liu, “Expectation-maximization attention networks for semantic segmentation,” in ICCV , 2019, pp. 9166–9175
2019
Cited alongside, same era.
H. Zhang, I. J. Goodfellow, D. N. Metaxas, and A. Odena, “Self-attention generative adversarial networks,” in ICML , vol. 97, 2019, pp. 7354–7363
2019
Cited alongside, same era.
Y. Yuan and J. Wang, “Ocnet: Object context network for scene parsing,” 2019
2019
Cited alongside, same era.
I. Bello, B. Zoph, Q. Le, A. Vaswani, and J. Shlens, “Attention augmented convolutional networks,” in ICCV , 2019, pp. 3285–3294
2019
Cited alongside, same era.
N. Parmar, P. Ramachandran, A. Vaswani, I. Bello, A. Levskaya, and J. Shlens, “Stand-alone self-attention in vision models,” in NeurIPS , 2019, pp. 68–80
2019
Cited alongside, same era.
2020
Later among the works it cites.
Y. Wang, Z. Xu, X. Wang, C. Shen, B. Cheng, H. Shen, and H. Xia, “End-to-end video instance segmentation with transformers,” 2020
2020
Later among the works it cites.
K. He, G. Gkioxari, P. Dollár, and R. B. Girshick, “Mask R-CNN,” IEEE Trans. Pattern Anal. Mach. Intell. , vol. 42, no. 2, pp. 386–397, 2020
2020
Later among the works it cites.
T. Lin, P. Goyal, R. B. Girshick, K. He, and P. Dollár, “Focal loss for dense object detection,” IEEE Trans. Pattern Anal. Mach. Intell. , vol. 42, no. 2, pp. 318–327, 2020
2020
Later among the works it cites.
Z. Zhong, Z. Q. Lin, R. Bidart, X. Hu, I. B. Daya, Z. Li, W. Zheng, J. Li, and A. Wong, “Squeeze-and-attention networks for semantic segmentation,” in CVPR , 2020, pp. 13 062–13 071
2020
Later among the works it cites.
X. Li, Y. Yang, Q. Zhao, T. Shen, Z. Lin, and H. Liu, “Spatial pyramid based graph reasoning for semantic segmentation,” in CVPR , 2020, pp. 8947–8956
2020
Later among the works it cites.
M. Contributors, “MMSegmentation: Openmmlab semantic segmentation toolbox and benchmark,” https://github.com/open-mmlab/mmsegmentation , 2020
2020
Later among the works it cites.
X. Yan, C. Zheng, Z. Li, S. Wang, and S. Cui, “Pointasnl: Robust point clouds processing using nonlocal neural networks with adaptive sampling,” in CVPR , 2020, pp. 5588–5597
2020
Later among the works it cites.
S.-M. Hu, D. Liang, G.-Y. Yang, G.-W. Yang, and W.-Y. Zhou, “Jittor: a novel deep learning framework with meta-operators and unified graph execution,” Information Sciences , vol. 63, no. 222103, pp. 1–21, 2020
2020
Later among the works it cites.
M. Kang and J. Park, “Contragan: Contrastive learning for conditional image generation,” in NeurIPS , 2020
2020
Later among the works it cites.
Z. Geng, M.-H. Guo, H. Chen, X. Li, K. Wei, and Z. Lin, “Is attention better than matrix decomposition?” in ICLR (Poster) , 2021. [Online]. Available: https://openreview.net/forum?id=1FvkSpWosOl
2021
Closest in time.
L. Yuan, Y. Chen, T. Wang, W. Yu, Y. Shi, F. E. H. Tay, J. Feng, and S. Yan, “Tokens-to-token vit: Training vision transformers from scratch on imagenet,” 2021
2021
Closest in time.
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, “An image is worth 16x16 words: Transformers for image recognition at scale,” in ICLR , 2021. [Online]. Available: https://openreview.net/forum?id=YicbFdNTTy
2021
Closest in time.
H. Touvron, M. Cord, M. Douze, F. Massa, A. Sablayrolles, and H. Jégou, “Training data-efficient image transformers & distillation through attention,” 2021
2021
Closest in time.
Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo, “Swin transformer: Hierarchical vision transformer using shifted windows,” in CVPR , 2021
2021
Closest in time.
H. Fan, B. Xiong, K. Mangalam, Y. Li, Z. Yan, J. Malik, and C. Feichtenhofer, “Multiscale vision transformers,” 2021
2021
Closest in time.
H. Touvron, M. Cord, M. Douze, F. Massa, A. Sablayrolles, and H. Jégou, “Training data-efficient image transformers & distillation through attention,” 2021
2021
Closest in time.
X. Chen, B. Yan, J. Zhu, D. Wang, X. Yang, and H. Lu, “Transformer tracking,” in CVPR , 2021
2021
Closest in time.
Y. Jiang, S. Chang, and Z. Wang, “Transgan: Two transformers can make one strong GAN,” 2021
2021
Closest in time.
R. Hu and A. Singh, “Transformer is all you need: Multimodal multitask learning with a unified transformer,” 2021
2021
Closest in time.
S. He, H. Luo, P. Wang, F. Wang, H. Li, and W. Jiang, “Transreid: Transformer-based object re-identification,” 2021
2021
Closest in time.
W. Liu, S. Chen, L. Guo, X. Zhu, and J. Liu, “Cptr: Full transformer network for image captioning,” 2021
2021
Closest in time.
M. Guo, J. Cai, Z. Liu, T. Mu, R. R. Martin, and S. Hu, “PCT: point cloud transformer,” Comput. Vis. Media , vol. 7, no. 2, pp. 187–199, 2021
2021
Closest in time.
X. Chen, S. Xie, and K. He, “An empirical study of training self-supervised vision transformers,” in CVPR , 2021
2021
Closest in time.
K. Han, Y. Wang, H. Chen, X. Chen, J. Guo, Z. Liu, Y. Tang, A. Xiao, C. Xu, Y. Xu, Z. Yang, Y. Zhang, and D. Tao, “A survey on visual transformer,” 2021
2021
Closest in time.
S. Khan, M. Naseer, M. Hayat, S. W. Zamir, F. S. Khan, and M. Shah, “Transformers in vision: A survey,” 2021
2021
Closest in time.
Z. Cai and N. Vasconcelos, “Cascade R-CNN: high quality object detection and instance segmentation,” IEEE Trans. Pattern Anal. Mach. Intell. , vol. 43, no. 5, pp. 1483–1498, 2021
2021
Closest in time.
K. Choromanski, V. Likhosherstov, D. Dohan, X. Song, A. Gane, T. Sarlós, P. Hawkins, J. Davis, A. Mohiuddin, L. Kaiser, D. Belanger, L. Colwell, and A. Weller, “Rethinking attention with performers,” in ICLR , 2021. [Online]. Available: https://openreview.net/forum?id=Ua6zuk0WRH
2021
Closest in time.
K. Xu, J. Ba, R. Kiros, K. Cho, A. C. Courville, R. Salakhutdinov, R. S. Zemel, and Y. Bengio, “Show, attend and tell: Neural image caption generation with visual attention,” in ICML , vol. 37, 2015, pp. 2048–2057
2057
Closest in time.