Fetching the paper…
Reading the bibliography…
This paper tackles the high computational/space complexity associated with Multi-Head Self-Attention (MHSA) in vanilla vision transformers.
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner, “Gradient-based learning applied to document recognition,” Proceedings of the IEEE , vol. 86, no. 11, pp. 2278–2324, 1998
1998
Earlier work this paper cites.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “ImageNet classification with deep convolutional neural networks,” in Adv. Neural Inform. Process. Syst. , 2012, pp. 1097–1105
2012
Earlier work this paper cites.
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick, “Microsoft COCO: Common objects in context,” in Eur. Conf. Comput. Vis. , 2014, pp. 740–755
2014
Earlier work this paper cites.
K. Simonyan and A. Zisserman, “Very deep convolutional networks for large-scale image recognition,” in Int. Conf. Learn. Represent. , 2015
2015
Earlier work this paper cites.
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein et al. , “ImageNet large scale visual recognition challenge,” Int. J. Comput. Vis. , vol. 115, no. 3, pp. 211–252, 2015
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich, “Going deeper with convolutions,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2015, pp. 1–9
2015
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2016, pp. 770–778
2016
Earlier work this paper cites.
S. Ren, K. He, R. Girshick, and J. Sun, “Faster R-CNN: Towards real-time object detection with region proposal networks,” IEEE Trans. Pattern Anal. Mach. Intell. , vol. 39, no. 6, pp. 1137–1149, 2016
2016
Earlier work this paper cites.
C. Szegedy, V. Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna, “Rethinking the inception architecture for computer vision,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2016, pp. 2818–2826
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
K. He, G. Gkioxari, P. Dollár, and R. Girshick, “Mask R-CNN,” in Int. Conf. Comput. Vis. , 2017, pp. 2961–2969
2017
Earlier work this paper cites.
H. Zhao, J. Shi, X. Qi, X. Wang, and J. Jia, “Pyramid scene parsing network,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2017, pp. 2881–2890
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” in Adv. Neural Inform. Process. Syst. , 2017, pp. 6000–6010
2017
Earlier work this paper cites.
C. Szegedy, S. Ioffe, V. Vanhoucke, and A. Alemi, “Inception-v4, Inception-ResNet and the impact of residual connections on learning,” in AAAI Conf. Artif. Intell. , 2017, pp. 4278–4284
2017
Earlier work this paper cites.
S. Xie, R. Girshick, P. Dollár, Z. Tu, and K. He, “Aggregated residual transformations for deep neural networks,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2017, pp. 1492–1500
2017
Earlier work this paper cites.
G. Huang, Z. Liu, L. Van Der Maaten, and K. Q. Weinberger, “Densely connected convolutional networks,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2017, pp. 4700–4708
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
L. Chen, H. Zhang, J. Xiao, L. Nie, J. Shao, W. Liu, and T.-S. Chua, “SCA-CNN: Spatial and channel-wise attention in convolutional networks for image captioning,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2017, pp. 5659–5667
2017
Earlier work this paper cites.
F. Wang, M. Jiang, C. Qian, S. Yang, C. Li, H. Zhang, X. Wang, and X. Tang, “Residual attention network for image classification,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2017, pp. 3156–3164
2017
Earlier work this paper cites.
——, “SGDR: Stochastic gradient descent with warm restarts,” in Int. Conf. Learn. Represent. , 2017
2017
Earlier work this paper cites.
B. Zhou, H. Zhao, X. Puig, S. Fidler, A. Barriuso, and A. Torralba, “Scene parsing through ADE20K dataset,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2017, pp. 633–641
2017
Earlier work this paper cites.
T.-Y. Lin, P. Goyal, R. Girshick, K. He, and P. Dollár, “Focal loss for dense object detection,” in Int. Conf. Comput. Vis. , 2017, pp. 2980–2988
2017
Earlier work this paper cites.
M. Sandler, A. Howard, M. Zhu, A. Zhmoginov, and L.-C. Chen, “MobileNetV2: Inverted residuals and linear bottlenecks,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2018, pp. 4510–4520
2018
Earlier work this paper cites.
X. Zhang, X. Zhou, M. Lin, and J. Sun, “ShuffleNet: An extremely efficient convolutional neural network for mobile devices,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2018, pp. 6848–6856
2018
Earlier work this paper cites.
N. Ma, X. Zhang, H.-T. Zheng, and J. Sun, “ShuffleNet V2: Practical guidelines for efficient CNN architecture design,” in Eur. Conf. Comput. Vis. , 2018, pp. 116–131
2018
Cited alongside, same era.
S. Woo, J. Park, J.-Y. Lee, and I. S. Kweon, “CBAM: Convolutional block attention module,” in Eur. Conf. Comput. Vis. , 2018, pp. 3–19
2018
Cited alongside, same era.
J. Park, S. Woo, J.-Y. Lee, and I. S. Kweon, “BAM: Bottleneck attention module,” in Brit. Mach. Vis. Conf. , 2018, p. 147
2018
Cited alongside, same era.
X. Wang, R. Girshick, A. Gupta, and K. He, “Non-local neural networks,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2018, pp. 7794–7803
2018
Cited alongside, same era.
S. Elfwing, E. Uchibe, and K. Doya, “Sigmoid-weighted linear units for neural network function approximation in reinforcement learning,” Neural Networks , vol. 107, pp. 3–11, 2018
Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo, “Swin transformer: Hierarchical vision transformer using shifted windows,” in Int. Conf. Comput. Vis. , 2021, pp. 10 012–10 022
2021
Closest in time.
W. Wang, E. Xie, X. Li, D.-P. Fan, K. Song, D. Liang, T. Lu, P. Luo, and L. Shao, “Pyramid vision transformer: A versatile backbone for dense prediction without convolutions,” in Int. Conf. Comput. Vis. , 2021, pp. 568–578
2021
Closest in time.
W. Xu, Y. Xu, T. Chang, and Z. Tu, “Co-Scale conv-attentional image transformers,” in Int. Conf. Comput. Vis. , 2021, pp. 9981–9990
2021
Closest in time.
H. Fan, B. Xiong, K. Mangalam, Y. Li, Z. Yan, J. Malik, and C. Feichtenhofer, “Multiscale vision transformers,” in Int. Conf. Comput. Vis. , 2021, pp. 6824–6835
2021
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2018
Cited alongside, same era.
H. Zhang, M. Cissé, Y. N. Dauphin, and D. Lopez-Paz, “mixup: Beyond empirical risk minimization,” in Int. Conf. Learn. Represent. , 2018
2018
Cited alongside, same era.
Y. Liu, M.-M. Cheng, X. Hu, J.-W. Bian, L. Zhang, X. Bai, and J. Tang, “Richer convolutional features for edge detection,” IEEE Trans. Pattern Anal. Mach. Intell. , vol. 41, no. 8, pp. 1939–1946, 2019
2019
Cited alongside, same era.
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of deep bidirectional transformers for language understanding,” in NAACL-HLT , 2019, pp. 4171–4186
2019
Cited alongside, same era.
Z. Dai, Z. Yang, Y. Yang, J. G. Carbonell, Q. Le, and R. Salakhutdinov, “Transformer-XL: Attentive language models beyond a fixed-length context,” in ACL , 2019, pp. 2978–2988
2019
Cited alongside, same era.
M. Tan, B. Chen, R. Pang, V. Vasudevan, M. Sandler, A. Howard, and Q. V. Le, “MnasNet: Platform-aware neural architecture search for mobile,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2019, pp. 2820–2828
2019
Cited alongside, same era.
M. Tan and Q. Le, “EfficientNet: Rethinking model scaling for convolutional neural networks,” in Int. Conf. Mach. Learn. , 2019, pp. 6105–6114
2019
Cited alongside, same era.
X. Li, W. Wang, X. Hu, and J. Yang, “Selective kernel networks,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2019, pp. 510–519
2019
Cited alongside, same era.
L. Yuan, Y. Chen, T. Wang, W. Yu, Y. Shi, Z.-H. Jiang, F. E. Tay, J. Feng, and S. Yan, “Tokens-to-token ViT: Training vision transformers from scratch on ImageNet,” in Int. Conf. Comput. Vis. , 2021, pp. 558–567
2021
Closest in time.
H. Touvron, M. Cord, A. Sablayrolles, G. Synnaeve, and H. Jégou, “Going deeper with image transformers,” in Int. Conf. Comput. Vis. , 2021, pp. 32–42
2021
Closest in time.
2021
Closest in time.
H. Touvron, M. Cord, M. Douze, F. Massa, A. Sablayrolles, and H. Jégou, “Training data-efficient image transformers & distillation through attention,” in Int. Conf. Mach. Learn. , 2021, pp. 10 347–10 357
2021
Closest in time.
A. Srinivas, T.-Y. Lin, N. Parmar, J. Shlens, P. Abbeel, and A. Vaswani, “Bottleneck transformers for visual recognition,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2021, pp. 16 519–16 529
2021
Closest in time.
I. O. Tolstikhin, N. Houlsby, A. Kolesnikov, L. Beyer, X. Zhai, T. Unterthiner, J. Yung, A. Steiner, D. Keysers, J. Uszkoreit et al. , “Mlp-Mixer: An all-MLP architecture for vision,” Adv. Neural Inform. Process. Syst. , pp. 24 261–24 272, 2021
2021
Closest in time.
H. Liu, Z. Dai, D. So, and Q. V. Le, “Pay attention to MLPs,” Adv. Neural Inform. Process. Syst. , pp. 9204–9215, 2021
2021
Closest in time.
K. Han, A. Xiao, E. Wu, J. Guo, C. Xu, and Y. Wang, “Transformer in transformer,” in Adv. Neural Inform. Process. Syst. , 2021, pp. 15 908–15 919
2021
Closest in time.
2021
Closest in time.
K. Yuan, S. Guo, Z. Liu, A. Zhou, F. Yu, and W. Wu, “Incorporating convolution designs into visual transformers,” in Int. Conf. Comput. Vis. , 2021, pp. 579–588
2021
Closest in time.
H. Wu, B. Xiao, N. Codella, M. Liu, X. Dai, L. Yuan, and L. Zhang, “CvT: Introducing convolutions to vision transformers,” in Int. Conf. Comput. Vis. , 2021, pp. 22–31
2021
Closest in time.
X. Chu, Z. Tian, Y. Wang, B. Zhang, H. Ren, X. Wei, H. Xia, and C. Shen, “Twins: Revisiting the design of spatial attention in vision transformers,” in Adv. Neural Inform. Process. Syst. , 2021, pp. 9355–9366
2021
Closest in time.
P. Zhang, X. Dai, J. Yang, B. Xiao, L. Yuan, L. Zhang, and J. Gao, “Multi-scale vision Longformer: A new vision transformer for high-resolution image encoding,” in Int. Conf. Comput. Vis. , 2021, pp. 2998–3008
2021
Closest in time.
Y. Liu, M.-M. Cheng, D.-P. Fan, L. Zhang, J.-W. Bian, and D. Tao, “Semantic edge detection with diverse deep supervision,” Int. J. Comput. Vis. , vol. 130, no. 1, pp. 179–198, 2022
2022
Closest in time.
H. Zhang, C. Wu, Z. Zhang, Y. Zhu, H. Lin, Z. Zhang, Y. Sun, T. He, J. Mueller, R. Manmatha et al. , “ResNeSt: Split-attention networks,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2022, pp. 2736–2746
2022
Closest in time.
Q. Hou, Z. Jiang, L. Yuan, M.-M. Cheng, S. Yan, and J. Feng, “Vision Permutator: A permutable MLP-like architecture for visual recognition,” IEEE Trans. Pattern Anal. Mach. Intell. , vol. 45, no. 1, pp. 1328–1334, 2022
2022
Closest in time.
Z. Wang, Y. Hao, X. Gao, H. Zhang, S. Wang, T. Mu, and X. He, “Parameterization of cross-token relations with relative positional encoding for vision MLP,” in ACM Int. Conf. Multimedia , 2022, pp. 6288–6299
2022
Closest in time.
W. Wang, E. Xie, X. Li, D.-P. Fan, K. Song, D. Liang, T. Lu, P. Luo, and L. Shao, “PVTv2: Improved baselines with pyramid vision transformer,” Computational Visual Media , vol. 8, no. 3, pp. 415–424, 2022
2022
Closest in time.
D. Bolya, C.-Y. Fu, X. Dai, P. Zhang, C. Feichtenhofer, and J. Hoffman, “Token merging: Your ViT but faster,” in Int. Conf. Learn. Represent. , 2023
2023
Closest in time.
M. Jaderberg, K. Simonyan, A. Zisserman, and K. Kavukcuoglu, “Spatial transformer networks,” in Adv. Neural Inform. Process. Syst. , 2015, pp. 2017–2025
2025
Closest in time.