Fetching the paper…
Reading the bibliography…
Transformers have shown superior performance on various vision tasks.
B. T. Polyak and A. B. Juditsky, “Acceleration of stochastic approximation by averaging,” SIAM journal on control and optimization , vol. 30, no. 4, pp. 838–855, 1992
1992
Earlier work this paper cites.
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” in CVPR . Ieee, 2009, pp. 248–255
2009
Earlier work this paper cites.
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick, “Microsoft coco: Common objects in context,” in ECCV . Springer, 2014, pp. 740–755
2014
Earlier work this paper cites.
N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov, “Dropout: A simple way to prevent neural networks from overfitting,” JMLR , 2014
2014
Earlier work this paper cites.
S. Ioffe and C. Szegedy, “Batch normalization: Accelerating deep network training by reducing internal covariate shift,” in ICML , 2015
2015
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in CVPR , 2016
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
G. Huang, Y. Sun, Z. Liu, D. Sedra, and K. Q. Weinberger, “Deep networks with stochastic depth,” in ECCV . Springer, 2016, pp. 646–661
2016
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in NeurIPS , 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
J. Dai, H. Qi, Y. Xiong, Y. Li, G. Zhang, H. Hu, and Y. Wei, “Deformable convolutional networks,” in ICCV , 2017
2017
Earlier work this paper cites.
T.-Y. Lin, P. Dollár, R. Girshick, K. He, B. Hariharan, and S. Belongie, “Feature pyramid networks for object detection,” in CVPR , 2017, pp. 2117–2125
2017
Earlier work this paper cites.
T.-Y. Lin, P. Goyal, R. Girshick, K. He, and P. Dollár, “Focal loss for dense object detection,” in ICCV , 2017, pp. 2980–2988
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
K. He, G. Gkioxari, P. Dollár, and R. Girshick, “Mask r-cnn,” in ICCV , 2017, pp. 2961–2969
2017
Earlier work this paper cites.
M. Sandler, A. Howard, M. Zhu, A. Zhmoginov, and L.-C. Chen, “Mobilenetv2: Inverted residuals and linear bottlenecks,” in CVPR , 2018
2018
Earlier work this paper cites.
G. Huang, S. Liu, L. Van der Maaten, and K. Q. Weinberger, “Condensenet: An efficient densenet using learned group convolutions,” in CVPR , 2018
2018
Earlier work this paper cites.
Z. Cai and N. Vasconcelos, “Cascade r-cnn: Delving into high quality object detection,” in CVPR , 2018, pp. 6154–6162
2018
Earlier work this paper cites.
T. Xiao, Y. Liu, B. Zhou, Y. Jiang, and J. Sun, “Unified perceptual parsing for scene understanding,” in ECCV , 2018
2018
Earlier work this paper cites.
G. Huang, Z. Liu, G. Pleiss, L. Van Der Maaten, and K. Weinberger, “Convolutional networks with dense connectivity,” IEEE TPAMI , 2019
2019
Earlier work this paper cites.
A. Howard, M. Sandler, G. Chu, L.-C. Chen, B. Chen, M. Tan, W. Wang, Y. Zhu, R. Pang, V. Vasudevan et al. , “Searching for mobilenetv3,” in ICCV , 2019
2019
Earlier work this paper cites.
Y. Cao, J. Xu, S. Lin, F. Wei, and H. Hu, “Gcnet: Non-local networks meet squeeze-excitation networks and beyond,” in ICCVW , 2019
2019
Earlier work this paper cites.
B. Zhou, H. Zhao, X. Puig, T. Xiao, S. Fidler, A. Barriuso, and A. Torralba, “Semantic understanding of scenes through the ade20k dataset,” IJCV , vol. 127, no. 3, pp. 302–321, 2019
2019
Earlier work this paper cites.
X. Zhu, H. Hu, S. Lin, and J. Dai, “Deformable convnets v2: More deformable, better results,” in CVPR , 2019
2019
Earlier work this paper cites.
S. Yun, D. Han, S. J. Oh, S. Chun, J. Choe, and Y. Yoo, “Cutmix: Regularization strategy to train strong classifiers with localizable features,” in ICCV , 2019, pp. 6023–6032
2019
Earlier work this paper cites.
A. Kirillov, R. Girshick, K. He, and P. Dollár, “Panoptic feature pyramid networks,” in CVPR , 2019, pp. 6399–6408
2019
Earlier work this paper cites.
M. Tan and Q. Le, “Efficientnet: Rethinking model scaling for convolutional neural networks,” in ICML , 2019
2019
Earlier work this paper cites.
P. Ramachandran, N. Parmar, A. Vaswani, I. Bello, A. Levskaya, and J. Shlens, “Stand-alone self-attention in vision models,” in NeurIPS , 2019
2019
Earlier work this paper cites.
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly et al. , “An image is worth 16x16 words: Transformers for image recognition at scale,” in ICLR , 2020
2020
Earlier work this paper cites.
I. Radosavovic, R. P. Kosaraju, R. Girshick, K. He, and P. Dollár, “Designing network design spaces,” in CVPR , 2020
2020
Earlier work this paper cites.
J. Wang, K. Sun, T. Cheng, B. Jiang, C. Deng, Y. Zhao, D. Liu, Y. Mu, M. Tan, X. Wang et al. , “Deep high-resolution representation learning for visual recognition,” IEEE TPAMI , vol. 43, no. 10, pp. 3349–3364, 2020
2020
Earlier work this paper cites.
H. Zhao, J. Jia, and V. Koltun, “Exploring self-attention for image recognition,” in CVPR , 2020
2020
Earlier work this paper cites.
N. Carion, F. Massa, G. Synnaeve, N. Usunier, A. Kirillov, and S. Zagoruyko, “End-to-end object detection with transformers,” in ECCV . Springer, 2020, pp. 213–229
2020
Earlier work this paper cites.
E. D. Cubuk, B. Zoph, J. Shlens, and Q. V. Le, “Randaugment: Practical automated data augmentation with a reduced search space,” in CVPRW , 2020, pp. 702–703
2020
Earlier work this paper cites.
Z. Zhong, L. Zheng, G. Kang, S. Li, and Y. Yang, “Random erasing data augmentation,” in AAAI , 2020
2020
Earlier work this paper cites.
Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo, “Swin transformer: Hierarchical vision transformer using shifted windows,” in ICCV , 2021
2021
Earlier work this paper cites.
W. Wang, E. Xie, X. Li, D.-P. Fan, K. Song, D. Liang, T. Lu, P. Luo, and L. Shao, “Pyramid vision transformer: A versatile backbone for dense prediction without convolutions,” in ICCV , 2021
2021
Earlier work this paper cites.
L. Yang, H. Jiang, R. Cai, Y. Wang, S. Song, G. Huang, and Q. Tian, “Condensenet v2: Sparse feature reactivation for deep networks,” in CVPR , 2021
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
J. Yang, C. Li, P. Zhang, X. Dai, B. Xiao, L. Yuan, and J. Gao, “Focal attention for long-range interactions in vision transformers,” in NeurIPS , 2021
2021
Earlier work this paper cites.
Y. Yuan, R. Fu, L. Huang, W. Lin, C. Zhang, X. Chen, and J. Wang, “Hrformer: High-resolution vision transformer for dense predict,” in NeurIPS , 2021
2021
Earlier work this paper cites.
H. Fan, B. Xiong, K. Mangalam, Y. Li, Z. Yan, J. Malik, and C. Feichtenhofer, “Multiscale vision transformers,” in ICCV , 2021
2021
Earlier work this paper cites.
X. Zhu, W. Su, L. Lu, B. Li, X. Wang, and J. Dai, “Deformable detr: Deformable transformers for end-to-end object detection,” in ICLR , 2021
2021
Earlier work this paper cites.
Z. Chen, Y. Zhu, C. Zhao, G. Hu, W. Zeng, J. Wang, and M. Tang, “Dpt: Deformable patch-based transformer for visual recognition,” in ACM MM , 2021
2021
Earlier work this paper cites.
X. Yue, S. Sun, Z. Kuang, M. Wei, P. Torr, W. Zhang, and D. Lin, “Vision transformer with progressive sampling,” in ICCV , 2021
2021
Cited alongside, same era.
2021
Cited alongside, same era.
Y. Xu, Q. Zhang, J. Zhang, and D. Tao, “Vitae: Vision transformer advanced by exploring intrinsic inductive bias,” in NeurIPS , 2021
2021
Cited alongside, same era.
2021
Cited alongside, same era.
A. Jaegle, F. Gimeno, A. Brock, O. Vinyals, A. Zisserman, and J. Carreira, “Perceiver: General perception with iterative attention,” in ICML , 2021
K. He, X. Chen, S. Xie, Y. Li, P. Dollár, and R. Girshick, “Masked autoencoders are scalable vision learners,” in CVPR , 2022
2022
Later among the works it cites.
J. Zhou, C. Wei, H. Wang, W. Shen, C. Xie, A. Yuille, and T. Kong, “Image BERT pre-training with online tokenizer,” in ICLR , 2022
2022
Later among the works it cites.
C. Wei, H. Fan, S. Xie, C.-Y. Wu, A. Yuille, and C. Feichtenhofer, “Masked feature prediction for self-supervised visual pre-training,” in CVPR , 2022
2022
Later among the works it cites.
M. Wortsman, G. Ilharco, S. Y. Gadre, R. Roelofs, R. Gontijo-Lopes, A. S. Morcos, H. Namkoong, A. Farhadi, Y. Carmon, S. Kornblith et al. , “Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time,” in ICML , 2022
2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2021
Cited alongside, same era.
H. Touvron, M. Cord, M. Douze, F. Massa, A. Sablayrolles, and H. Jegou, “Training data-efficient image transformers & distillation through attention,” in ICML , vol. 139, July 2021, pp. 10 347–10 357
2021
Cited alongside, same era.
H. Touvron, M. Cord, A. Sablayrolles, G. Synnaeve, and H. Jégou, “Going deeper with image transformers,” in ICCV , 2021
2021
Cited alongside, same era.
2021
Cited alongside, same era.
2021
Cited alongside, same era.
X. Chen, S. Xie, and K. He, “An empirical study of training self-supervised vision transformers,” in ICCV , 2021
2021
Cited alongside, same era.
M. Caron, H. Touvron, I. Misra, H. Jégou, J. Mairal, P. Bojanowski, and A. Joulin, “Emerging properties in self-supervised vision transformers,” in ICCV , 2021
2021
Cited alongside, same era.
H. Bao, L. Dong, S. Piao, and F. Wei, “Beit: Bert pre-training of image transformers,” in ICLR , 2021
2021
Cited alongside, same era.
2022
Later among the works it cites.
L. Meng, H. Li, B.-C. Chen, S. Lan, Z. Wu, Y.-G. Jiang, and S.-N. Lim, “Adavit: Adaptive vision transformers for efficient image recognition,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 12 309–12 318
2022
Later among the works it cites.
H. Yin, A. Vahdat, J. Alvarez, A. Mallya, J. Kautz, and P. Molchanov, “A-ViT: Adaptive tokens for efficient vision transformer,” in CVPR , 2022
2022
Later among the works it cites.
Y. Liang, C. GE, Z. Tong, Y. Song, J. Wang, and P. Xie, “EVit: Expediting vision transformers via token reorganizations,” in ICLR , 2022
2022
Later among the works it cites.
S. Tang, J. Zhang, S. Zhu, and P. Tan, “Quadtree attention for vision transformers,” in ICLR , 2022
2022
Later among the works it cites.
2022
Later among the works it cites.
Q. Zhang, Y. Xu, J. Zhang, and D. Tao, “Vsa: Learning varied-size window attention in vision transformers,” in ECCV , 2022
2022
Later among the works it cites.
J. Guo, K. Han, H. Wu, Y. Tang, X. Chen, Y. Wang, and C. Xu, “Cmt: Convolutional neural networks meet vision transformers,” in CVPR , 2022
2022
Later among the works it cites.
X. Pan, C. Ge, R. Lu, S. Song, G. Chen, Z. Huang, and G. Huang, “On the integration of self-attention and convolution,” in CVPR , 2022
2022
Later among the works it cites.
Q. Han, Z. Fan, Q. Dai, L. Sun, M.-M. Cheng, J. Liu, and J. Wang, “On the connection between local attention and dynamic depth-wise convolution,” in ICLR , 2022
2022
Later among the works it cites.
M. Arar, A. Shamir, and A. H. Bermano, “Learned queries for efficient local attention,” in CVPR , 2022
2022
Later among the works it cites.
2022
Later among the works it cites.
Z. Liu, H. Mao, C.-Y. Wu, C. Feichtenhofer, T. Darrell, and S. Xie, “A convnet for the 2020s,” in CVPR , 2022
2022
Later among the works it cites.
X. Ding, X. Zhang, J. Han, and G. Ding, “Scaling up your kernels to 31x31: Revisiting large kernel design in cnns,” in CVPR , 2022
2022
Later among the works it cites.
Y. Rao, W. Zhao, Y. Tang, J. Zhou, S. N. Lim, and J. Lu, “Hornet: Efficient high-order spatial interactions with recursive gated convolutions,” in NeurIPS , 2022
2022
Later among the works it cites.
W. Wang, J. Dai, Z. Chen, Z. Huang, Z. Li, X. Zhu, X. Hu, T. Lu, L. Lu, H. Li et al. , “Internimage: Exploring large-scale vision foundation models with deformable convolutions,” in CVPR , 2022
2022
Later among the works it cites.
G. Huang, Y. Wang, K. Lv, H. Jiang, W. Huang, P. Qi, and S. Song, “Glance and focus networks for dynamic visual recognition,” IEEE TPAMI , 2022
2022
Later among the works it cites.
Z. Pan, J. Cai, and B. Zhuang, “Fast vision transformers with hilo attention,” in NeurIPS , 2022
2022
Later among the works it cites.
2023
Closest in time.
H. Huang, X. Zhou, J. Cao, R. He, and T. Tan, “Vision transformer with super token sampling,” in CVPR , 2023
2023
Closest in time.
L. Zhu, X. Wang, Z. Ke, W. Zhang, and R. Lau, “Biformer: Vision transformer with bi-level routing attention,” in CVPR , 2023
2023
Closest in time.
Q. Zhang, Y. Xu, J. Zhang, and D. Tao, “Vitaev2: Vision transformer advanced by exploring inductive bias for image recognition and beyond,” International Journal of Computer Vision , pp. 1–22, 2023
2023
Closest in time.
2023
Closest in time.
Y. Li, H. Fan, R. Hu, C. Feichtenhofer, and K. He, “Scaling language-image pre-training via masking,” in CVPR , 2023
2023
Closest in time.
W. Su, X. Zhu, C. Tao, L. Lu, B. Li, G. Huang, Y. Qiao, X. Wang, J. Zhou, and J. Dai, “Towards all-in-one pre-training via maximizing multi-modal mutual information,” in CVPR , 2023
2023
Closest in time.
Y. Fang, W. Wang, B. Xie, Q. Sun, L. Wu, X. Wang, T. Huang, X. Wang, and Y. Cao, “Eva: Exploring the limits of masked visual representation learning at scale,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 19 358–19 369
2023
Closest in time.
M. Assran, Q. Duval, I. Misra, P. Bojanowski, P. Vincent, M. Rabbat, Y. LeCun, and N. Ballas, “Self-supervised learning from images with a joint-embedding predictive architecture,” in CVPR , 2023
2023
Closest in time.
M. Dehghani, J. Djolonga, B. Mustafa, P. Padlewski, J. Heek, J. Gilmer, A. P. Steiner, M. Caron, R. Geirhos, I. Alabdulmohsin et al. , “Scaling vision transformers to 22 billion parameters,” in ICML , 2023
2023
Closest in time.
Y. Rao, Z. Liu, W. Zhao, J. Zhou, and J. Lu, “Dynamic spatial sparsification for efficient vision transformers and convolutional neural networks,” IEEE TPAMI , 2023
2023
Closest in time.
D. Bolya, C.-Y. Fu, X. Dai, P. Zhang, C. Feichtenhofer, and J. Hoffman, “Token merging: Your ViT but faster,” in ICLR , 2023
2023
Closest in time.
M. Ding, Y. Shen, L. Fan, Z. Chen, Z. Chen, P. Luo, J. B. Tenenbaum, and C. Gan, “Visual dependency transformers: Dependency tree emerges from reversed attention,” in CVPR , 2023
2023
Closest in time.
2023
Closest in time.
X. Pan, T. Ye, Z. Xia, S. Song, and G. Huang, “Slide-transformer: Hierarchical vision transformer with local self-attention,” in CVPR , 2023
2023
Closest in time.
X. Chu, Z. Tian, B. Zhang, X. Wang, and C. Shen, “Conditional positional encodings for vision transformers,” in ICLR , 2023
2023
Closest in time.
A. Hassani, S. Walton, J. Li, S. Li, and H. Shi, “Neighborhood attention transformer,” in CVPR , 2023
2023
Closest in time.
S. Liu, T. Chen, X. Chen, X. Chen, Q. Xiao, B. Wu, T. Kärkkäinen, M. Pechenizkiy, D. C. Mocanu, and Z. Wang, “More convnets in the 2020s: Scaling up kernels beyond 51x51 using sparsity,” in ICLR , 2023
2023
Closest in time.
Y. Rao, W. Zhao, Z. Zhu, J. Zhou, and J. Lu, “Gfnet: Global filter networks for visual recognition,” IEEE TPAMI , 2023
2023
Closest in time.
2023
Closest in time.
Y. Pu, Y. Wang, Z. Xia, Y. Han, Y. Wang, W. Gan, Z. Wang, S. Song, and G. Huang, “Adaptive rotated convolution for rotated object detection,” in ICCV , 2023
2023
Closest in time.
Y. Han, D. Han, Z. Liu, Y. Wang, X. Pan, Y. Pu, C. Deng, J. Feng, S. Song, and G. Huang, “Dynamic perceiver for efficient visual recognition,” in ICCV , 2023
2023
Closest in time.
A. Hatamizadeh, H. Yin, G. Heinrich, J. Kautz, and P. Molchanov, “Global context vision transformers,” in ICML , 2023
2023
Closest in time.