Fetching the paper…
Reading the bibliography…
Window-based attention has become a popular choice in vision transformers due to its superior performance, lower computational complexity, and less memory footprint.
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . Ieee, 2009, pp. 248–255
2009
Earlier work this paper cites.
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick, “Microsoft coco: Common objects in context,” in Proceedings of the European Conference on Computer Vision (ECCV) . Springer, 2014, pp. 740–755
2014
Earlier work this paper cites.
2015
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2016, pp. 770–778
2016
Earlier work this paper cites.
F. Yu and V. Koltun, “Multi-scale context aggregation by dilated convolutions,” in International Conference on Learning Representations , 2016
2016
Earlier work this paper cites.
J. Dai, H. Qi, Y. Xiong, Y. Li, G. Zhang, H. Hu, and Y. Wei, “Deformable convolutional networks,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) , 2017, pp. 764–773
2017
Earlier work this paper cites.
B. Zhou, H. Zhao, X. Puig, S. Fidler, A. Barriuso, and A. Torralba, “Scene parsing through ade20k dataset,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2017, pp. 633–641
2017
Earlier work this paper cites.
K. He, G. Gkioxari, P. Dollár, and R. Girshick, “Mask r-cnn,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) , 2017, pp. 2961–2969
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
S. Xie, R. Girshick, P. Dollár, Z. Tu, and K. He, “Aggregated residual transformations for deep neural networks,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2017, pp. 1492–1500
2017
Earlier work this paper cites.
T. Xiao, Y. Liu, B. Zhou, Y. Jiang, and J. Sun, “Unified perceptual parsing for scene understanding,” in Proceedings of the European Conference on Computer Vision (ECCV) , 2018, pp. 418–434
2018
Earlier work this paper cites.
B. Xiao, H. Wu, and Y. Wei, “Simple baselines for human pose estimation and tracking,” in Proceedings of the European conference on computer vision (ECCV) , 2018, pp. 466–481
2018
Earlier work this paper cites.
I. Loshchilov and F. Hutter, “Decoupled weight decay regularization.” in International Conference on Learning Representations , 2018
2018
Earlier work this paper cites.
Z. Cai and N. Vasconcelos, “Cascade r-cnn: Delving into high quality object detection,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2018, pp. 6154–6162
2018
Earlier work this paper cites.
X. Zhu, H. Hu, S. Lin, and J. Dai, “Deformable convnets v2: More deformable, better results,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2019, pp. 9308–9316
2019
Earlier work this paper cites.
B. Zhou, H. Zhao, X. Puig, T. Xiao, S. Fidler, A. Barriuso, and A. Torralba, “Semantic understanding of scenes through the ade20k dataset,” International Journal of Computer Vision , vol. 127, no. 3, pp. 302–321, 2019
2019
Earlier work this paper cites.
C. Lin, M. Guo, C. Li, X. Yuan, W. Wu, J. Yan, D. Lin, and W. Ouyang, “Online hyper-parameter learning for auto-augmentation strategy,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) , 2019, pp. 6579–6588
2019
Earlier work this paper cites.
S. Yun, D. Han, S. J. Oh, S. Chun, J. Choe, and Y. Yoo, “Cutmix: Regularization strategy to train strong classifiers with localizable features,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) , 2019, pp. 6023–6032
2019
Earlier work this paper cites.
Z. Cai and N. Vasconcelos, “Cascade r-cnn: High quality object detection and instance segmentation,” IEEE Transactions on Pattern Analysis and Machine Intelligence , 2019
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
Y. Jing, X. Liu, Y. Ding, X. Wang, E. Ding, M. Song, and S. Wen, “Dynamic instance normalization for arbitrary style transfer,” in AAAI , 2020
2020
Earlier work this paper cites.
M. Contributors, “Openmmlab semantic segmentation toolbox and benchmark,” 2020
2020
Earlier work this paper cites.
J. Huang, Z. Zhu, F. Guo, and G. Huang, “The devil is in the details: Delving into unbiased data processing for human pose estimation,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2020, pp. 5700–5709
2020
Cited alongside, same era.
2020
Cited alongside, same era.
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, “An image is worth 16x16 words: Transformers for image recognition at scale,” in International Conference on Learning Representations , 2021
2021
Cited alongside, same era.
Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo, “Swin transformer: Hierarchical vision transformer using shifted windows,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) , 2021, pp. 10 012–10 022
H. Bao, L. Dong, and F. Wei, “BEiT: BERT pre-training of image transformers,” 2021
2021
Later among the works it cites.
P. Zhang, X. Dai, J. Yang, B. Xiao, L. Yuan, L. Zhang, and J. Gao, “Multi-scale vision longformer: A new vision transformer for high-resolution image encoding,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) , October 2021, pp. 2998–3008
2021
Later among the works it cites.
B. Heo, S. Yun, D. Han, S. Chun, J. Choe, and S. J. Oh, “Rethinking spatial dimensions of vision transformers,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) , 2021
2021
Later among the works it cites.
K. Han, A. Xiao, E. Wu, J. Guo, C. Xu, and Y. Wang, “Transformer in transformer,” Advances in Neural Information Processing Systems , vol. 34, 2021
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2021
Cited alongside, same era.
J. Yang, C. Li, P. Zhang, X. Dai, B. Xiao, L. Yuan, and J. Gao, “Focal attention for long-range interactions in vision transformers,” in Advances in Neural Information Processing Systems , 2021
2021
Cited alongside, same era.
2021
Cited alongside, same era.
Y. Xu, Q. ZHANG, J. Zhang, and D. Tao, “ViTAE: Vision transformer advanced by exploring intrinsic inductive bias,” in Advances in Neural Information Processing Systems , 2021
2021
Cited alongside, same era.
2021
Cited alongside, same era.
2021
Cited alongside, same era.
2021
Cited alongside, same era.
H. Touvron, M. Cord, M. Douze, F. Massa, A. Sablayrolles, and H. Jegou, “Training data-efficient image transformers; distillation through attention,” in International Conference on Machine Learning . PMLR, 2021
2021
Cited alongside, same era.
Z. Dai, H. Liu, Q. V. Le, and M. Tan, “Coatnet: Marrying convolution and attention for all data sizes,” in Advances in Neural Information Processing Systems , 2021
2021
Cited alongside, same era.
X. Chu, Z. Tian, Y. Wang, B. Zhang, H. Ren, X. Wei, H. Xia, and C. Shen, “Twins: Revisiting the design of spatial attention in vision transformers,” in Advances in Neural Information Processing Systems , 2021
2021
Later among the works it cites.
L. Yuan, Y. Chen, T. Wang, W. Yu, Y. Shi, Z.-H. Jiang, F. E. Tay, J. Feng, and S. Yan, “Tokens-to-token vit: Training vision transformers from scratch on imagenet,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) , 2021, pp. 558–567
2021
Later among the works it cites.
Y. Yuan, R. Fu, L. Huang, W. Lin, C. Zhang, X. Chen, and J. Wang, “Hrformer: High-resolution vision transformer for dense predict,” Advances in Neural Information Processing Systems , vol. 34, pp. 7281–7293, 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
Q. Zhang and Y.-B. Yang, “Rest: An efficient transformer for visual recognition,” Advances in Neural Information Processing Systems , vol. 34, 2021
2021
Later among the works it cites.
T. Xiao, M. Singh, E. Mintun, T. Darrell, P. Dollár, and R. Girshick, “Early convolutions help transformers see better,” Advances in Neural Information Processing Systems , vol. 34, pp. 30 392–30 400, 2021
2021
Later among the works it cites.
Y. Li, H. Mao, R. Girshick, and K. He, “Exploring plain vision transformer backbones for object detection,” in Computer Vision–ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23–27, 2022, Proceedings, Part IX . Springer, 2022, pp. 280–296
2022
Later among the works it cites.
2022
Later among the works it cites.
Z. Yang, D. Liu, C. Wang, J. Yang, and D. Tao, “Modeling image composition for complex scene generation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 7764–7773
2022
Later among the works it cites.
2022
Later among the works it cites.
S. Wu, T. Wu, H. Tan, and G. Guo, “Pale transformer: A general vision transformer backbone with pale-shaped attention,” in Proceedings of the AAAI Conference on Artificial Intelligence , 2022
2022
Later among the works it cites.
W. Wang, L. Yao, L. Chen, B. Lin, D. Cai, X. He, and W. Liu, “Crossformer: A versatile vision transformer hinging on cross-scale attention,” in International Conference on Learning Representations , 2022
2022
Later among the works it cites.
2022
Later among the works it cites.
P. Ren, C. Li, G. Wang, Y. Xiao, Q. Du, X. Liang, and X. Chang, “Beyond fixation: Dynamic window visual transformer,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 11 987–11 997
2022
Later among the works it cites.
C.-F. Chen, R. Panda, and Q. Fan, “Regionvit: Regional-to-local attention for vision transformers,” in International Conference on Learning Representations , 2022
2022
Later among the works it cites.
Z. Xia, X. Pan, S. Song, L. E. Li, and G. Huang, “Vision transformer with deformable attention,” 2022
2022
Later among the works it cites.