Fetching the paper…
Reading the bibliography…
Capturing long-range dependencies while preserving high-resolution visual representations is crucial for dense prediction tasks such as human pose estimation.
K. Li, S. Wang, X. Zhang, Y. Xu, W. Xu, and Z. Tu, “Pose recognition with cascade transformers,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2021, pp. 1944–1953
1953
Earlier work this paper cites.
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” in 2009 IEEE conference on computer vision and pattern recognition . Ieee, 2009, pp. 248–255
2009
Earlier work this paper cites.
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick, “Microsoft coco: Common objects in context,” in Computer Vision–ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part V 13 , 2014, pp. 740–755
2014
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 770–778
2016
Earlier work this paper cites.
J. Hu, L. Shen, and G. Sun, “Squeeze-and-Excitation networks,” in CVPR , 2018, pp. 7132–7141
2018
Earlier work this paper cites.
J. Wang, K. Sun, T. Cheng, B. Jiang, C. Deng, Y. Zhao, D. Liu, Y. Mu, M. Tan, X. Wang et al. , “Deep high-resolution representation learning for visual recognition,” IEEE transactions on pattern analysis and machine intelligence , vol. 43, no. 10, pp. 3349–3364, 2020
2020
Earlier work this paper cites.
M. Contributors, “Openmmlab pose estimation toolbox and benchmark,” https://github.com/open-mmlab/mmpose , 2020
2020
Earlier work this paper cites.
Y. Yuan, R. Fu, L. Huang, W. Lin, C. Zhang, X. Chen, and J. Wang, “Hrformer: high-resolution transformer for dense prediction,” in Proceedings of the 35th International Conference on Neural Information Processing Systems , 2021, pp. 7281–7293
2021
Earlier work this paper cites.
Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo, “Swin transformer: Hierarchical vision transformer using shifted windows,” in Proceedings of the IEEE/CVF international conference on computer vision , 2021, pp. 10 012–10 022
2021
Earlier work this paper cites.
H. Touvron, M. Cord, M. Douze, F. Massa, A. Sablayrolles, and H. Jégou, “Training data-efficient image transformers & distillation through attention,” in International conference on machine learning . PMLR, 2021, pp. 10 347–10 357
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
C. Yu, B. Xiao, C. Gao, L. Yuan, L. Zhang, N. Sang, and J. Wang, “Lite-hrnet: A lightweight high-resolution network,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2021, pp. 10 440–10 450
2021
Earlier work this paper cites.
S. Yang, Z. Quan, M. Nie, and W. Yang, “Transpose: Keypoint localization via transformer,” in Proceedings of the IEEE/CVF international conference on computer vision , 2021, pp. 11 802–11 812
2021
Earlier work this paper cites.
Y. Li, S. Zhang, Z. Wang, S. Yang, W. Yang, S.-T. Xia, and E. Zhou, “Tokenpose: Learning keypoint tokens for human pose estimation,” in Proceedings of the IEEE/CVF International conference on computer vision , 2021, pp. 11 313–11 322
2021
Earlier work this paper cites.
T. Xiao, Y. Liu, B. Zhou, Y. Jiang, and J. Sun, “Unified perceptual parsing for scene understanding,” in Proceedings of the European conference on computer vision (ECCV) , 2018, pp. 418–434
2021
Earlier work this paper cites.
Z. Liu, H. Mao, C.-Y. Wu, C. Feichtenhofer, T. Darrell, and S. Xie, “A convnet for the 2020s,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2022, pp. 11 976–11 986
2022
Earlier work this paper cites.
2022
Cited alongside, same era.
2022
Cited alongside, same era.
E. Nguyen, K. Goel, A. Gu, G. Downs, P. Shah, T. Dao, S. Baccus, and C. Ré, “S4nd: Modeling images and videos as multidimensional signals with state spaces,” Advances in neural information processing systems , vol. 35, pp. 2846–2861, 2022
2022
Cited alongside, same era.
Q. Li, Z. Zhang, F. Xiao, F. Zhang, and B. Bhanu, “Dite-hrnet: Dynamic lightweight high-resolution network for human pose estimation,” in Proceedings of the Thirty-First International Joint Conference on Artificial Intelligence, IJCAI 2022, Vienna, Austria, 23-29 July 2022 , L. D. Raedt, Ed. ijcai.org, 2022, pp. 1095–1101
Y. Liu, Y. Tian, Y. Zhao, H. Yu, L. Xie, Y. Wang, Q. Ye, J. Jiao, and Y. Liu, “Vmamba: Visual state space model,” Advances in neural information processing systems , vol. 37, pp. 103 031–103 063, 2024
2024
Closest in time.
T. Huang, X. Pei, S. You, F. Wang, C. Qian, and C. Xu, “Localmamba: Visual state space model with windowed selective scan,” in European Conference on Computer Vision . Springer, 2024, pp. 12–22
2024
Closest in time.
K. Li, X. Li, Y. Wang, Y. He, Y. Wang, L. Wang, and Y. Qiao, “Videomamba: State space model for efficient video understanding,” in European Conference on Computer Vision , 2024, pp. 237–255
2024
Closest in time.
Y. Xiong, Z. Li, Y. Chen, F. Wang, X. Zhu, J. Luo, W. Wang, T. Lu, H. Li, Y. Qiao et al. , “Efficient deformable convnets: Rethinking dynamic and sparse operator for vision applications,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 5652–5661
2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2022
Cited alongside, same era.
W. Wang, E. Xie, X. Li, D.-P. Fan, K. Song, D. Liang, T. Lu, P. Luo, and L. Shao, “Pvt v2: Improved baselines with pyramid vision transformer,” Computational Visual Media , vol. 8, no. 3, pp. 415–424, 2022
2022
Cited alongside, same era.
J. Guo, K. Han, H. Wu, Y. Tang, X. Chen, Y. Wang, and C. Xu, “Cmt: Convolutional neural networks meet vision transformers,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2022, pp. 12 175–12 185
2022
Cited alongside, same era.
J. Zhang, D. Zhang, H. Yang, Y. Liu, J. Ren, X. Xu, F. Jia, and Y. Zhang, “Mvpose: Realtime multi-person pose estimation using motion vector on mobile devices,” IEEE Transactions on Mobile Computing , vol. 22, no. 6, pp. 3508–3524, 2023
2023
Cited alongside, same era.
H. Zhang, S. Lai, Y. Wang, Z. Da, Y. Dun, and X. Qian, “Scgnet: Shifting and cascaded group network,” IEEE Transactions on Circuits and Systems for Video Technology , vol. 33, no. 9, pp. 4997–5008, 2023
2023
Cited alongside, same era.
A. Shaker, M. Maaz, H. Rasheed, S. Khan, M.-H. Yang, and F. S. Khan, “Swiftformer: Efficient additive attention for transformer-based real-time mobile vision applications,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 17 425–17 436
2023
Cited alongside, same era.
2023
Cited alongside, same era.
Y. Xu, J. Zhang, Q. Zhang, and D. Tao, “Vitpose++: Vision transformer for generic body pose estimation,” IEEE Trans. Pattern Anal. Mach. Intell. , vol. 46, no. 2, pp. 1212–1230, 2024
2024
Cited alongside, same era.
H. Zhang, L. Xu, S. Lai, W. Shao, N. Zheng, P. Luo, Y. Qiao, and K. Zhang, “Open-vocabulary animal keypoint detection with semantic-feature matching,” International Journal of Computer Vision , vol. 132, no. 12, pp. 5741–5758, 2024
2024
Cited alongside, same era.
H. Zhang, W. Shao, H. Liu, Y. Ma, P. Luo, Y. Qiao, N. Zheng, and K. Zhang, “B-avibench: Towards evaluating the robustness of large vision-language model on black-box adversarial visual-instructions,” IEEE Transactions on Information Forensics and Security , 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
S. Yun and Y. Ro, “Shvit: Single-head vision transformer with memory efficient macro design,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 5756–5767
2024
Closest in time.
X. Ma, X. Dai, J. Yang, B. Xiao, Y. Chen, Y. Fu, and L. Yuan, “Efficient modulation for vision networks,” in The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024 . OpenReview.net, 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
E. J. Roh, H. Baek, D. Kim, and J. Kim, “Fast quantum convolutional neural networks for low-complexity object detection in autonomous driving applications,” IEEE Transactions on Mobile Computing , vol. 24, no. 2, pp. 1031–1042, 2025
2025
Closest in time.
J. Gong, Y. Liu, T. Li, J. Ding, Z. Wang, and D. Jin, “STTF: A spatiotemporal transformer framework for multi-task mobile network prediction,” IEEE Transactions on Mobile Computing , vol. 24, no. 5, pp. 4072–4085, 2025
2025
Closest in time.
A. Shaker, S. T. Wasim, S. Khan, J. Gall, and F. S. Khan, “Groupmamba: Efficient group-based visual state space model,” in Proceedings of the Computer Vision and Pattern Recognition Conference , 2025, pp. 14 912–14 922
2025
Closest in time.
A. Hatamizadeh and J. Kautz, “Mambavision: A hybrid mamba-transformer vision backbone,” in Proceedings of the Computer Vision and Pattern Recognition Conference , 2025, pp. 25 261–25 270
2025
Closest in time.
W. Yu and X. Wang, “Mambaout: Do we really need mamba for vision?” in Proceedings of the Computer Vision and Pattern Recognition Conference , 2025, pp. 4484–4496
2025
Closest in time.