Fetching the paper…
Reading the bibliography…
The Segment Anything Model (SAM) has gained significant attention for its impressive performance in image segmentation.
L. Lan, X. Wang, G. Hua, T. S. Huang, and D. Tao, “Semi-online multi-people tracking by re-identification,” in International Journal of Computer Vision , vol. 128, no. 7, 2020, pp. 1937–1955
1955
Earlier work this paper cites.
H. Jhuang, J. Gall, S. Zuffi, C. Schmid, and M. J. Black, “Towards understanding action recognition,” in Proceedings of the IEEE international conference on computer vision , 2013, pp. 3192–3199
2013
Earlier work this paper cites.
C. Xu, S.-H. Hsieh, C. Xiong, and J. J. Corso, “Can humans fly? action understanding with multiple classes of actors,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2015, pp. 2264–2273
2015
Earlier work this paper cites.
L. Yu, P. Poirson, S. Yang, A. C. Berg, and T. L. Berg, “Modeling context in referring expressions,” in Computer Vision–ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11-14, 2016, Proceedings, Part II 14 . Springer, 2016, pp. 69–85
2016
Earlier work this paper cites.
V. K. Nagaraja, V. I. Morariu, and L. S. Davis, “Modeling context between objects for referring expression understanding,” in European Conference on Computer Vision . Springer, 2016, pp. 792–807
2016
Earlier work this paper cites.
J. Mao, J. Huang, A. Toshev, O. Camburu, A. L. Yuille, and K. Murphy, “Generation and comprehension of unambiguous object descriptions,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2016, pp. 11–20
2016
Earlier work this paper cites.
2017
Earlier work this paper cites.
K. Gavrilyuk, A. Ghodrati, Z. Li, and C. G. Snoek, “Actor and action video segmentation from a sentence,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2018, pp. 5958–5966
2018
Earlier work this paper cites.
L. Lan, X. Wang, S. Zhang, D. Tao, W. Gao, and T. S. Huang, “Interacting tracklets for multi-object tracking,” IEEE Transactions on Image Processing , vol. 27, no. 9, pp. 4585–4597, 2018
2018
Earlier work this paper cites.
A. Khoreva, A. Rohrbach, and B. Schiele, “Video object segmentation with language referring expressions,” in Computer Vision–ACCV 2018: 14th Asian Conference on Computer Vision , 2019, pp. 123–141
2019
Earlier work this paper cites.
L. Ye, M. Rochan, Z. Liu, and Y. Wang, “Cross-modal self-attention network for referring image segmentation,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2019, pp. 10 502–10 511
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
J. Zhang and D. Tao, “Empowering things with intelligence: A survey of the progress, challenges, and opportunities in artificial intelligence of things,” IEEE Internet of Things Journal , vol. 8, no. 10, pp. 7789–7817, 2020
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu, “Exploring the limits of transfer learning with a unified text-to-text transformer,” The Journal of Machine Learning Research , vol. 21, no. 1, pp. 5485–5551, 2020
2020
Earlier work this paper cites.
T. Wolf, L. Debut, V. Sanh, J. Chaumond, C. Delangue, A. Moi, P. Cistac, T. Rault, R. Louf, M. Funtowicz et al. , “Transformers: State-of-the-art natural language processing,” in Proceedings of the 2020 conference on empirical methods in natural language processing: system demonstrations , 2020, pp. 38–45
2020
Earlier work this paper cites.
G. Luo, Y. Zhou, R. Ji, X. Sun, J. Su, C.-W. Lin, and Q. Tian, “Cascade grouped attention network for referring expression segmentation,” in Proceedings of the 28th ACM International Conference on Multimedia , 2020, pp. 1274–1282
2020
Earlier work this paper cites.
Z. Ding, T. Hui, S. Huang, S. Liu, X. Luo, J. Huang, and X. Wei, “Progressive multimodal interaction network for referring video object segmentation,” The 3rd Large-scale Video Object Segmentation Challenge , p. 7, 2021
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo, “Swin transformer: Hierarchical vision transformer using shifted windows,” in Proceedings of the IEEE international conference on computer vision , 2021, pp. 10 012–10 022
2021
Earlier work this paper cites.
H. Ding, C. Liu, S. Wang, and X. Jiang, “Vision-language transformer and query generation for referring segmentation,” in Proceedings of the IEEE international conference on computer vision , 2021, pp. 16 321–16 330
2021
Earlier work this paper cites.
H. Tan, X. Zhang, Z. Zhang, L. Lan, W. Zhang, and Z. Luo, “Nocal-siam: Refining visual features and response with advanced non-local blocks for real-time siamese tracking,” IEEE Transactions on Image Processing , vol. 30, pp. 2656–2668, 2021
2021
Earlier work this paper cites.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark et al. , “Learning transferable visual models from natural language supervision,” in International conference on machine learning , 2021, pp. 8748–8763
2021
Earlier work this paper cites.
Y. Jing, T. Kong, W. Wang, L. Wang, L. Li, and T. Tan, “Locate then segment: A strong pipeline for referring image segmentation,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2021, pp. 9858–9867
2021
Cited alongside, same era.
M. Li and L. Sigal, “Referring transformer: A one-step approach to multi-task visual grounding,” in Advances in neural information processing systems , vol. 34, 2021, pp. 19 652–19 664
2021
Cited alongside, same era.
A. Botach, E. Zheltonozhskii, and C. Baskin, “End-to-end referring video object segmentation with multimodal transformers,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2022, pp. 4985–4995
2022
Cited alongside, same era.
J. Wu, Y. Jiang, P. Sun, Z. Yuan, and P. Luo, “Language as queries for referring video object segmentation,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2022, pp. 4974–4984
2022
Cited alongside, same era.
2023
Closest in time.
2023
Closest in time.
Q. Zhang, Y. Xu, J. Zhang, and D. Tao, “Vitaev2: Vision transformer advanced by exploring inductive bias for image recognition and beyond,” in International Journal of Computer Vision , 2023, pp. 1–22
2023
Closest in time.
2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Z. Ding, T. Hui, J. Huang, X. Wei, J. Han, and S. Liu, “Language-bridged spatial-temporal interaction for referring video object segmentation,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2022, pp. 4964–4973
2022
Cited alongside, same era.
W. Zhao, K. Wang, X. Chu, F. Xue, X. Wang, and Y. You, “Modeling motion with multi-modal features for text-based video segmentation,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2022, pp. 11 737–11 746
2022
Cited alongside, same era.
S. Seo, J.-Y. Lee, and B. Han, “Urvos: Unified referring video object segmentation network with a large-scale benchmark,” in European Conference on Computer Vision , 2022, pp. 208–223
2022
Cited alongside, same era.
D. Wu, X. Dong, L. Shao, and J. Shen, “Multi-level representation learning with semantic alignment for referring video object segmentation,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2022, pp. 4996–5005
2022
Cited alongside, same era.
W. Wang, J. Zhang, Y. Cao, Y. Shen, and D. Tao, “Towards data-efficient detection transformers,” in European Conference on Computer Vision , 2022, pp. 88–105
2022
Cited alongside, same era.
Y. Xu, J. Zhang, Q. Zhang, and D. Tao, “Vitpose: Simple vision transformer baselines for human pose estimation,” in Advances in neural information processing systems , vol. 35, 2022, pp. 38 571–38 584
2022
Cited alongside, same era.
M. Lan, J. Zhang, F. He, and L. Zhang, “Siamese network with interactive transformer for video object segmentation,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 36, no. 2, 2022, pp. 1228–1236
2022
Cited alongside, same era.
K. He, X. Chen, S. Xie, Y. Li, P. Dollár, and R. Girshick, “Masked autoencoders are scalable vision learners,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2022, pp. 16 000–16 009
2022
Cited alongside, same era.
2023
Closest in time.
S. Julka and M. Granitzer, “Knowledge distillation with segment anything (sam) model for planetary geological mapping,” in International Conference on Machine Learning, Optimization, and Data Science . Springer, 2023, pp. 68–77
2023
Closest in time.
J. Cen, Z. Zhou, J. Fang, W. Shen, L. Xie, D. Jiang, X. Zhang, Q. Tian et al. , “Segment anything in 3d with nerfs,” Advances in Neural Information Processing Systems , vol. 36, pp. 25 971–25 990, 2023
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
D. Wu, T. Wang, Y. Zhang, X. Zhang, and J. Shen, “Onlinerefer: A simple online baseline for referring video object segmentation,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 2761–2770
2023
Closest in time.
B. Miao, M. Bennamoun, Y. Gao, and A. Mian, “Spectrum-guided multi-granularity referring video object segmentation,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 920–930
2023
Closest in time.
Y. Lu, R. Quan, L. Zhu, and Y. Yang, “Zero-shot video grounding with pseudo query lookup and verification,” IEEE Transactions on Image Processing , vol. 33, pp. 1643–1654, 2024
2024
Closest in time.
C. Zhao, Y. Wang, X. Jiang, Y. Shen, K. Song, D. Li, and D. Miao, “Learning domain invariant prompt for vision-language models,” IEEE Transactions on Image Processing , 2024
2024
Closest in time.
Y. Zhang, Q. Li, Y. Pan, X. Zhao, and M. Tan, “Multi-stage image-language cross-generative fusion network for video-based referring expression comprehension,” IEEE Transactions on Image Processing , 2024
2024
Closest in time.
W. Yue, J. Zhang, K. Hu, Y. Xia, J. Luo, and Z. Wang, “Surgicalsam: Efficient class promptable surgical instrument segmentation,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 38, no. 7, 2024, pp. 6890–6898
2024
Closest in time.
D. Wang, J. Zhang, B. Du, M. Xu, L. Liu, D. Tao, and L. Zhang, “Samrs: Scaling-up remote sensing segmentation dataset with segment anything model,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
Z. Xie, S. Wang, Q. Yu, X. Tan, and Y. Xie, “Csfwinformer: Cross-space-frequency window transformer for mirror detection,” IEEE Transactions on Image Processing , 2024
2024
Closest in time.
Z. Li, X. Wang, X. Liu, and J. Jiang, “Binsformer: Revisiting adaptive bins for monocular depth estimation,” IEEE Transactions on Image Processing , 2024
2024
Closest in time.
L. Ke, M. Ye, M. Danelljan, Y.-W. Tai, C.-K. Tang, F. Yu et al. , “Segment anything in high quality,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
2024
Closest in time.
Z. Luo, Y. Xiao, Y. Liu, S. Li, Y. Wang, Y. Tang, X. Li, and Y. Yang, “Soc: Semantic-assisted object cluster for referring video object segmentation,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
Y. Li, J. Wan, X. Teng, and L. Lan, “Fine-grained text-video fusion for referring video object segmentation,” in 2024 IEEE 4th International Conference on Power, Electronics and Computer Applications (ICPECA) . IEEE, 2024, pp. 788–793
2024
Closest in time.
J. Wu, X. Li, X. Li, H. Ding, Y. Tong, and D. Tao, “Towards robust referring image segmentation,” IEEE Transactions on Image Processing , 2024
2024
Closest in time.