Fetching the paper…
Reading the bibliography…
Referring Remote Sensing Image Segmentation is a complex and challenging task that integrates the paradigms of computer vision and natural language processing.
S. Kazemzadeh, V. Ordonez, M. Matten, and T. Berg, “Referitgame: Referring to objects in photographs of natural scenes,” in Proceedings of the 2014 conference on empirical methods in natural language processing (EMNLP) , 2014, pp. 787–798
2014
Earlier work this paper cites.
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick, “Microsoft coco: Common objects in context,” in Computer Vision–ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part V 13 . Springer, 2014, pp. 740–755
2014
Earlier work this paper cites.
A. Vaswani, “Attention is all you need,” Advances in Neural Information Processing Systems , 2017
2017
Earlier work this paper cites.
R. Li, K. Li, Y.-C. Kuo, M. Shu, X. Qi, X. Shen, and J. Jia, “Referring image segmentation via recurrent refinement networks,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2018, pp. 5745–5753
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
L. Ye, M. Rochan, Z. Liu, and Y. Wang, “Cross-modal self-attention network for referring image segmentation,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2019, pp. 10 502–10 511
2019
Earlier work this paper cites.
D. Z. Chen, A. X. Chang, and M. Nießner, “Scanrefer: 3d object localization in rgb-d scans using natural language,” in European conference on computer vision . Springer, 2020, pp. 202–221
2020
Earlier work this paper cites.
S. Huang, T. Hui, S. Liu, G. Li, Y. Wei, J. Han, L. Liu, and B. Li, “Referring image segmentation via cross-modal progressive comprehension,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2020, pp. 10 488–10 497
2020
Earlier work this paper cites.
Z. Hu, G. Feng, J. Sun, L. Zhang, and H. Lu, “Bi-directional relationship inferring network for referring image segmentation,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2020, pp. 4424–4433
2020
Earlier work this paper cites.
2021
Earlier work this paper cites.
S. Liu, T. Hui, S. Huang, Y. Wei, B. Li, and G. Li, “Cross-modal progressive comprehension for referring segmentation,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 44, no. 9, pp. 4761–4775, 2021
2021
Earlier work this paper cites.
Y. Jing, T. Kong, W. Wang, L. Wang, L. Li, and T. Tan, “Locate then segment: A strong pipeline for referring image segmentation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2021, pp. 9858–9867
2021
Earlier work this paper cites.
Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo, “Swin transformer: Hierarchical vision transformer using shifted windows,” in Proceedings of the IEEE/CVF international conference on computer vision , 2021, pp. 10 012–10 022
2021
Earlier work this paper cites.
2022
Earlier work this paper cites.
Z. Yang, J. Wang, Y. Tang, K. Chen, H. Zhao, and P. H. Torr, “Lavt: Language-aware vision transformer for referring image segmentation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 18 155–18 165
2022
Earlier work this paper cites.
D. Di Palma, “Retrieval-augmented recommender system: Enhancing recommender systems with large language models,” in Proceedings of the 17th ACM Conference on Recommender Systems , 2023, pp. 1369–1373
2023
Cited alongside, same era.
C. Liu, H. Ding, and X. Jiang, “Gres: Generalized referring expression segmentation,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2023, pp. 23 592–23 601
2023
Cited alongside, same era.
H. Ding, C. Liu, S. He, X. Jiang, and C. C. Loy, “Mevis: A large-scale benchmark for video segmentation with motion expressions,” in Proceedings of the IEEE/CVF international conference on computer vision , 2023, pp. 2694–2703
2023
Cited alongside, same era.
A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y. Lo et al. , “Segment anything,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 4015–4026
2023
Cited alongside, same era.
X. Wang, S. Li, K. Kallidromitis, Y. Kato, K. Kozuka, and T. Darrell, “Hierarchical open-vocabulary universal image segmentation,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Later among the works it cites.
S. Lei, X. Xiao, T. Zhang, H.-C. Li, Z. Shi, and Q. Zhu, “Exploring fine-grained image-text alignment for referring remote sensing image segmentation,” IEEE Transactions on Geoscience and Remote Sensing , 2024
2024
Later among the works it cites.
S. Ha, C. Kim, D. Kim, J. Lee, S. Lee, and J. Lee, “Finding nemo: Negative-mined mosaic augmentation for referring image segmentation,” in European Conference on Computer Vision . Springer, 2024, pp. 121–137
2024
Later among the works it cites.
Y. Yang, C. Ma, J. Yao, Z. Zhong, Y. Zhang, and Y. Wang, “Remamber: Referring image segmentation with mamba twister,” in European Conference on Computer Vision . Springer, 2024, pp. 108–126
2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y. Zhan, Z. Xiong, and Y. Yuan, “Rsvg: Exploring data and models for visual grounding on remote sensing data,” IEEE Transactions on Geoscience and Remote Sensing , vol. 61, pp. 1–13, 2023
2023
Cited alongside, same era.
F. Liu, Y. Liu, Y. Kong, K. Xu, L. Zhang, B. Yin, G. Hancke, and R. Lau, “Referring image segmentation using text supervision,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 22 124–22 134
2023
Cited alongside, same era.
Z. Wei, L. Chen, Y. Jin, X. Ma, T. Liu, P. Ling, B. Wang, H. Chen, and J. Zheng, “Stronger fewer & superior: Harnessing vision foundation models for domain generalized semantic segmentation,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2024, pp. 28 619–28 630
2024
Cited alongside, same era.
B. Xie, J. Cao, J. Xie, F. S. Khan, and Y. Pang, “Sed: A simple encoder-decoder for open-vocabulary semantic segmentation,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2024, pp. 3426–3436
2024
Cited alongside, same era.
2024
Cited alongside, same era.
C. Shang, Z. Song, H. Qiu, L. Wang, F. Meng, and H. Li, “Prompt-driven referring image segmentation with instance contrasting,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 4124–4134
2024
Cited alongside, same era.
Y. X. Chng, H. Zheng, Y. Han, X. Qiu, and G. Huang, “Mask grounding for referring image segmentation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 26 573–26 583
2024
Cited alongside, same era.
Q. Li, M. Zhang, Z. Yang, Y. Yuan, and Q. Wang, “Edge-guided perceptual network for infrared small target detection,” IEEE Transactions on Geoscience and Remote Sensing , 2024
2024
Cited alongside, same era.
P. Yue, J. Lin, S. Zhang, J. Hu, Y. Lu, H. Niu, H. Ding, Y. Zhang, G. Jiang, L. Cao et al. , “Adaptive selection based referring image segmentation,” in Proceedings of the 32nd ACM International Conference on Multimedia , 2024, pp. 1101–1110
2024
Later among the works it cites.
X. Lai, Z. Tian, Y. Chen, Y. Li, Y. Yuan, S. Liu, and J. Jia, “Lisa: Reasoning segmentation via large language model,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 9579–9589
2024
Later among the works it cites.
Z. Ren, Z. Huang, Y. Wei, Y. Zhao, D. Fu, J. Feng, and X. Jin, “Pixellm: Pixel reasoning with large multimodal model,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 26 374–26 383
2024
Later among the works it cites.
Z. Yang, W. Zhang, Q. Li, W. Ni, J. Wu, and Q. Wang, “C 2 net: Road extraction via context perception and cross spatial-scale feature interaction,” IEEE Transactions on Geoscience and Remote Sensing , 2024
2024
Later among the works it cites.
Z. Yang, Q. Li, Y. Yuan, and Q. Wang, “Hcnet: Hierarchical feature aggregation and cross-modal feature alignment for remote sensing image captioning,” IEEE Transactions on Geoscience and Remote Sensing , 2024
2024
Later among the works it cites.
Q. Li, W. Zhang, W. Lu, and Q. Wang, “Multi-branch mutual-guiding learning for infrared small target detection,” IEEE Transactions on Geoscience and Remote Sensing , 2025
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
H. Nguyen-Truong, E.-R. Nguyen, T.-A. Vu, M.-T. Tran, B.-S. Hua, and S.-K. Yeung, “Vision-aware text features in referring image segmentation: From object understanding to context understanding,” in 2025 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) . IEEE, 2025, pp. 4988–4998
2025
Closest in time.
C. Yan, H. Wang, S. Yan, X. Jiang, Y. Hu, G. Kang, W. Xie, and E. Gavves, “Visa: Reasoning video object segmentation via large language models,” in European Conference on Computer Vision . Springer, 2025, pp. 98–115
2025
Closest in time.