Fetching the paper…
Reading the bibliography…
As a novel and challenging task, referring segmentation combines computer vision and natural language processing to localize and segment objects based on textual descriptions.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L. (2009) · 2009
Earlier work this paper cites.
Microsoft coco: Common objects in context
Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., and Zitnick, C. L. (2014) · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Simonyan, K. and Zisserman, A. (2014) · 2014
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation
Ronneberger, O., Fischer, P., and Brox, T. (2015) · 2015
Earlier work this paper cites.
Segmentation from natural language expressions
Hu, R., Rohrbach, M., and Darrell, T. (2016) · 2016
Earlier work this paper cites.
Modeling context between objects for referring expression understanding
Nagaraja, V. K., Morariu, V. I., and Davis, L. S. (2016) · 2016
Earlier work this paper cites.
Modeling context in referring expressions
Yu, L., Poirson, P., Yang, S., Berg, A. C., and Berg, T. L. (2016) · 2016
Earlier work this paper cites.
Recurrent multimodal interaction for referring image segmentation
Liu, C., Lin, Z., Shen, X., Yang, J., Lu, X., and Yuille, A. (2017) · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F. (2017) · 2017
Earlier work this paper cites.
Referring image segmentation via recurrent refinement networks
Li, R., Li, K., Kuo, Y.-C., Shu, M., Qi, X., Shen, X., and Jia, J. (2018) · 2018
Earlier work this paper cites.
Dynamic multimodal instance segmentation guided by natural language queries
Margffoy-Tuay, E., Pérez, J. C., Botero, E., and Arbeláez, P. (2018) · 2018
Earlier work this paper cites.
Key-word-aware network for referring expression image segmentation
Shi, H., Li, H., Meng, F., and Wu, Q. (2018) · 2018
Earlier work this paper cites.
Mattnet: Modular attention network for referring expression comprehension
Yu, L., Lin, Z., Shen, X., Yang, J., Lu, X., Bansal, M., and Berg, T. L. (2018) · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. (2019) · 2019
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., et al. (2019) · 2019
Earlier work this paper cites.
Cross-modal self-attention network for referring image segmentation
Ye, L., Rochan, M., Liu, Z., and Wang, Y. (2019) · 2019
Earlier work this paper cites.
Bi-directional relationship inferring network for referring image segmentation
Hu, Z., Feng, G., Sun, J., Zhang, L., and Lu, H. (2020) · 2020
Earlier work this paper cites.
Uavid: A semantic segmentation dataset for uav imagery
Lyu, Y., Vosselman, G., Xia, G.-S., Yilmaz, A., and Yang, M. Y. (2020) · 2020
Cited alongside, same era.
Swin transformer: Hierarchical vision transformer using shifted windows
Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., and Guo, B. (2021) · 2021
Cited alongside, same era.
Assessing city-scale green roof development potential using unmanned aerial vehicle (uav) imagery
Shao, H., Song, P., Mu, B., Tian, G., Chen, Q., He, R., and Kim, G. (2021) · 2021
Cited alongside, same era.
High-resolution remote sensing image captioning based on structured attention
Zhao, R., Shi, Z., and Zou, Z. (2021) · 2021
Cited alongside, same era.
Vlt: Vision-language transformer and query generation for referring segmentation
Ding, H., Liu, C., Wang, S., and Jiang, X. (2022) · 2022
Cited alongside, same era.
Rsvg: Exploring data and models for visual grounding on remote sensing data
Zhan, Y., Xiong, Z., and Yuan, Y. (2023) · 2023
Later among the works it cites.
Efficient large-scale oblique image matching based on cascade hashing and match data scheduling
Zhang, Q., Zheng, S., Zhang, C., Wang, X., and Li, R. (2023) · 2023
Later among the works it cites.
Visual selection and multi-stage reasoning for rsvg
Ding, Y., Xu, H., Wang, D., Li, K., and Tian, Y. (2024) · 2024
Later among the works it cites.
Cross-modal bidirectional interaction model for referring remote sensing image segmentation
Dong, Z., Sun, Y., Gu, Y., and Liu, T. (2024) · 2024
Later among the works it cites.
Lisa: Reasoning segmentation via large language model
Lai, X., Tian, Z., Chen, Y., Li, Y., Yuan, Y., Liu, S., and Jia, J. (2024) · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Land cover classification of treeline ecotones along a 1100 km latitudinal transect using spectral-and three-dimensional information from uav-based aerial imagery
Mienna, I. M., Klanderud, K., Ørka, H. O., Bryn, A., and Bollandsås, O. M. (2022) · 2022
Cited alongside, same era.
Uav based long range environment monitoring system with industry 5.0 perspectives for smart city infrastructure
Sharma, R. and Arya, R. (2022) · 2022
Cited alongside, same era.
Visual grounding in remote sensing images
Sun, Y., Feng, S., Li, X., Ye, Y., Kang, J., and Huang, X. (2022) · 2022
Cited alongside, same era.
Uav-borne, lidar-based elevation modelling: A method for improving local-scale urban flood risk assessment
Trepekli, K., Balstrøm, T., Friborg, T., Fog, B., Allotey, A. N., Kofie, R. Y., and Møller-Jensen, L. (2022) · 2022
Cited alongside, same era.
Unetformer: A unet-like transformer for efficient semantic segmentation of remote sensing urban scene imagery
Wang, L., Li, R., Zhang, C., Fang, S., Duan, C., Meng, X., and Atkinson, P. M. (2022) · 2022
Cited alongside, same era.
Lavt: Language-aware vision transformer for referring image segmentation
Yang, Z., Wang, J., Tang, Y., Chen, K., Zhao, H., and Torr, P. H. (2022) · 2022
Cited alongside, same era.
Vdd: Varied drone dataset for semantic segmentation
Cai, W., Jin, K., Hou, J., Guo, C., Wu, L., and Yang, W. (2023) · 2023
Cited alongside, same era.
Exploring fine-grained image-text alignment for referring remote sensing image segmentation
Lei, S., Xiao, X., Zhang, T., Li, H.-C., Shi, Z., and Zhu, Q. (2024) · 2024
Later among the works it cites.
Language-guided progressive attention for visual grounding in remote sensing images
Li, K., Wang, D., Xu, H., Zhong, H., and Wang, C. (2024) · 2024
Later among the works it cites.
Lswinsr: Uav imagery super-resolution based on linear swin transformer
Li, R. and Zhao, X. (2024) · 2024
Later among the works it cites.
Rotated multi-scale interaction network for referring remote sensing image segmentation
Liu, S., Ma, Y., Zhang, X., Wang, H., Ji, J., Sun, X., and Ji, R. (2024) · 2024
Later among the works it cites.
Rrsis: Referring remote sensing image segmentation
Yuan, Z., Mou, L., Hua, Y., and Zhu, X. X. (2024) · 2024
Later among the works it cites.
Bai, S., Chen, K., Liu, X., Wang, J., Ge, W., Song, S., Dang, K., Wang, P., Wang, S., Tang, J., et al. (2025) · 2025
Closest in time.
Rsrefseg: Referring remote sensing image segmentation with foundation models
Chen, K., Zhang, J., Liu, C., Zou, Z., and Shi, Z. (2025) · 2025
Closest in time.
Scale-wise bidirectional alignment network for referring remote sensing image segmentation
Li, K., Vosselman, G., and Yang, M. Y. (2025) · 2025
Closest in time.
Multimodal-aware fusion network for referring remote sensing image segmentation
Shi, L. and Zhang, J. (2025) · 2025
Closest in time.
Referring remote sensing image segmentation via bidirectional alignment guided joint prediction
Zhang, T., Wen, Z., Kong, B., Liu, K., Zhang, Y., Zhuang, P., and Li, J. (2025) · 2025
Closest in time.
Rethinking the implicit optimization paradigm with dual alignments for referring remote sensing image segmentation
Pan, Y., Sun, R., Wang, Y., Zhang, T., and Zhang, Y. (2024) · 2040
Closest in time.