Fetching the paper…
Reading the bibliography…
Recent advances in large multimodal models (LMMs) have recognized fine-grained grounding as an imperative factor of visual understanding and dialogue.
Referring image segmentation via recurrent refinement networks
Li, R., Li, K., Kuo, Y.-C., Shu, M., Qi, X., Shen, X., and Jia, J · 2018
Earlier work this paper cites.
Towards better validity: Dispersion based clustering for unsupervised person re-identification
Ding, G., Khan, S., Tang, Z., Zhang, J., and Porikli, F · 2019
Earlier work this paper cites.
isaid: A large-scale dataset for instance segmentation in aerial images
Waqas Zamir, S., Arora, A., Gupta, A., Khan, S., Sun, G., Shahbaz Khan, F., Zhu, F., Shao, L., Xia, G.-S., and Bai, X · 2019
Earlier work this paper cites.
Cross-modal self-attention network for referring image segmentation
Ye, L., Rochan, M., Liu, Z., and Wang, Y · 2019
Earlier work this paper cites.
Minigpt-4: Enhancing vision-language understanding with advanced large language models
Zhu, D., Chen, J., Shen, X., Li, X., and Elhoseiny, M · 2019
Earlier work this paper cites.
Bi-directional relationship inferring network for referring image segmentation
Hu, Z., Feng, G., Sun, J., Zhang, L., and Lu, H · 2020
Earlier work this paper cites.
Referring image segmentation via cross-modal progressive comprehension
Huang, S., Hui, T., Liu, S., Li, G., Wei, Y., Han, J., Liu, L., and Li, B · 2020
Earlier work this paper cites.
Linguistic structure guided context modeling for referring image segmentation
Hui, T., Liu, S., Huang, S., Li, G., Yu, S., Zhang, F., and Han, J · 2020
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N · 2021
Earlier work this paper cites.
Cross-modal progressive comprehension for referring segmentation
Liu, S., Hui, T., Huang, S., Wei, Y., Li, B., and Li, G · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al · 2021
Earlier work this paper cites.
Masked autoencoders are scalable vision learners
He, K., Chen, X., Xie, S., Li, Y., Dollár, P., and Girshick, R · 2022
Earlier work this paper cites.
Shikra: Unleashing multimodal llm’s referential dialogue magic
Chen, K., Zhang, Z., Zeng, W., Zhang, R., Zhu, F., and Zhao, R · 2023
Earlier work this paper cites.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, March 2023
Chiang, W.-L., Li, Z., Lin, Z., Sheng, Y., Wu, Z., Zhang, H., Zheng, L., Zhuang, S., Zhuang, Y., Gonzalez, J. E., Stoica, I., and Xing, E. P · 2023
Earlier work this paper cites.
Instructblip: Towards general-purpose vision-language models with instruction tuning
Dai, W., Li, J., LI, D., Tiong, A., Zhao, J., Wang, W., Li, B., Fung, P. N., and Hoi, S · 2023
Earlier work this paper cites.
Rsgpt: A remote sensing vision language model and benchmark
Hu, Y., Yuan, J., Wen, C., Lu, X., and Li, X · 2023
Earlier work this paper cites.
Phi-2: the surprising power of small language models (2023)
Javaheripi, M., Bubeck, S., et al · 2023
Earlier work this paper cites.
Segment anything
Kirillov, A., Mintun, E., Ravi, N., Mao, H., Rolland, C., Gustafson, L., Xiao, T., Whitehead, S., Berg, A. C., Lo, W.-Y., et al · 2023
Earlier work this paper cites.
Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Li, J., Li, D., Savarese, S., and Hoi, S · 2023
Cited alongside, same era.
Lin, Z., Liu, C., Zhang, R., Gao, P., Qiu, L., Xiao, H., Qiu, H., Lin, C., Shao, W., Chen, K., et al · 2023
Cited alongside, same era.
Hiera: A hierarchical vision transformer without the bells-and-whistles
Ryali, C., Hu, Y.-T., Bolya, D., Wei, C., Fan, H., Huang, P.-Y., Aggarwal, V., Chowdhury, A., Poursaeed, O., Hoffman, J., Malik, J., Li, Y., and Feichtenhofer, C · 2023
Cited alongside, same era.
Gemini: a family of highly capable multimodal models
Team, G., Anil, R., Borgeaud, S., Alayrac, J.-B., Yu, J., Soricut, R., Schalkwyk, J., Dai, A. M., Hauth, A., Millican, K., et al · 2023
Cited alongside, same era.
Rotated multi-scale interaction network for referring remote sensing image segmentation
Liu, S., Ma, Y., Zhang, X., Wang, H., Ji, J., Sun, X., and Ji, R · 2024
Later among the works it cites.
Luo, J., Pang, Z., Zhang, Y., Wang, T., Wang, L., Dang, B., Lao, J., Wang, J., Chen, J., Tan, Y., et al · 2024
Later among the works it cites.
Lhrs-bot: Empowering remote sensing with vgi-enhanced large multimodal language model
Muhtar, D., Li, Z., Gu, F., Zhang, X., and Xiao, P · 2024
Later among the works it cites.
Chatgpt: Language model for dialogue applications
OpenAI · 2024
Later among the works it cites.
H2rsvlm: Towards helpful and honest remote sensing large vision language model
Pang, C., Wu, J., Li, J., Liu, Y., Sun, J., Li, W., Weng, X., Wang, S., Feng, L., Xia, G.-S., et al · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., et al · 2023
Cited alongside, same era.
Set-of-mark prompting unleashes extraordinary visual grounding in gpt-4v
Yang, J., Zhang, H., Li, F., Zou, X., Li, C., and Gao, J · 2023
Cited alongside, same era.
Ferret: Refer and ground anything anywhere at any granularity
You, H., Zhang, H., Gan, Z., Du, X., Zhang, B., Wang, Z., Cao, L., Chang, S.-F., and Yang, Y · 2023
Cited alongside, same era.
Gpt4roi: Instruction tuning large language model on region-of-interest
Zhang, S., Sun, P., Chen, S., Xiao, M., Shao, W., Zhang, W., Liu, Y., Chen, K., and Luo, P · 2023
Cited alongside, same era.
Bubogpt: Enabling visual grounding in multi-modal llms
Zhao, Y., Lin, Z., Zhou, D., Huang, Z., Feng, J., and Kang, B · 2023
Cited alongside, same era.
Rs-llava: A large vision-language model for joint captioning and question answering in remote sensing imagery
Bazi, Y., Bashmal, L., Al Rahhal, M. M., Ricci, R., and Melgani, F · 2024
Cited alongside, same era.
Internlm2 technical report, 2024
Cai, Z., Cao, M., Chen, H., Chen, K., Chen, K., Chen, X., Chen, X., Chen, Z., Chen, Z., Chu, P., Dong, X., Duan, H., et al · 2024
Cited alongside, same era.
Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., et al · 2024
Cited alongside, same era.
Later among the works it cites.
Grounding multimodal large language models to the world
Peng, Z., Wang, W., Dong, L., Hao, Y., Huang, S., Ma, S., Ye, Q., and Wei, F · 2024
Later among the works it cites.
Glamm: Pixel grounding large multimodal model
Rasheed, H., Maaz, M., Shaji, S., Shaker, A., Khan, S., Cholakkal, H., Anwer, R. M., Xing, E., Yang, M.-H., and Khan, F. S · 2024
Later among the works it cites.
Sam 2: Segment anything in images and videos
Ravi, N., Gabeur, V., Hu, Y.-T., Hu, R., Ryali, C., Ma, T., Khedr, H., Rädle, R., Rolland, C., Gustafson, L., Mintun, E., Pan, J., Alwala, K. V., Carion, N., Wu, C.-Y., Girshick, R., Dollár, P., and Feichtenhofer, C · 2024
Later among the works it cites.
Pixellm: Pixel reasoning with large multimodal model
Ren, Z., Huang, Z., Wei, Y., Zhao, Y., Fu, D., Feng, J., and Jin, X · 2024
Later among the works it cites.
Earthdial: Turning multi-sensory earth observations to interactive dialogues
Soni, S., Dudhane, A., Debary, H., Fiaz, M., Munir, M. A., Danish, M. S., Fraccaro, P., Watson, C. D., Klein, L. J., Khan, F. S., et al · 2024
Later among the works it cites.
Gsva: Generalized segmentation via multimodal large language models
Xia, Z., Han, D., Han, Y., Pan, X., Song, S., and Huang, G · 2024
Later among the works it cites.
Pink: Unveiling the power of referential comprehension for multi-modal llms
Xuan, S., Guo, Q., Yang, M., and Zhang, S · 2024
Later among the works it cites.
Language-aware vision transformer for referring segmentation
Yang, Z., Wang, J., Ye, X., Tang, Y., Chen, K., Zhao, H., and Torr, P. H · 2024
Later among the works it cites.
Rrsis: Referring remote sensing image segmentation
Yuan, Z., Mou, L., Hua, Y., and Zhu, X. X · 2024
Later among the works it cites.
Zhan, Y., Xiong, Z., and Yuan, Y · 2024
Later among the works it cites.
Earthgpt: A universal multimodal large language model for multisensor image comprehension in remote sensing domain
Zhang, W., Cai, M., Zhang, T., Zhuang, Y., and Mao, X · 2024
Later among the works it cites.
Groma: Localized visual tokenization for grounding multimodal large language models
Ma, C., Jiang, Y., Wu, J., Yuan, Z., and Qi, X · 2025
Closest in time.
Vary: Scaling up the vision vocabulary for large vision-language model
Wei, H., Kong, L., Chen, J., Zhao, L., Ge, Z., Yang, J., Sun, J., Han, C., and Zhang, X · 2025
Closest in time.