Fetching the paper…
Reading the bibliography…
We introduce a new benchmark designed to advance the development of general-purpose, large-scale vision-language models for remote sensing images.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Lin Chin-Yew · 2004
Earlier work this paper cites.
Meteor: An automatic metric for mt evaluation with improved correlation with human judgments
Satanjeev Banerjee and Alon Lavie · 2005
Earlier work this paper cites.
Cider: Consensus-based image description evaluation
Ramakrishna Vedantam, C Lawrence Zitnick, and Devi Parikh · 2015
Earlier work this paper cites.
Deep semantic understanding of high resolution remote sensing image
Bo Qu, Xuelong Li, Dacheng Tao, and Xiaoqiang Lu · 2016
Earlier work this paper cites.
Gaussian error linear units (gelus)
Dan Hendrycks and Kevin Gimpel · 2016
Earlier work this paper cites.
Can a machine generate humanlike language descriptions for a remote sensing image?
Zhenwei Shi and Zhengxia Zou · 2017
Earlier work this paper cites.
Exploring models and data for remote sensing image caption generation
Xiaoqiang Lu, Binqiang Wang, Xiangtao Zheng, and Xuelong Li · 2017
Earlier work this paper cites.
Natural language description of remote sensing images based on deep learning
Xiangrong Zhang, Xiang Li, Jinliang An, Li Gao, Biao Hou, and Chen Li · 2017
Earlier work this paper cites.
Description generation for remote sensing images using attribute attention mechanism
Xiangrong Zhang, Xin Wang, Xu Tang, Huiyu Zhou, and Chen Li · 2019
Earlier work this paper cites.
A multi-level attention model for remote sensing image captions
Yangyang Li, Shuangkang Fang, Licheng Jiao, Ruijiao Liu, and Ronghua Shang · 2020
Earlier work this paper cites.
Word–sentence framework for remote sensing image captioning
Qi Wang, Wei Huang, Xueting Zhang, and Xuelong Li · 2020
Earlier work this paper cites.
Truncation cross entropy loss for remote sensing image captioning
Xuelong Li, Xueting Zhang, Wei Huang, and Qi Wang · 2020
Earlier work this paper cites.
Rsvqa: Visual question answering for remote sensing data
Sylvain Lobry, Diego Marcos, Jesse Murray, and Devis Tuia · 2020
Earlier work this paper cites.
Object detection in optical remote sensing images: A survey and a new benchmark
Ke Li, Gang Wan, Gong Cheng, Liqiu Meng, and Junwei Han · 2020
Earlier work this paper cites.
High-resolution remote sensing image captioning based on structured attention
Rui Zhao, Zhenwei Shi, and Zhengxia Zou · 2021
Earlier work this paper cites.
Mutual attention inception network for remote sensing visual question answering
Xiangtao Zheng, Binqiang Wang, Xingqian Du, and Xiaoqiang Lu · 2021
Cited alongside, same era.
Object detection in aerial images: A large-scale benchmark and challenges
Jian Ding, Nan Xue, Gui-Song Xia, Xiang Bai, Wen Yang, Michael Ying Yang, Serge Belongie, Jiebo Luo, Mihai Datcu, Marcello Pelillo, et al · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Cited alongside, same era.
Lora: Low-rank adaptation of large language models
Edward J Hu, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, Weizhu Chen, et al · 2021
Cited alongside, same era.
Transforming remote sensing images to textual descriptions
Usman Zia, M Mohsin Riaz, and Abdul Ghafoor · 2022
Cited alongside, same era.
Gpt-4 technical report, 2023
OpenAI · 2023
Later among the works it cites.
Improved baselines with visual instruction tuning
Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee · 2023
Later among the works it cites.
Minigpt-v2: Large language model as a unified interface for vision-language multi-task learning
Jun Chen, Deyao Zhu, Xiaoqian Shen, Xiang Li, Zechun Liu, Pengchuan Zhang, Raghuraman Krishnamoorthi, Vikas Chandra, Yunyang Xiong, and Mohamed Elhoseiny · 2023
Later among the works it cites.
Clair: Evaluating image captions with large language models
David Chan, Suzanne Petryk, Joseph E Gonzalez, Trevor Darrell, and John Canny · 2023
Later among the works it cites.
Visual instruction tuning
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yuxi Sun, Shanshan Feng, Xutao Li, Yunming Ye, Jian Kang, and Xu Huang · 2022
Cited alongside, same era.
Language transformers for remote sensing visual question answering
Christel Chappuis, Vincent Mendez, Eliot Walt, Sylvain Lobry, Bertrand Le Saux, and Devis Tuia · 2022
Cited alongside, same era.
Open-ended remote sensing visual question answering with transformers
Mohamad M Al Rahhal, Yakoub Bazi, Sara O Alsaleh, Muna Al-Razgan, Mohamed Lamine Mekhalfi, Mansour Al Zuair, and Naif Alajlan · 2022
Cited alongside, same era.
Bi-modal transformer-based approach for visual question answering in remote sensing imagery
Yakoub Bazi, Mohamad Mahmoud Al Rahhal, Mohamed Lamine Mekhalfi, Mansour Abdulaziz Al Zuair, and Farid Melgani · 2022
Cited alongside, same era.
From easy to hard: Learning language-guided curriculum for visual question answering on remote sensing data
Zhenghang Yuan, Lichao Mou, Qi Wang, and Xiao Xiang Zhu · 2022
Cited alongside, same era.
Qwen-vl: A frontier large vision-language model with versatile abilities
Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou · 2023
Cited alongside, same era.
Rs-clip: Zero shot remote sensing scene classification via contrastive vision-language supervision
Xiang Li, Congcong Wen, Yuan Hu, and Nan Zhou · 2023
Cited alongside, same era.
Minigpt-4: Enhancing vision-language understanding with advanced large language models
Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny · 2024
Closest in time.
Vision-language models in remote sensing: Current progress and future trends
Xiang Li, Congcong Wen, Yuan Hu, Zhenghang Yuan, and Xiao Xiang Zhu · 2024
Closest in time.
Remoteclip: A vision language foundation model for remote sensing
Fan Liu, Delong Chen, Zhangqingyun Guan, Xiaocong Zhou, Jiale Zhu, Qiaolin Ye, Liyong Fu, and Jun Zhou · 2024
Closest in time.
Rotated multi-scale interaction network for referring remote sensing image segmentation
Sihan Liu, Yiwei Ma, Xiaoqing Zhang, Haowei Wang, Jiayi Ji, Xiaoshuai Sun, and Rongrong Ji · 2024
Closest in time.
Geochat: Grounded large vision-language model for remote sensing
Kartik Kuckreja, Muhammad Sohail Danish, Muzammal Naseer, Abhijit Das, Salman Khan, and Fahad Shahbaz Khan · 2024
Closest in time.
Wei Zhang, Miaoxin Cai, Tong Zhang, Yin Zhuang, and Xuerui Mao · 2024
Closest in time.
Rrsis: Referring remote sensing image segmentation
Zhenghang Yuan, Lichao Mou, Yuansheng Hua, and Xiao Xiang Zhu · 2024
Closest in time.
A survey on hallucination in large vision-language models
Hanchao Liu, Wenyuan Xue, Yifei Chen, Dapeng Chen, Xiutian Zhao, Ke Wang, Liping Hou, Rongjun Li, and Wei Peng · 2024
Closest in time.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, May 2024
Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing · 2024
Closest in time.
Mini-gemini: Mining the potential of multi-modality vision language models
Yanwei Li, Yuechen Zhang, Chengyao Wang, Zhisheng Zhong, Yixin Chen, Ruihang Chu, Shaoteng Liu, and Jiaya Jia · 2024
Closest in time.