Fetching the paper…
Reading the bibliography…
Large language models (LLMs) have recently been extended to the vision-language realm, obtaining impressive general multi-modal capabilities.
Deep semantic understanding of high resolution remote sensing image
Bo Qu, Xuelong Li, Dacheng Tao, and Xiaoqiang Lu · 2016
Earlier work this paper cites.
Exploring models and data for remote sensing image caption generation
Xiaoqiang Lu, Binqiang Wang, Xiangtao Zheng, and Xuelong Li · 2018
Earlier work this paper cites.
Dota: A large-scale dataset for object detection in aerial images
Gui-Song Xia, Xiang Bai, Jian Ding, Zhen Zhu, Serge Belongie, Jiebo Luo, Mihai Datcu, Marcello Pelillo, and Liangpei Zhang · 2018
Earlier work this paper cites.
A fast and accurate one-stage approach to visual grounding
Zhengyuan Yang, Boqing Gong, Liwei Wang, Wenbing Huang, Dong Yu, and Jiebo Luo · 2019
Earlier work this paper cites.
Object detection in optical remote sensing images: A survey and a new benchmark
Ke Li, Gang Wan, Gong Cheng, Liqiu Meng, and Junwei Han · 2020
Earlier work this paper cites.
Rsvqa: Visual question answering for remote sensing data
Sylvain Lobry, Diego Marcos, Jesse Murray, and Devis Tuia · 2020
Earlier work this paper cites.
Sound active attention framework for remote sensing image captioning
Xiaoqiang Lu, Binqiang Wang, and Xiangtao Zheng · 2020
Earlier work this paper cites.
Era: A data set and deep learning benchmark for event recognition in aerial videos [software and data sets]
Lichao Mou, Yuansheng Hua, Pu Jin, and Xiao Xiang Zhu · 2020
Earlier work this paper cites.
Improving one-stage visual grounding by recursive sub-query construction
Zhengyuan Yang, Tianlang Chen, Liwei Wang, and Jiebo Luo · 2020
Earlier work this paper cites.
Look before you leap: Learning landmark features for one-stage visual grounding
Binbin Huang, Dongze Lian, Weixin Luo, and Shenghua Gao · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Earlier work this paper cites.
Nwpu-captions dataset and mlca-net for remote sensing image captioning
Qimin Cheng, Haiyan Huang, Yuan Xu, Yuzhuo Zhou, Huanying Li, and Zhongyuan Wang · 2022
Earlier work this paper cites.
Visual grounding in remote sensing images
Yuxi Sun, Shanshan Feng, Xutao Li, Yunming Ye, Jian Kang, and Xu Huang · 2022
Earlier work this paper cites.
Earthnets: Empowering AI in earth observation
Zhitong Xiong, Fahong Zhang, Yi Wang, Yilei Shi, and Xiao Xiang Zhu · 2022
Cited alongside, same era.
From easy to hard: Learning language-guided curriculum for visual question answering on remote sensing data
Zhenghang Yuan, Lichao Mou, Qi Wang, and Xiao Xiang Zhu · 2022
Cited alongside, same era.
Exploring a fine-grained multiscale method for cross-modal remote sensing image retrieval
Zhiqiang Yuan, Wenkai Zhang, Kun Fu, Xuan Li, Chubo Deng, Hongqi Wang, and Xian Sun · 2022
Cited alongside, same era.
Mutual attention inception network for remote sensing visual question answering
Xiangtao Zheng, Binqiang Wang, Xingqian Du, and Xiaoqiang Lu · 2022
Cited alongside, same era.
Capera: Captioning events in aerial videos
Laila Bashmal, Yakoub Bazi, Mohamad Mahmoud Al Rahhal, Mansour Zuair, and Farid Melgani · 2023
Cited alongside, same era.
Visual instruction tuning
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee · 2023
Later among the works it cites.
Gpt-4 technical report
OpenAI · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al · 2023
Later among the works it cites.
Visionllm: Large language model is also an open-ended decoder for vision-centric tasks
Wenhai Wang, Zhe Chen, Xiaokang Chen, Jiannan Wu, Xizhou Zhu, Gang Zeng, Ping Luo, Tong Lu, Jie Zhou, Yu Qiao, et al · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jun Chen, Deyao Zhu, Xiaoqian Shen, Xiang Li, Zechun Liu, Pengchuan Zhang, Raghuraman Krishnamoorthi, Vikas Chandra, Yunyang Xiong, and Mohamed Elhoseiny · 2023
Cited alongside, same era.
Shikra: Unleashing multimodal llm’s referential dialogue magic
Keqin Chen, Zhao Zhang, Weili Zeng, Richong Zhang, Feng Zhu, and Rui Zhao · 2023
Cited alongside, same era.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality
Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E Gonzalez, et al · 2023
Cited alongside, same era.
Instructblip: Towards general-purpose vision-language models with instruction tuning
Wenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong, Junqi Zhao, Weisheng Wang, Boyang Li, Pascale Fung, and Steven Hoi · 2023
Cited alongside, same era.
Eva: Exploring the limits of masked visual representation learning at scale
Yuxin Fang, Wen Wang, Binhui Xie, Quan Sun, Ledell Wu, Xinggang Wang, Tiejun Huang, Xinlong Wang, and Yue Cao · 2023
Cited alongside, same era.
Improving image captioning systems with postprocessing strategies
Genc Hoxha, Giacomo Scuccato, and Farid Melgani · 2023
Cited alongside, same era.
Rsgpt: A remote sensing vision language model and benchmark
Yuan Hu, Jianlong Yuan, Congcong Wen, Xiaonan Lu, and Xiang Li · 2023
Cited alongside, same era.
Later among the works it cites.
u-llava: Unifying multi-modal tasks via large language model
Jinjin Xu, Liwu Xu, Yuzhe Yang, Xiang Li, Yanchun Xie, Yi-Jie Huang, and Yaqian Li · 2023
Later among the works it cites.
Parameter-efficient transfer learning for remote sensing image–text retrieval
Yuan Yuan, Yang Zhan, and Zhitong Xiong · 2023
Later among the works it cites.
Overcoming language bias in remote sensing visual question answering via adversarial training
Zhenghang Yuan, Lichao Mou, and Xiao Xiang Zhu · 2023
Later among the works it cites.
Rsvg: Exploring data and models for visual grounding on remote sensing data
Yang Zhan, Zhitong Xiong, and Yuan Yuan · 2023
Later among the works it cites.
Mono3dvg: 3d visual grounding in monocular images
Yang Zhan, Yuan Yuan, and Zhitong Xiong · 2023
Later among the works it cites.
A spatial hierarchical reasoning network for remote sensing visual question answering
Zixiao Zhang, Licheng Jiao, Lingling Li, Xu Liu, Puhua Chen, Fang Liu, Yuxuan Li, and Zhicheng Guo · 2023
Later among the works it cites.
Minigpt-4: Enhancing vision-language understanding with advanced large language models
Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny · 2023
Later among the works it cites.