Fetching the paper…
Reading the bibliography…
This paper develops a Versatile and Honest vision language Model (VHM) for remote sensing image analysis.
A literature survey on algorithms for multi-label learning
Sorower, M. S. 2010 · 2010
Earlier work this paper cites.
Satellite Image Classification via Two-Layer Sparse Coding With Biased Image Representation
Dai, D.; and Yang, W. 2011 · 2011
Earlier work this paper cites.
International Society for Photogrammetry and Remote Sensing: 2D semantic labeling challenge
ISPRS. 2016 · 2016
Earlier work this paper cites.
A scene change detection framework for multi-temporal very high resolution remote sensing images
Wu, C.; Zhang, L.; and Zhang, L. 2016 · 2016
Earlier work this paper cites.
Bag-of-visual-words scene classifier with local and global features for high spatial resolution remote sensing imagery
Zhu, Q.; Zhong, Y.; Zhao, B.; Xia, G.-S.; and Zhang, L. 2016 · 2016
Earlier work this paper cites.
Remote sensing image scene classification: Benchmark and state of the art
Cheng, G.; Han, J.; and Lu, X. 2017 · 2017
Earlier work this paper cites.
AID: A benchmark data set for performance evaluation of aerial scene classification
Xia, G.-S.; Hu, J.; Hu, F.; Shi, B.; Bai, X.; Zhong, Y.; Zhang, L.; and Lu, X. 2017 · 2017
Earlier work this paper cites.
Functional map of the world
Christie, G.; Fendley, N.; Wilson, J.; and Mukherjee, R. 2018 · 2018
Earlier work this paper cites.
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
Sharma, P.; Ding, N.; Goodman, S.; and Soricut, R. 2018 · 2018
Earlier work this paper cites.
DOTA: A large-scale dataset for object detection in aerial images
Xia, G.-S.; Bai, X.; Ding, J.; Zhu, Z.; Belongie, S.; Luo, J.; Datcu, M.; Pelillo, M.; and Zhang, L. 2018 · 2018
Earlier work this paper cites.
Lending orientation to neural networks for cross-view geo-localization
Liu, L.; and Li, H. 2019 · 2019
Earlier work this paper cites.
Object detection in optical remote sensing images: A survey and a new benchmark
Li, K.; Wan, G.; Cheng, G.; Meng, L.; and Han, J. 2020 · 2020
Earlier work this paper cites.
RSVQA: Visual question answering for remote sensing data
Lobry, S.; Marcos, D.; Murray, J.; and Tuia, D. 2020 · 2020
Earlier work this paper cites.
Land-cover classification with high-resolution remote sensing images using transferable deep models
Tong, X.-Y.; Xia, G.-S.; Lu, Q.; Shen, H.; Li, S.; You, S.; and Zhang, L. 2020 · 2020
Earlier work this paper cites.
Crowdai mapping challenge
CrowdAI. 2018 · 2021
Earlier work this paper cites.
On creating benchmark dataset for aerial image interpretation: Reviews, guidances, and million-aid
Long, Y.; Xia, G.-S.; Li, S.; Yang, W.; Yang, M. Y.; Zhu, X. X.; Zhang, L.; and Li, D. 2021 · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. 2021 · 2021
Cited alongside, same era.
Sun, X.; Wang, P.; Yan, Z.; Xu, F.; Wang, R.; Diao, W.; Chen, J.; Li, J.; Feng, Y.; Xu, T.; Weinmann, M.; Hinz, S.; Wang, C.; and Fu, K. 2021 · 2021
Cited alongside, same era.
LoveDA: A remote sensing land-cover dataset for domain adaptive semantic segmentation
Wang, J.; Zheng, Z.; Ma, A.; Lu, X.; and Zhong, Y. 2021 · 2021
Cited alongside, same era.
Machine-learned regularization and polygonization of building segmentation masks
Gemini: a family of highly capable multimodal models
Team, G.; Anil, R.; Borgeaud, S.; Wu, Y.; Alayrac, J.-B.; Yu, J.; Soricut, R.; Schalkwyk, J.; Dai, A. M.; Hauth, A.; et al. 2023 · 2023
Later among the works it cites.
Enabling country-scale land cover mapping with meter-resolution satellite imagery
Tong, X.-Y.; Xia, G.-S.; and Zhu, X. X. 2023 · 2023
Later among the works it cites.
Cogvlm: Visual expert for pretrained language models
Wang, W.; Lv, Q.; Yu, W.; Hong, W.; Qi, J.; Wang, Y.; Ji, J.; Yang, Z.; Zhao, L.; Song, X.; et al. 2023 · 2023
Later among the works it cites.
RSVG: Exploring Data and Models for Visual Grounding on Remote Sensing Data
Zhan, Y.; Xiong, Z.; and Yuan, Y. 2023 · 2023
Later among the works it cites.
Rs5m: A large scale vision-language dataset for remote sensing vision-language foundation model
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zorzi, S.; Bittner, K.; and Fraundorfer, F. 2021 · 2021
Cited alongside, same era.
Cloud/shadow segmentation based on multi-level feature enhanced network for remote sensing imagery
Miao, S.; Xia, M.; Qian, M.; Zhang, Y.; Liu, J.; and Lin, H. 2022 · 2022
Cited alongside, same era.
CRTransSar: A visual transformer based on contextual joint representation learning for SAR ship detection
Xia, R.; Chen, J.; Huang, Z.; Wan, H.; Wu, B.; Sun, L.; Yao, B.; Xiang, H.; and Xing, M. 2022 · 2022
Cited alongside, same era.
METER-ML: A Multi-sensor Earth Observation Benchmark for Automated Methane Source Mapping
Zhu, B.; Lui, N.; Irvin, J.; Le, J.; Tadwalkar, S.; Wang, C.; Ouyang, Z.; Liu, F. Y.; Ng, A. Y.; and Jackson, R. B. 2022 · 2022
Cited alongside, same era.
Qwen-vl: A frontier large vision-language model with versatile abilities
Bai, J.; Bai, S.; Yang, S.; Wang, S.; Tan, S.; Wang, P.; Lin, J.; Zhou, C.; and Zhou, J. 2023 · 2023
Cited alongside, same era.
Minigpt-v2: large language model as a unified interface for vision-language multi-task learning
Chen, J.; Zhu, D.; Shen, X.; Li, X.; Liu, Z.; Zhang, P.; Krishnamoorthi, R.; Chandra, V.; Xiong, Y.; and Elhoseiny, M. 2023 · 2023
Cited alongside, same era.
Vicuna: An Open-Source Chatbot Impressing GPT-4 with 90%* ChatGPT Quality
Chiang, W.-L.; Li, Z.; Lin, Z.; Sheng, Y.; Wu, Z.; Zhang, H.; Zheng, L.; Zhuang, S.; Zhuang, Y.; Gonzalez, J. E.; Stoica, I.; and Xing, E. P. 2023 · 2023
Cited alongside, same era.
Rsgpt: A remote sensing vision language model and benchmark
Hu, Y.; Yuan, J.; Wen, C.; Lu, X.; and Li, X. 2023 · 2023
Cited alongside, same era.
Zhang, Z.; Zhao, T.; Guo, Y.; and Yin, J. 2023 · 2023
Later among the works it cites.
Rs-llava: A large vision-language model for joint captioning and question answering in remote sensing imagery
Bazi, Y.; Bashmal, L.; Al Rahhal, M. M.; Ricci, R.; and Melgani, F. 2024 · 2024
Closest in time.
Geochat: Grounded large vision-language model for remote sensing
Kuckreja, K.; Danish, M. S.; Naseer, M.; Das, A.; Khan, S.; and Khan, F. S. 2024 · 2024
Closest in time.
Visual instruction tuning
Liu, H.; Li, C.; Wu, Q.; and Lee, Y. J. 2024 · 2024
Closest in time.
Luo, J.; Pang, Z.; Zhang, Y.; Wang, T.; Wang, L.; Dang, B.; Lao, J.; Wang, J.; Chen, J.; Tan, Y.; et al. 2024 · 2024
Closest in time.
LHRS-Bot: Empowering Remote Sensing with VGI-Enhanced Large Multimodal Language Model
Muhtar, D.; Li, Z.; Gu, F.; Zhang, X.; and Xiao, P. 2024 · 2024
Closest in time.
Skyscript: A large and semantically diverse vision-language dataset for remote sensing
Wang, Z.; Prabha, R.; Huang, T.; Wu, J.; and Rajagopal, R. 2024 · 2024
Closest in time.
RS-GPT4V: A Unified Multimodal Instruction-Following Dataset for Remote Sensing Image Understanding
Xu, L.; Zhao, L.; Guo, W.; Li, Q.; Long, K.; Zou, K.; Wang, Y.; and Li, H. 2024 · 2024
Closest in time.
Zhan, Y.; Xiong, Z.; and Yuan, Y. 2024 · 2024
Closest in time.
Good at captioning, bad at counting: Benchmarking GPT-4V on Earth observation data
Zhang, C.; and Wang, S. 2024 · 2024
Closest in time.
Earthgpt: A universal multi-modal large language model for multi-sensor image comprehension in remote sensing domain
Zhang, W.; Cai, M.; Zhang, T.; Zhuang, Y.; and Mao, X. 2024 · 2024
Closest in time.