Fetching the paper…
Reading the bibliography…
Large Vision-Language Models (VLMs) have demonstrated impressive performance on complex tasks involving visual input with natural language instructions.
Bleu: a method for automatic evaluation of machine translation
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu · 2002
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
C.-Y. Lin · 2004
Earlier work this paper cites.
METEOR: An automatic metric for MT evaluation with improved correlation with human judgments
S. Banerjee and A. Lavie · 2005
Earlier work this paper cites.
Wikidata: a free collaborative knowledgebase
D. Vrandečić and M. Krötzsch · 2014
Earlier work this paper cites.
Cider: Consensus-based image description evaluation
R. Vedantam, C. L. Zitnick, and D. Parikh · 2015
Earlier work this paper cites.
Combining satellite imagery and machine learning to predict poverty
N. Jean, M. Burke, M. Xie, W. M. Davis, D. B. Lobell, and S. Ermon · 2016
Earlier work this paper cites.
A large contextual dataset for classification, detection and counting of cars with deep learning
T. N. Mundhenk, G. Konjevod, W. A. Sakla, and K. Boakye · 2016
Earlier work this paper cites.
Exploring models and data for remote sensing image caption generation
X. Lu, B. Wang, X. Zheng, and X. Li · 2017
Earlier work this paper cites.
Functional map of the world
G. Christie, N. Fendley, J. Wilson, and R. Mukherjee · 2018
Earlier work this paper cites.
Datasheets for datasets
T. Gebru, J. Morgenstern, B. Vecchione, J. W. Vaughan, H. M. Wallach, H. Daumé, and K. Crawford · 2018
Earlier work this paper cites.
Patternnet: A benchmark dataset for performance evaluation of remote sensing image retrieval
W. Zhou, S. Newsam, C. Li, and Z. Shao · 2018
Earlier work this paper cites.
Mmrotate: A rotated object detection benchmark using pytorch
Y. Zhou, X. Yang, G. Zhang, J. Wang, Y. Liu, L. Hou, X. Jiang, X. Liu, J. Yan, C. Lyu, W. Zhang, and K. Chen · 2018
Earlier work this paper cites.
Improving the precision and accuracy of animal population estimates with aerial image object detection
J. A. Eikelboom, J. Wind, E. van de Ven, L. M. Kenana, B. Schroder, H. J. de Knegt, F. van Langevelde, and H. H. Prins · 2019
Earlier work this paper cites.
xbd: A dataset for assessing building damage from satellite imagery
R. Gupta, R. Hosfelt, S. Sajeev, N. Patel, B. Goodman, J. Doshi, E. Heim, H. Choset, and M. Gaston · 2019
Earlier work this paper cites.
Object detection in optical remote sensing images: A survey and a new benchmark
K. Li, G. Wan, G. Cheng, L. Meng, and J. Han · 2019
Earlier work this paper cites.
Bigearthnet: A large-scale benchmark archive for remote sensing image understanding
G. Sumbul, M. Charfuelan, B. Demir, and V. Markl · 2019
Earlier work this paper cites.
Counting cows: Tracking illegal cattle ranching from high-resolution satellite imagery
I. Laradji, P. Rodriguez, F. Kalaitzis, D. Vazquez, R. Young, E. Davey, and A. Lacoste · 2020
Earlier work this paper cites.
Meta-learning for few-shot land cover classification
M. Rußwurm*, S. Wang*, M. Körner, and D. B. Lobell · 2020
Cited alongside, same era.
Crop yield prediction using machine learning: A systematic literature review
T. van Klompenburg, A. Kassahun, and C. Catal · 2020
Cited alongside, same era.
A benchmark dataset for individual tree crown delineation in co-registered airborne rgb, lidar and hyperspectral imagery from the national ecological observation network
B. G. Weinstein, S. J. Graves, S. Marconi, A. Singh, A. Zare, D. Stewart, S. A. Bohlman, and E. P. White · 2020
Cited alongside, same era.
Google landmarks dataset v2 - a large-scale benchmark for instance-level recognition and retrieval
T. Weyand, A. Araújo, B. Cao, and J. Sim · 2020
Cited alongside, same era.
CLIPScore: a reference-free evaluation metric for image captioning
J. Hessel, A. Holtzman, M. Forbes, R. L. Bras, and Y. Choi · 2021
Cited alongside, same era.
Rsgpt: A remote sensing vision language model and benchmark
Y. Hu, J. Yuan, C. Wen, X. Lu, and X. Li · 2023
Later among the works it cites.
Geochat: Grounded large vision-language model for remote sensing
K. Kuckreja, M. S. Danish, M. Naseer, A. Das, S. Khan, and F. S. Khan · 2023
Later among the works it cites.
OpenAI · 2023
Later among the works it cites.
Glamm: Pixel grounding large multimodal model
H. Rasheed, M. Maaz, S. Shaji, A. Shaker, S. Khan, H. Cholakkal, R. M. Anwer, E. Xing, M.-H. Yang, and F. S. Khan · 2023
Later among the works it cites.
Charting new territories: Exploring the geographic and geospatial capabilities of multimodal llms
J. Roberts, T. Lüddecke, R. Sheikh, K. Han, and S. Albanie · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
WILDS: A benchmark of in-the-wild distribution shifts
P. W. Koh, S. Sagawa, H. Marklund, S. M. Xie, M. Zhang, A. Balsubramani, W. Hu, M. Yasunaga, R. L. Phillips, I. Gao, T. Lee, E. David, I. Stavness, W. Guo, B. A. Earnshaw, I. S. Haque, S. Beery, J. Leskovec, A. Kundaje, E. Pierson, S. Levine, C. Finn, and P. Liang · 2021
Cited alongside, same era.
Continental-scale building detection from high resolution satellite imagery
W. Sirko, S. Kashubin, M. Ritter, A. Annkah, Y. S. E. Bouchareb, Y. Dauphin, D. Keysers, M. Neumann, M. Cisse, and J. Quinn · 2021
Cited alongside, same era.
A benchmark dataset for canopy crown detection and delineation in co-registered airborne rgb, lidar and hyperspectral imagery from the national ecological observation network
B. G. Weinstein, S. J. Graves, S. Marconi, A. Singh, A. Zare, D. Stewart, S. A. Bohlman, and E. P. White · 2021
Cited alongside, same era.
xview3-sar: Detecting dark fishing activity using synthetic aperture radar imagery
F. Paolo, T.-t. T. Lin, R. Gupta, B. Goodman, N. Patel, D. Kuster, D. Kroodsma, and J. Dunnmon · 2022
Cited alongside, same era.
Extending the wilds benchmark for unsupervised adaptation
S. Sagawa, P. W. Koh, T. Lee, I. Gao, S. M. Xie, K. Shen, A. Kumar, W. Hu, M. Yasunaga, H. Marklund, S. Beery, E. David, I. Stavness, W. Guo, J. Leskovec, K. Saenko, T. Hashimoto, S. Levine, C. Finn, and P. Liang · 2022
Cited alongside, same era.
Robust fine-tuning of zero-shot models
M. Wortsman, G. Ilharco, M. Li, J. W. Kim, H. Hajishirzi, A. Farhadi, H. Namkoong, and L. Schmidt · 2022
Cited alongside, same era.
Qwen-vl: A frontier large vision-language model with versatile abilities
J. Bai, S. Bai, S. Yang, S. Wang, S. Tan, P. Wang, J. Lin, C. Zhou, and J. Zhou · 2023
Cited alongside, same era.
W. Shi, A. Ajith, M. Xia, Y. Huang, D. Liu, T. Blevins, D. Chen, and L. Zettlemoyer · 2023
Later among the works it cites.
On the promises and challenges of multimodal foundation models for geographical, environmental, agricultural, and urban planning applications
C. Tan, Q. Cao, Y. Li, J. Zhang, X. Yang, H. Zhao, Z. Wu, Z. Liu, H. Yang, N. Wu, T. Tang, X. Ye, L. Chai, N. Liu, C. Li, L. Mu, T. Liu, and G. Mai · 2023
Later among the works it cites.
Gemini technical report
G. G. Team · 2023
Later among the works it cites.
Florence-2: Advancing a unified representation for a variety of vision tasks
B. Xiao, H. Wu, W. Xu, X. Dai, H. Hu, Y. Lu, M. Zeng, C. Liu, and L. Yuan · 2023
Later among the works it cites.
The dawn of lmms: Preliminary explorations with gpt-4v (ision)
Z. Yang, L. Li, K. Lin, J. Wang, C.-C. Lin, Z. Liu, and L. Wang · 2023
Later among the works it cites.
Mm-vet: Evaluating large multimodal models for integrated capabilities
W. Yu, Z. Yang, L. Li, J. Wang, K. Lin, Z. Liu, X. Wang, and L. Wang · 2023
Later among the works it cites.
Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi
X. Yue, Y. Ni, K. Zhang, T. Zheng, R. Liu, G. Zhang, S. Stevens, D. Jiang, W. Ren, Y. Sun, C. Wei, B. Yu, R. Yuan, R. Sun, M. Yin, B. Zheng, Z. Yang, Y. Liu, W. Huang, H. Sun, Y. Su, and W. Chen · 2023
Later among the works it cites.
Rsvg: Exploring data and models for visual grounding on remote sensing data
Y. Zhan, Z. Xiong, and Y. Yuan · 2023
Later among the works it cites.
Text2seg: Remote sensing image semantic segmentation via text-guided visual foundation models
J. Zhang, Z. Zhou, G. Mai, L. Mu, M. Hu, and S. Li · 2023
Later among the works it cites.
Skyeyegpt: Unifying remote sensing vision-language tasks via instruction tuning with large language model
Y. Zhan, Z. Xiong, and Y. Yuan · 2024
Closest in time.
Earthgpt: A universal multi-modal large language model for multi-sensor image comprehension in remote sensing domain
W. Zhang, M. Cai, T. Zhang, Y. Zhuang, and X. Mao · 2024
Closest in time.