Fetching the paper…
Reading the bibliography…
While numerous recent benchmarks focus on evaluating generic Vision-Language Models (VLMs), they do not effectively address the specific challenges of geospatial applications.
Deformable detr: Deformable transformers for end-to-end object detection
Xizhou Zhu, Weijie Su, Lewei Lu, Bin Li, Xiaogang Wang, and Jifeng Dai · 2010
Earlier work this paper cites.
Combining satellite imagery and machine learning to predict poverty
Neal Jean, Marshall Burke, Michael Xie, W Matthew Davis, David B Lobell, and Stefano Ermon · 2016
Earlier work this paper cites.
A large contextual dataset for classification, detection and counting of cars with deep learning
T. Nathan Mundhenk, Goran Konjevod, Wesam A. Sakla, and Kofi Boakye · 2016
Earlier work this paper cites.
Remote sensing image scene classification: Benchmark and state of the art
Gong Cheng, Junwei Han, and Xiaoqiang Lu · 2017
Earlier work this paper cites.
Patternnet: A benchmark dataset for performance evaluation of remote sensing image retrieval
Weixun Zhou, Shawn D. Newsam, Congmin Li, and Zhenfeng Shao · 2017
Earlier work this paper cites.
Deepglobe 2018: A challenge to parse the earth through satellite images
Ilke Demir, Krzysztof Koperski, David Lindenbaum, Guan Pang, Jing Huang, Saikat Basu, Forest Hughes, Devis Tuia, and Ramesh Raskar · 2018
Earlier work this paper cites.
Dota: A large-scale dataset for object detection in aerial images
Gui-Song Xia, Xiang Bai, Jian Ding, Zhen Zhu, Serge Belongie, Jiebo Luo, Mihai Datcu, Marcello Pelillo, and Liangpei Zhang · 2018
Earlier work this paper cites.
xbd: A dataset for assessing building damage from satellite imagery
Ritwik Gupta, Richard Hosfelt, Sandra Sajeev, Nirav Patel, Bryce Goodman, Jigar Doshi, Eric T. Heim, Howie Choset, and Matthew E. Gaston · 2019
Earlier work this paper cites.
Object detection in optical remote sensing images: A survey and a new benchmark
Ke Li, Gang Wan, Gong Cheng, Liqiu Meng, and Junwei Han · 2020
Earlier work this paper cites.
Airound and cv-brct: Novel multi-view datasets for scene classification
Gabriel L. S. Machado, Edemir Ferreira, Keiller Nogueira, Hugo N. Oliveira, Pedro H. T. Gama, and Jefersson A. dos Santos · 2020
Earlier work this paper cites.
Meta-learning for few-shot land cover classification
Marc Rußwurm, Sherrie Wang, Marco Korner, and David Lobell · 2020
Earlier work this paper cites.
Rareplanes dataset, 2020
Jacob Shermeyer, Thomas Hossler, Adam Van Etten, Daniel Hogan, Ryan Lewis, and Daeil Kim · 2020
Earlier work this paper cites.
Bertscore: Evaluating text generation with bert
Tianyi Zhang*, Varsha Kishore*, Felix Wu*, Kilian Q. Weinberger, and Yoav Artzi · 2020
Earlier work this paper cites.
Forest damages – larch casebearer 1.0, 2021
Swedish Forest Agency · 2021
Earlier work this paper cites.
Anchor-free oriented proposal generator for object detection
Gong Cheng, Jiabao Wang, Ke Li, Xingxing Xie, Chunbo Lang, Yanqing Yao, and Junwei Han · 2021
Earlier work this paper cites.
A public dataset for fine-grained ship classification in optical remote sensing images
Yanghua Di, Zhiguo Jiang, and Haopeng Zhang · 2021
Earlier work this paper cites.
Panoptic segmentation of satellite image time series with convolutional temporal attention networks
Vivien Sainte Fare Garnot and Loic Landrieu · 2021
Earlier work this paper cites.
Marine debris dataset for object detection in planetscope imagery, 2021
A. Shah, L. Thomas, and M. Maskey · 2021
Earlier work this paper cites.
Fair1m: A benchmark dataset for fine-grained object recognition in high-resolution remote sensing imagery
Xian Sun, Peijin Wang, Zhiyuan Yan, F. Xu, Ruiping Wang, W. Diao, Jin Chen, Jihao Li, Yingchao Feng, Tao Xu, M. Weinmann, S. Hinz, Cheng Wang, and K. Fu · 2021
Cited alongside, same era.
Segformer: Simple and efficient design for semantic segmentation with transformers
Enze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar, Jose M Alvarez, and Ping Luo · 2021
Cited alongside, same era.
Broaden the vision: Geo-diverse visual commonsense reasoning
Da Yin, Liunian Harold Li, Ziniu Hu, Nanyun Peng, and Kai-Wei Chang · 2021
Cited alongside, same era.
Synthesizing optical and sar imagery from land cover maps and auxiliary raster data
Gerald Baier, Antonin Deschemps, Michael Schmitt, and Naoto Yokoya · 2022
Cited alongside, same era.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al · 2023
Junwei Luo, Zhen Pang, Yongjun Zhang, Tingzhu Wang, Linlin Wang, Bo Dang, Jiangwei Lao, Jian Wang, Jingdong Chen, Yihua Tan, et al · 2024
Closest in time.
Lhrs-bot: Empowering remote sensing with vgi-enhanced large multimodal language model
Dilxat Muhtar, Zhenshi Li, Feng Gu, Xueliang Zhang, and Pengfeng Xiao · 2024
Closest in time.
Benchmarking vision language models for cultural understanding
Shravan Nayak, Kanishk Jain, Rabiul Awal, Siva Reddy, Sjoerd van Steenkiste, Lisa Anne Hendricks, Karolina Stańczak, and Aishwarya Agrawal · 2024
Closest in time.
Hello gpt-4o, 2024
OpenAI · 2024
Closest in time.
Glamm: Pixel grounding large multimodal model
Hanoona Rasheed, Muhammad Maaz, Sahal Shaji, Abdelrahman Shaker, Salman Khan, Hisham Cholakkal, Rao M Anwer, Eric Xing, Ming-Hsuan Yang, and Fahad S Khan · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Qwen-vl: A frontier large vision-language model with versatile abilities
Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou · 2023
Cited alongside, same era.
Sparks of artificial general intelligence: Early experiments with gpt-4. arxiv
Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, et al · 2023
Cited alongside, same era.
Rsgpt: A remote sensing vision language model and benchmark
Yuan Hu, Jianlong Yuan, Congcong Wen, Xiaonan Lu, and Xiang Li · 2023
Cited alongside, same era.
Seed-bench: Benchmarking multimodal llms with generative comprehension
Bohao Li, Rui Wang, Guangzhi Wang, Yuying Ge, Yixiao Ge, and Ying Shan · 2023
Cited alongside, same era.
Ziyi Lin, Chris Liu, Renrui Zhang, Peng Gao, Longtian Qiu, Han Xiao, Han Qiu, Chen Lin, Wenqi Shao, Keqin Chen, et al · 2023
Cited alongside, same era.
Firerisk: A remote sensing dataset for fire risk assessment with benchmarks using supervised and self-supervised learning, 2023
Shuchang Shen, Sachith Seneviratne, Xinye Wanyan, and Michael Kirley · 2023
Cited alongside, same era.
Fpcd: An open aerial vhr dataset for farm pond change detection
Chintan Tundia, Rajiv Kumar, Om Damani, and G. Sivakumar · 2023
Cited alongside, same era.
Quakeset: A dataset and low-resource models to monitor earthquakes through sentinel-1
Daniele Rege Cambrin and Paolo Garza · 2024
Closest in time.
Earthdial: Turning multi-sensory earth observations to interactive dialogues
Sagar Soni, Akshay Dudhane, Hiyam Debary, Mustansar Fiaz, Muhammad Akhtar Munir, Muhammad Sohail Danish, Paolo Fraccaro, Campbell D Watson, Levente J Klein, Fahad Shahbaz Khan, et al · 2024
Closest in time.
Mar20: A benchmark for military aircraft recognition in remote sensing images
YU Wenqi, CHENG Gong, WANG Meijun, YAO Yanqing, XIE Xingxing, YAO Xiwen, and HAN Junwei · 2024
Closest in time.
Florence-2: Advancing a unified representation for a variety of vision tasks
Bin Xiao, Haiping Wu, Weijian Xu, Xiyang Dai, Houdong Hu, Yumao Lu, Michael Zeng, Ce Liu, and Lu Yuan · 2024
Closest in time.
Exploring diverse in-context configurations for image captioning
Xu Yang, Yongliang Wu, Mingzhuo Yang, Haokun Chen, and Xin Geng · 2024
Closest in time.
Lamm: Language-assisted multi-modal instruction-tuning dataset, framework, and benchmark
Zhenfei Yin, Jiong Wang, Jianjian Cao, Zhelun Shi, Dingning Liu, Mukai Li, Xiaoshui Huang, Zhiyong Wang, Lu Sheng, Lei Bai, et al · 2024
Closest in time.
Rrsis: Referring remote sensing image segmentation
Zhenghang Yuan, Lichao Mou, Yuansheng Hua, and Xiao Xiang Zhu · 2024
Closest in time.
Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi
Xiang Yue, Yuansheng Ni, Kai Zhang, Tianyu Zheng, Ruoqi Liu, Ge Zhang, Samuel Stevens, Dongfu Jiang, Weiming Ren, Yuxuan Sun, et al · 2024
Closest in time.
Yang Zhan, Zhitong Xiong, and Yuan Yuan · 2024
Closest in time.
Good at captioning, bad at counting: Benchmarking gpt-4v on earth observation data
Chenhui Zhang and Sherrie Wang · 2024
Closest in time.
Panoptic perception: A novel task and fine-grained dataset for universal remote sensing image interpretation
Danpei Zhao, Bo Yuan, Ziqiang Chen, Tian Li, Zhuoran Liu, Wentao Li, and Yue Gao · 2024
Closest in time.
Mmbench: Is your multi-modal model an all-around player?
Yuan Liu, Haodong Duan, Yuanhan Zhang, Bo Li, Songyang Zhang, Wangbo Zhao, Yike Yuan, Jiaqi Wang, Conghui He, Ziwei Liu, et al · 2025
Closest in time.