Fetching the paper…
Reading the bibliography…
Multi-modal large language models (MLLMs) have achieved remarkable success in image- and region-level remote sensing (RS) image understanding tasks, such as image captioning, visual question answering, and visual grounding.
“Bleu: a method for automatic evaluation of machine translation”
Kishore Papineni, Salim Roukos, Todd Ward and Wei-Jing Zhu · 2002
Earlier work this paper cites.
“Rouge: A package for automatic evaluation of summaries”
Chin-Yew Lin · 2004
Earlier work this paper cites.
“METEOR: An Automatic Metric for MT Evaluation with Improved Correlation with Human Judgments”
Satanjeev Banerjee and Alon Lavie · 2005
Earlier work this paper cites.
“CIDEr: Consensus-based Image Description Evaluation”
Ramakrishna Vedantam, C. Zitnick and Devi Parikh · 2015
Earlier work this paper cites.
“AID: A Benchmark Data Set for Performance Evaluation of Aerial Scene Classification”
Gui-Song Xia et al · 2017
Earlier work this paper cites.
“Memory matching networks for genomic sequence classification”
Jack Lanchantin, Ritambhara Singh and Yanjun Qi · 2017
Earlier work this paper cites.
“A new approach to automatic memory banking using trace-based address mining”
Yuan Zhou, Khalid Al-Hawaj and Zhiru Zhang · 2017
Earlier work this paper cites.
“Unsupervised learning using pretrained CNN and associative memory bank”
Qun Liu and Supratik Mukhopadhyay · 2018
Earlier work this paper cites.
“Object detection in optical remote sensing images: A survey and a new benchmark”
Ke Li et al · 2019
Earlier work this paper cites.
“Eurosat: A novel dataset and deep learning benchmark for land use and land cover classification”
Patrick Helber, Benjamin Bischke, Andreas Dengel and Damian Borth · 2019
Earlier work this paper cites.
“NWPU-Crowd: A Large-Scale Benchmark for Crowd Counting and Localization”
Qi Wang, Junyu Gao, Wei Lin and Xuelong Li · 2020
Earlier work this paper cites.
“RSVQA: Visual Question Answering for Remote Sensing Data”
Sylvain Lobry, Diego Marcos, Jesse Murray and Devis Tuia · 2020
Earlier work this paper cites.
“Object Detection in Aerial Images: A Large-Scale Benchmark and Challenges”
Jian Ding et al · 2021
Earlier work this paper cites.
“FAIR1M: A Benchmark Dataset for Fine-grained Object Recognition in High-Resolution Remote Sensing Imagery”
Xian Sun et al · 2021
Earlier work this paper cites.
“Mamba: Multi-level aggregation via memory bank for video object detection”
Guanxiong Sun, Yang Hua, Guosheng Hu and Neil Robertson · 2021
Earlier work this paper cites.
“LoRA: Low-Rank Adaptation of Large Language Models”, 2021
Edward. Hu et al · 2021
Earlier work this paper cites.
“Semi-supervised semantic segmentation using unreliable pseudo-labels”
Yuchao Wang et al · 2022
Earlier work this paper cites.
“Weakly supervised semantic segmentation by pixel-to-prototype contrast”
Ye Du, Zehua Fu, Qingjie Liu and Yunhong Wang · 2022
Earlier work this paper cites.
“Memory-Based Cross-Image Contexts for Weakly Supervised Semantic Segmentation”
Junsong Fan and Zhaoxiang Zhang · 2022
Earlier work this paper cites.
“RSGPT: A Remote Sensing Vision Language Model and Benchmark”
Yuan Hu et al · 2023
Cited alongside, same era.
“RSVG: Exploring Data and Models for Visual Grounding on Remote Sensing Data”
Yang Zhan, Zhitong Xiong and Yuan Yuan · 2023
Cited alongside, same era.
“SAMRS: Scaling-up Remote Sensing Segmentation Dataset with Segment Anything Model”
Di Wang et al · 2023
Cited alongside, same era.
“Visual Instruction Tuning”
Haotian Liu, Chunyuan Li, Qingyang Wu and Yong Lee · 2023
Cited alongside, same era.
“Improved Baselines with Visual Instruction Tuning”
Haotian Liu, Chunyuan Li, Yuheng Li and Yong Lee · 2023
Cited alongside, same era.
“MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models”
Yang Zhan, Zhitong Xiong and Yuan Yuan · 2024
Later among the works it cites.
“Earthgpt: A universal multi-modal large language model for multi-sensor image comprehension in remote sensing domain”
Wei Zhang et al · 2024
Later among the works it cites.
“BB-GeoGPT: A framework for learning a large language model for geographic information science”
Yifan Zhang et al · 2024
Later among the works it cites.
“Lhrs-bot: Empowering remote sensing with vgi-enhanced large multimodal language model”
Dilxat Muhtar et al · 2024
Later among the works it cites.
“MTP: Advancing remote sensing foundation model via multi-task pretraining”
Di Wang et al · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Deyao Zhu et al · 2023
Cited alongside, same era.
“MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning”
Jun Chen et al · 2023
Cited alongside, same era.
Jinze Bai et al · 2023
Cited alongside, same era.
“Cogvlm: Visual expert for pretrained language models”
Weihan Wang et al · 2023
Cited alongside, same era.
“InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning”, 2023
Wenliang Dai et al · 2023
Cited alongside, same era.
“Otter: A Multi-Modal Model with In-Context Instruction Tuning”, 2023
Bo Li et al · 2023
Cited alongside, same era.
“Instruction Tuning with GPT-4”, 2023
Baolin Peng et al · 2023
Cited alongside, same era.
Junwei Luo et al · 2024
Later among the works it cites.
“TEOChat: A Large Vision-Language Assistant for Temporal Earth Observation Data”
Jeremy Irvin et al · 2024
Later among the works it cites.
“EarthMarker: A Visual Prompting Multi-modal Large Language Model for Remote Sensing”
Wei Zhang et al · 2024
Later among the works it cites.
“RRSIS: Referring Remote Sensing Image Segmentation”
Zhenghang Yuan, Lichao Mou, Yuansheng Hua and Xiao Zhu · 2024
Later among the works it cites.
“Rotated Multi-Scale Interaction Network for Referring Remote Sensing Image Segmentation”, 2024
Sihan Liu et al · 2024
Later among the works it cites.
“VRSBench: A Versatile Vision-Language Benchmark Dataset for Remote Sensing Image Understanding”
Jian Xiang and Mohamed Elhoseiny · 2024
Later among the works it cites.
“LLaVA-NeXT: Improved reasoning, OCR, and world knowledge”, 2024
Haotian Liu et al · 2024
Later among the works it cites.
“LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention”, 2024
Renrui Zhang et al · 2024
Later among the works it cites.
“PixelLM: Pixel Reasoning with Large Multimodal Model”, 2024
Zhongwei Ren et al · 2024
Later among the works it cites.
“Point Cloud Classification via Learnable Memory Bank”
Lisa Liu, William Wang and Pingping Cai · 2024
Later among the works it cites.
“SAM 2: Segment Anything in Images and Videos”, 2024
Nikhila Ravi et al · 2024
Later among the works it cites.
Xin Guo et al · 2024
Later among the works it cites.
“Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models”, 2024
Yanwei Li et al · 2024
Later among the works it cites.
“GeoPixel: Pixel Grounding Large Multimodal Model in Remote Sensing”, 2025
Akashah Shabbir et al · 2025
Closest in time.