Fetching the paper…
Reading the bibliography…
Recent advances in prompt learning have allowed users to interact with artificial intelligence (AI) tools in multi-turn dialogue, enabling an interactive understanding of images.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Chin-Yew Lin · 2004
Earlier work this paper cites.
Meteor: An automatic metric for mt evaluation with improved correlation with human judgments
Satanjeev Banerjee and Alon Lavie · 2005
Earlier work this paper cites.
Satellite image classification via two-layer sparse coding with biased image representation
Dengxin Dai and Wen Yang · 2010
Earlier work this paper cites.
Bag-of-visual-words and spatial extensions for land-use classification
Yi Yang and Shawn Newsam · 2010
Earlier work this paper cites.
Referitgame: Referring to objects in photographs of natural scenes
Sahar Kazemzadeh, Vicente Ordonez, Mark Matten, and Tamara Berg · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Microsoft coco captions: Data collection and evaluation server
Xinlei Chen, Hao Fang, Tsung-Yi Lin, Ramakrishna Vedantam, Saurabh Gupta, Piotr Dollár, and C Lawrence Zitnick · 2015
Earlier work this paper cites.
Orientation robust object detection in aerial images using deep convolutional neural network
Haigang Zhu, Xiaogang Chen, Weiqun Dai, Kun Fu, Qixiang Ye, and Jianbin Jiao · 2015
Earlier work this paper cites.
Cider: Consensus-based image description evaluation
Ramakrishna Vedantam, C Lawrence Zitnick, and Devi Parikh · 2015
Earlier work this paper cites.
What’s the point: Semantic segmentation with point supervision
Amy Bearman, Olga Russakovsky, Vittorio Ferrari, and Li Fei-Fei · 2016
Earlier work this paper cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations, 2016
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A. Shamma, Michael S. Bernstein, and Fei-Fei Li · 2016
Earlier work this paper cites.
Modeling context in referring expressions, 2016
Licheng Yu, Patrick Poirson, Shan Yang, Alexander C. Berg, and Tamara L. Berg · 2016
Earlier work this paper cites.
Deep semantic understanding of high resolution remote sensing image
Bo Qu, Xuelong Li, Dacheng Tao, and Xiaoqiang Lu · 2016
Earlier work this paper cites.
Learning rotation-invariant convolutional neural networks for object detection in vhr optical remote sensing images
Gong Cheng, Peicheng Zhou, and Junwei Han · 2016
Earlier work this paper cites.
Spice: Semantic propositional image caption evaluation
Peter Anderson, Basura Fernando, Mark Johnson, and Stephen Gould · 2016
Earlier work this paper cites.
Deep learning in remote sensing: A comprehensive review and list of resources
Xiao Xiang Zhu, Devis Tuia, Lichao Mou, Gui-Song Xia, Liangpei Zhang, Feng Xu, and Friedrich Fraundorfer · 2017
Earlier work this paper cites.
Remote sensing image scene classification: Benchmark and state of the art
Gong Cheng, Junwei Han, and Xiaoqiang Lu · 2017
Earlier work this paper cites.
Accurate object localization in remote sensing images based on convolutional neural networks
Yang Long, Yiping Gong, Zhifeng Xiao, and Qing Liu · 2017
Earlier work this paper cites.
Random access memories: A new paradigm for target detection in high resolution aerial remote sensing images
Zhengxia Zou and Zhenwei Shi · 2017
Earlier work this paper cites.
A high resolution optical satellite image dataset for ship recognition and some new baselines
Zikun Liu, Liu Yuan, Lubin Weng, and Yiping Yang · 2017
Earlier work this paper cites.
Aid: A benchmark data set for performance evaluation of aerial scene classification
Gui-Song Xia, Jingwen Hu, Fan Hu, Baoguang Shi, Xiang Bai, Yanfei Zhong, Liangpei Zhang, and Xiaoqiang Lu · 2017
Earlier work this paper cites.
Exploring models and data for remote sensing image caption generation
Xiaoqiang Lu, Binqiang Wang, Xiangtao Zheng, and Xuelong Li · 2017
Earlier work this paper cites.
Scene classification with recurrent attention of vhr remote sensing images
Qi Wang, Shaoteng Liu, Jocelyn Chanussot, and Xuelong Li · 2018
Earlier work this paper cites.
Fully convolutional networks for multisource building extraction from an open aerial and satellite imagery data set
Shunping Ji, Shiqing Wei, and Meng Lu · 2018
Earlier work this paper cites.
Hierarchical and robust convolutional neural network for very high-resolution remote sensing object detection
Yuanlin Zhang, Yuan Yuan, Yachuang Feng, and Xiaoqiang Lu · 2019
Earlier work this paper cites.
isaid: A large-scale dataset for instance segmentation in aerial images
Syed Waqas Zamir, Aditya Arora, Akshita Gupta, Salman Khan, Guolei Sun, Fahad Shahbaz Khan, Fan Zhu, Ling Shao, Gui-Song Xia, and Xiang Bai · 2019
Earlier work this paper cites.
Generalized intersection over union: A metric and a loss for bounding box regression
Hamid Rezatofighi, Nathan Tsoi, JunYoung Gwak, Amir Sadeghian, Ian Reid, and Silvio Savarese · 2019
Earlier work this paper cites.
Semantic descriptions of high-resolution remote sensing images
Binqiang Wang, Xiaoqiang Lu, Xiangtao Zheng, and Xuelong Li · 2019
Earlier work this paper cites.
Description generation for remote sensing images using attribute attention mechanism
Xiangrong Zhang, Xin Wang, Xu Tang, Huiyu Zhou, and Chen Li · 2019
Earlier work this paper cites.
Autoprompt: Eliciting knowledge from language models with automatically generated prompts
Taylor Shin, Yasaman Razeghi, Robert L Logan IV, Eric Wallace, and Sameer Singh · 2020
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
Pcams: Weakly supervised semantic segmentation using point supervision
R Austin McEver and BS Manjunath · 2020
Earlier work this paper cites.
Hi-ucd: A large-scale dataset for urban semantic change detection in remote sensing imagery
Shiqi Tian, Ailong Ma, Zhuo Zheng, and Yanfei Zhong · 2020
Earlier work this paper cites.
Uavid: A semantic segmentation dataset for uav imagery
Ye Lyu, George Vosselman, Gui-Song Xia, Alper Yilmaz, and Michael Ying Yang · 2020
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Cited alongside, same era.
Object detection in aerial images: A large-scale benchmark and challenges
Jian Ding, Nan Xue, Gui-Song Xia, Xiang Bai, Wen Yang, Michael Yang, Serge Belongie, Jiebo Luo, Mihai Datcu, Marcello Pelillo, and Liangpei Zhang · 2021
Cited alongside, same era.
Detection and tracking meet drones challenge
Pengfei Zhu, Longyin Wen, Dawei Du, Xiao Bian, Heng Fan, Qinghua Hu, and Haibin Ling · 2021
Cited alongside, same era.
Fine-grained recognition for oriented ship against complex scenes in optical remote sensing images
Yaqi Han, Xinyi Yang, Tian Pu, and Zhenming Peng · 2021
Cited alongside, same era.
Shikra: Unleashing multimodal llm’s referential dialogue magic
Keqin Chen, Zhao Zhang, Weili Zeng, Richong Zhang, Feng Zhu, and Rui Zhao · 2023
Later among the works it cites.
Posterior instance injection detector for arbitrary-oriented object detection from optical remote sensing imagery
Tong Zhang, Yin Zhuang, He Chen, Guanqun Wang, Lihui Ge, Liang Chen, Hao Dong, and Lianlin Li · 2023
Later among the works it cites.
Learning remote sensing object detection with single point supervision
Shitian He, Huanxin Zou, Yingqian Wang, Boyang Li, Xu Cao, and Ning Jing · 2023
Later among the works it cites.
Point label meets remote sensing change detection: A consistency-aligned regional growth network
Leyuan Fang, Yiqi Jiang, Hongfeng Yu, Yingying Zhang, and Jun Yue · 2023
Later among the works it cites.
Dinov2: Learning robust visual features without supervision
Maxime Oquab, Timothée Darcet, Théo Moutakanni, Huy Vo, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, et al · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Urban modelling and semantic labelling benchmark
isprs · 2021
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale, 2021
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2021
Cited alongside, same era.
Learning to prompt for vision-language models
Kaiyang Zhou, Jingkang Yang, Chen Change Loy, and Ziwei Liu · 2022
Cited alongside, same era.
Visualgpt: Data-efficient adaptation of pretrained language models for image captioning
Jun Chen, Han Guo, Kai Yi, Boyang Li, and Mohamed Elhoseiny · 2022
Cited alongside, same era.
Flamingo: a visual language model for few-shot learning
Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, et al · 2022
Cited alongside, same era.
Learning to prompt for open-vocabulary object detection with vision-language model
Yu Du, Fangyun Wei, Zihe Zhang, Miaojing Shi, Yue Gao, and Guoqi Li · 2022
Cited alongside, same era.
Promptdet: Towards open-vocabulary detection using uncurated images
Chengjian Feng, Yujie Zhong, Zequn Jie, Xiangxiang Chu, Haibing Ren, Xiaolin Wei, Weidi Xie, and Lin Ma · 2022
Cited alongside, same era.
Later among the works it cites.
Set-of-mark prompting unleashes extraordinary visual grounding in gpt-4v
Jianwei Yang et al · 2023
Later among the works it cites.
Qwen-vl: A frontier large vision-language model with versatile abilities
Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou · 2023
Later among the works it cites.
Minigpt-v2: large language model as a unified interface for vision-language multi-task learning
Jun Chen, Deyao Zhu, Xiaoqian Shen, Xiang Li, Zechun Liu, Pengchuan Zhang, Raghuraman Krishnamoorthi, Vikas Chandra, Yunyang Xiong, and Mohamed Elhoseiny · 2023
Later among the works it cites.
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee · 2023
Later among the works it cites.
Glamm: Pixel grounding large multimodal model
Hanoona Rasheed, Muhammad Maaz, Sahal Shaji, Abdelrahman Shaker, Salman Khan, Hisham Cholakkal, Rao M Anwer, Erix Xing, Ming-Hsuan Yang, and Fahad S Khan · 2023
Later among the works it cites.
Learning visual prompts for guiding the attention of vision transformers
Razieh Rezaei, Masoud Jalili Sabet, Jindong Gu, Daniel Rueckert, Philip Torr, and Ashkan Khakzar · 2024
Closest in time.
Popeye: A unified visual-language model for multisource ship detection from remote sensing imagery
W. Zhang, M. Cai, T. Zhang, G. Lei, Y. Zhuang, and X. Mao · 2024
Closest in time.
Position-enhanced visual instruction tuning for multimodal large language models
Chi Chen, Ruoyu Qin, Fuwen Luo, Xiaoyue Mi, Peng Li, Maosong Sun, and Yang Liu · 2024
Closest in time.
Geochat: Grounded large vision-language model for remote sensing
Kartik Kuckreja, Muhammad Sohail Danish, Muzammal Naseer, Abhijit Das, Salman Khan, and Fahad Shahbaz Khan · 2024
Closest in time.
Rs5m and georsclip: A large scale vision-language dataset and a large vision-language model for remote sensing
Zilun Zhang, Tiancheng Zhao, Yulong Guo, and Jianwei Yin · 2024
Closest in time.
Draw-and-understand: Leveraging visual prompts to enable mllms to comprehend what you want
Weifeng Lin, Xinyu Wei, Ruichuan An, Peng Gao, Bocheng Zou, Yulin Luo, Siyuan Huang, Shanghang Zhang, and Hongsheng Li · 2024
Closest in time.
Earthgpt: A universal multi-modal large language model for multi-sensor image comprehension in remote sensing domain
Wei Zhang, Miaoxin Cai, Tong Zhang, Yin Zhuang, and Xuerui Mao · 2024
Closest in time.
A comparative performance analysis of popular deep learning models and segment anything model (sam) for river water segmentation in close-range remote sensing imagery
Armin Moghimi, Mario Welzel, Turgay Celik, and Torsten Schlurmann · 2024
Closest in time.
Rsprompter: Learning to prompt for remote sensing instance segmentation based on visual foundation model
Keyan Chen, Chenyang Liu, Hao Chen, Haotian Zhang, Wenyuan Li, Zhengxia Zou, and Zhenwei Shi · 2024
Closest in time.
Visionllm: Large language model is also an open-ended decoder for vision-centric tasks
Wenhai Wang, Zhe Chen, Xiaokang Chen, Jiannan Wu, Xizhou Zhu, Gang Zeng, Ping Luo, Tong Lu, Jie Zhou, Yu Qiao, et al · 2024
Closest in time.
Yang Zhan, Zhitong Xiong, and Yuan Yuan · 2024
Closest in time.
Lhrs-bot: Empowering remote sensing with vgi-enhanced large multimodal language model
Dilxat Muhtar, Zhenshi Li, Feng Gu, Xueliang Zhang, and Pengfeng Xiao · 2024
Closest in time.
Regionplc: Regional point-language contrastive learning for open-world 3d scene understanding
Jihan Yang, Runyu Ding, Weipeng Deng, Zhe Wang, and Xiaojuan Qi · 2024
Closest in time.
Hypersor: Context-aware graph hypernetwork for salient object ranking
Minglang Qiao, Mai Xu, Lai Jiang, Peng Lei, Shijie Wen, Yunjin Chen, and Leonid Sigal · 2024
Closest in time.
Earthvqa: Towards queryable earth via relational reasoning-based remote sensing visual question answering
Junjue Wang, Zhuo Zheng, Zihang Chen, Ailong Ma, and Yanfei Zhong · 2024
Closest in time.
A unified theory of scene representation learning and object representation learning, 2024
Takayuki Komatsu, Yoshiyuki Ohmura, and Yasuo Kuniyoshi · 2024
Closest in time.
Sphinx-x: Scaling data and parameters for a family of multi-modal large language models
Peng Gao, Renrui Zhang, Chris Liu, Longtian Qiu, Siyuan Huang, Weifeng Lin, Shitian Zhao, Shijie Geng, Ziyi Lin, Peng Jin, et al · 2024
Closest in time.
Multi-objective evolutionary multi-tasking band selection algorithm for hyperspectral image classification
Qijun Wang, Yong Liu, Ke Xu, Yanni Dong, Fan Cheng, Ye Tian, Bo Du, and Xingyi Zhang · 2024
Closest in time.
Mar20: A benchmark for military aircraft recognition in remote sensing images
YU Wenqi, CHENG Gong, WANG Meijun, YAO Yanqing, XIE Xingxing, YAO Xiwen, and HAN Junwei · 2024
Closest in time.
Samrs: Scaling-up remote sensing segmentation dataset with segment anything model
Di Wang, Jing Zhang, Bo Du, Minqiang Xu, Lin Liu, Dacheng Tao, and Liangpei Zhang · 2024
Closest in time.
Heterogeneous prototype distillation with support-query correlative guidance for few-shot remote sensing scene classification
Yin Zhuang, Yuqing Liu, Tong Zhang, Liang Chen, He Chen, and Lianlin Li · 2024
Closest in time.
Vip-llava: Making large multimodal models understand arbitrary visual prompts
Mu Cai, Haotian Liu, Siva Karthik Mustikovela, Gregory P Meyer, Yuning Chai, Dennis Park, and Yong Jae Lee · 2024
Closest in time.
Jinpeng Xu, Lin Bai, Xin Xie, and Lin Zhou · 2024
Closest in time.