Fetching the paper…
Reading the bibliography…
The domain gap between remote sensing imagery and natural images has recently received widespread attention and Vision-Language Models (VLMs) have demonstrated excellent generalization performance in remote sensing multimodal tasks.
R. Caruna, “Multitask learning: A knowledge-based source of inductive bias,” in Machine learning: Proceedings of the tenth international conference , 1993, pp. 41–48
1993
Earlier work this paper cites.
R. Caruana, “Multitask learning,” Machine learning , vol. 28, pp. 41–75, 1997
1997
Earlier work this paper cites.
K. Kafle and C. Kanan, “Answer-type prediction for visual question answering,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 4976–4984
2016
Earlier work this paper cites.
G. Cheng, J. Han, and X. Lu, “Remote sensing image scene classification: Benchmark and state of the art,” Proceedings of the IEEE , vol. 105, no. 10, pp. 1865–1883, 2017
2017
Earlier work this paper cites.
G.-S. Xia, X. Bai, J. Ding, Z. Zhu, S. Belongie, J. Luo, M. Datcu, M. Pelillo, and L. Zhang, “Dota: A large-scale dataset for object detection in aerial images,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2018, pp. 3974–3983
2018
Earlier work this paper cites.
K. Marino, M. Rastegari, A. Farhadi, and R. Mottaghi, “Ok-vqa: A visual question answering benchmark requiring external knowledge,” in Proceedings of the IEEE/cvf conference on computer vision and pattern recognition , 2019, pp. 3195–3204
2019
Earlier work this paper cites.
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell et al. , “Language models are few-shot learners,” Advances in neural information processing systems , vol. 33, pp. 1877–1901, 2020
2020
Earlier work this paper cites.
S. Lobry, D. Marcos, J. Murray, and D. Tuia, “Rsvqa: Visual question answering for remote sensing data,” IEEE Transactions on Geoscience and Remote Sensing , vol. 58, no. 12, pp. 8555–8566, 2020
2020
Earlier work this paper cites.
L. Mou, Y. Hua, P. Jin, and X. X. Zhu, “Era: A data set and deep learning benchmark for event recognition in aerial videos [software and data sets],” IEEE Geoscience and Remote Sensing Magazine , vol. 8, no. 4, pp. 125–133, 2020
2020
Earlier work this paper cites.
H. Chen and Z. Shi, “A spatial-temporal attention-based method and a new dataset for remote sensing image change detection,” Remote Sensing , vol. 12, no. 10, p. 1662, 2020
2020
Earlier work this paper cites.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark et al. , “Learning transferable visual models from natural language supervision,” in International conference on machine learning . PMLR, 2021, pp. 8748–8763
2021
Earlier work this paper cites.
M. Rahnemoonfar, T. Chowdhury, A. Sarkar, D. Varshney, M. Yari, and R. R. Murphy, “Floodnet: A high resolution aerial imagery dataset for post flood scene understanding,” IEEE Access , vol. 9, pp. 89 644–89 654, 2021
2021
Earlier work this paper cites.
P. Jin, L. Mou, Y. Hua, G.-S. Xia, and X. X. Zhu, “Temporal relations matter: A two-pathway network for aerial video recognition,” in 2021 IEEE International Geoscience and Remote Sensing Symposium IGARSS . IEEE, 2021, pp. 8221–8224
2021
Earlier work this paper cites.
J.-B. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, Y. Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynolds et al. , “Flamingo: a visual language model for few-shot learning,” Advances in neural information processing systems , vol. 35, pp. 23 716–23 736, 2022
2022
Earlier work this paper cites.
J. Li, D. Li, C. Xiong, and S. Hoi, “Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation,” in International conference on machine learning . PMLR, 2022, pp. 12 888–12 900
2022
Earlier work this paper cites.
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray et al. , “Training language models to follow instructions with human feedback,” Advances in neural information processing systems , vol. 35, pp. 27 730–27 744, 2022
2022
Earlier work this paper cites.
C. Schuhmann, R. Beaumont, R. Vencu, C. Gordon, R. Wightman, M. Cherti, T. Coombes, A. Katta, C. Mullis, M. Wortsman et al. , “Laion-5b: An open large-scale dataset for training next generation image-text models,” Advances in Neural Information Processing Systems , vol. 35, pp. 25 278–25 294, 2022
2022
Earlier work this paper cites.
Y. Bazi, M. M. Al Rahhal, M. L. Mekhalfi, M. A. Al Zuair, and F. Melgani, “Bi-modal transformer-based approach for visual question answering in remote sensing imagery,” IEEE Transactions on Geoscience and Remote Sensing , vol. 60, pp. 1–11, 2022
2022
Earlier work this paper cites.
P. Jin, L. Mou, Y. Hua, G.-S. Xia, and X. X. Zhu, “Futh-net: fusing temporal relations and holistic features for aerial video classification,” IEEE Transactions on Geoscience and Remote Sensing , vol. 60, pp. 1–13, 2022
2022
Earlier work this paper cites.
C. Liu, R. Zhao, H. Chen, Z. Zou, and Z. Shi, “Remote sensing image change captioning with dual-branch transformers: A new method and a large scale dataset,” IEEE Transactions on Geoscience and Remote Sensing , vol. 60, pp. 1–20, 2022
2022
Cited alongside, same era.
Q. Cheng, H. Huang, Y. Xu, Y. Zhou, H. Li, and Z. Wang, “Nwpu-captions dataset and mlca-net for remote sensing image captioning,” IEEE Transactions on Geoscience and Remote Sensing , vol. 60, pp. 1–19, 2022
2022
Cited alongside, same era.
F. Yang, J. Zhang, Y. Zhao, A. Qin, and C. Gao, “Multiscale spatio-temporal network for aerial video event recognition,” in IGARSS 2022-2022 IEEE International Geoscience and Remote Sensing Symposium . IEEE, 2022, pp. 7835–7838
2022
Cited alongside, same era.
G. Cheng, J. Wang, K. Li, X. Xie, C. Lang, Y. Yao, and J. Han, “Anchor-free oriented proposal generator for object detection,” IEEE Transactions on Geoscience and Remote Sensing , vol. 60, pp. 1–11, 2022
2022
2023
Later among the works it cites.
Z. Zhang, L. Jiao, L. Li, X. Liu, P. Chen, F. Liu, Y. Li, and Z. Guo, “A spatial hierarchical reasoning network for remote sensing visual question answering,” IEEE Transactions on Geoscience and Remote Sensing , vol. 61, pp. 1–15, 2023
2023
Later among the works it cites.
H. Liu, C. Li, Q. Wu, and Y. J. Lee, “Visual instruction tuning,” Advances in neural information processing systems , vol. 36, 2024
2024
Closest in time.
J. Lin, H. Yin, W. Ping, P. Molchanov, M. Shoeybi, and S. Han, “Vila: On pre-training for visual language models,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 26 689–26 699
2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
X. Sun, P. Wang, Z. Yan, F. Xu, R. Wang, W. Diao, J. Chen, J. Li, Y. Feng, T. Xu et al. , “Fair1m: A benchmark dataset for fine-grained object recognition in high-resolution remote sensing imagery,” ISPRS Journal of Photogrammetry and Remote Sensing , vol. 184, pp. 116–130, 2022
2022
Cited alongside, same era.
2023
Cited alongside, same era.
J. Li, D. Li, S. Savarese, and S. Hoi, “Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,” in International conference on machine learning . PMLR, 2023, pp. 19 730–19 742
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
M. Zhang, F. Chen, and B. Li, “Multi-step question-driven visual question answering for remote sensing,” IEEE Transactions on Geoscience and Remote Sensing , 2023
2023
Cited alongside, same era.
S. Chang and P. Ghamisi, “Changes to captions: An attentive network for remote sensing change captioning,” IEEE Transactions on Image Processing , 2023
2023
Cited alongside, same era.
C. Liu, J. Yang, Z. Qi, Z. Zou, and Z. Shi, “Progressive scale-aware network for remote sensing image change captioning,” in IGARSS 2023-2023 IEEE International Geoscience and Remote Sensing Symposium . IEEE, 2023, pp. 6668–6671
2023
Cited alongside, same era.
C. Li, C. Wong, S. Zhang, N. Usuyama, H. Liu, J. Yang, T. Naumann, H. Poon, and J. Gao, “Llava-med: Training a large language-and-vision assistant for biomedicine in one day,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
K. Kuckreja, M. S. Danish, M. Naseer, A. Das, S. Khan, and F. S. Khan, “Geochat: Grounded large vision-language model for remote sensing,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 27 831–27 840
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Y. Bazi, L. Bashmal, M. M. Al Rahhal, R. Ricci, and F. Melgani, “Rs-llava: A large vision-language model for joint captioning and question answering in remote sensing imagery,” Remote Sensing , vol. 16, no. 9, p. 1477, 2024
2024
Closest in time.
W. Zhang, M. Cai, T. Zhang, Y. Zhuang, and X. Mao, “Earthgpt: A universal multi-modal large language model for multi-sensor image comprehension in remote sensing domain,” IEEE Transactions on Geoscience and Remote Sensing , 2024
2024
Closest in time.
D. Luo, J. Huang, S. Gong, H. Jin, and Y. Liu, “Zero-shot video moment retrieval from frozen vision-language models,” in Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision , 2024, pp. 5464–5473
2024
Closest in time.
M. Springstein, S. Schneider, J. Rahnama, J. Stalter, M. Kristen, E. Müller-Budack, and R. Ewerth, “Visual narratives: Large-scale hierarchical classification of art-historical images,” in Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision , 2024, pp. 7220–7230
2024
Closest in time.
H. Liu, C. Li, Y. Li, and Y. J. Lee, “Improved baselines with visual instruction tuning,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 26 296–26 306
2024
Closest in time.
2024
Closest in time.
S. Jaradat, R. Nayak, A. Paz, H. I. Ashqar, and M. Elhenawy, “Multitask learning for crash analysis: A fine-tuned llm framework using twitter data,” Smart Cities , vol. 7, no. 5, pp. 2422–2465, 2024
2024
Closest in time.
B. Liu, C. Chen, Z. Gong, C. Liao, H. Wang, Z. Lei, M. Liang, D. Chen, M. Shen, H. Zhou et al. , “Mftcoder: Boosting code llms with multitask fine-tuning,” in Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining , 2024, pp. 5430–5441
2024
Closest in time.