Fetching the paper…
Reading the bibliography…
Image captioning and cross-modal retrieval are examples of tasks that involve the joint analysis of visual and linguistic information.
S. Hochreiter and J. Schmidhuber, “Long Short-term Memory,” Neural computation , vol. 9, pp. 1735–1780, 1997
1997
Earlier work this paper cites.
2014
Earlier work this paper cites.
B. Qu, X. Li, D. Tao, and X. Lu, “Deep semantic understanding of high resolution remote sensing image,” in Proceedings of the International Conference on Computer, Information and Telecommunication Systems , 2016
2016
Earlier work this paper cites.
Z. Shi and Z. Zou, “Can a Machine Generate Humanlike Language Descriptions for a Remote Sensing Image?” IEEE Transactions on Geoscience and Remote Sensing , vol. 55, 2017
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in neural information processing systems , vol. 30, 2017
2017
Earlier work this paper cites.
X. Lu, B. Wang, X. Zheng, and X. Li, “Exploring models and data for remote sensing image caption generation,” IEEE Transactions on Geoscience and Remote Sensing , vol. 56, 2018
2018
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever et al. , “Language models are unsupervised multitask learners,” Open AI Blog , 2019
2019
Earlier work this paper cites.
N. Reimers and I. Gurevych, “Sentence-BERT: Sentence embeddings using Siamese BERT-networks,” in Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP) , K. Inui, J. Jiang, V. Ng, and X. Wan, Eds. Hong Kong, China: Association for Computational Linguistics, Nov. 2019, pp. 3982–3992
2019
Earlier work this paper cites.
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga et al. , “Pytorch: An imperative style, high-performance deep learning library,” Advances in neural information processing systems , vol. 32, 2019
2019
Earlier work this paper cites.
S. Lobry, D. Marcos, J. Murray, and D. Tuia, “RSVQA: Visual Question Answering for Remote Sensing Data,” IEEE Transactions on Geoscience and Remote Sensing , vol. 58, 2020
2020
Earlier work this paper cites.
Y. Li, S. Fang, L. Jiao, R. Liu, and R. Shang, “A multi-level attention model for remote sensing image captions,” Remote Sensing , vol. 12, no. 6, 2020
2020
Earlier work this paper cites.
Z. Yuan, X. Li, and Q. Wang, “Exploring multi-level attention and semantic relationship for remote sensing image captioning,” IEEE Access , vol. 8, pp. 2608–2620, 2020
2020
Earlier work this paper cites.
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell et al. , “Language models are few-shot learners,” Advances in neural information processing systems , vol. 33, pp. 1877–1901, 2020
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
D. Tuia, R. Roscher, J. D. Wegner, N. Jacobs, X. Zhu, and G. Camps-Valls, “Toward a collective agenda on AI for earth science data analysis,” IEEE Geoscience and Remote Sensing Magazine , vol. 9, no. 2, 2021
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
W. Huang, Q. Wang, and X. Li, “Denoising-based multiscale feature fusion for remote sensing image captioning,” IEEE Geoscience and Remote Sensing Letters , vol. 18, no. 3, pp. 436–440, 2021
2021
Earlier work this paper cites.
R. Zhao, Z. Shi, and Z. Zou, “High-resolution remote sensing image captioning based on structured attention,” IEEE Transactions on Geoscience and Remote Sensing , vol. 60, pp. 1–14, 2021
2021
Earlier work this paper cites.
M. Tsimpoukelli, J. L. Menick, S. Cabi, S. Eslami, O. Vinyals, and F. Hill, “Multimodal few-shot learning with frozen language models,” Advances in Neural Information Processing Systems , vol. 34, pp. 200–212, 2021
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
B. Lester, R. Al-Rfou, and N. Constant, “The power of scale for parameter-efficient prompt tuning,” in Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, EMNLP 2021, Virtual Event / Punta Cana, Dominican Republic, 7-11 November, 2021 . Association for Computational Linguistics, 2021, pp. 3045–3059
2021
Earlier work this paper cites.
G. Sumbul, S. Nayak, and B. Demir, “Sd-rsic: Summarization-driven deep remote sensing image captioning,” IEEE Transactions on Geoscience and Remote Sensing , vol. 59, no. 8, pp. 6922–6934, 2021
2021
Earlier work this paper cites.
X. Li, X. Zhang, W. Huang, and Q. Wang, “Truncation cross entropy loss for remote sensing image captioning,” IEEE Transactions on Geoscience and Remote Sensing , vol. 59, no. 6, pp. 5246–5257, 2021
2021
Earlier work this paper cites.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever, “Learning Transferable Visual Models From Natural Language Supervision,” in Proceedings of the 38th International Conference on Machine Learning, ICML 2021, 18-24 July 2021, Virtual Event , vol. 139. PMLR, 2021, pp. 8748–8763
2021
Cited alongside, same era.
S. Pal, A. Arutiunian, G. Venkatesh, R. Ghosh, D. Vidhani, and M. Bhaskar, “Fine tuning CLIP with Remote Sensing (Satellite) images and captions,” HuggingFace Blog , 2021
2021
Cited alongside, same era.
Y. Long, G.-S. Xia, S. Li, W. Yang, M. Y. Yang, X. X. Zhu, L. Zhang, and D. Li, “On creating benchmark dataset for aerial image interpretation: Reviews, guidances and million-aid,” IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing , vol. 14, pp. 4205–4230, 2021
2021
Cited alongside, same era.
2022
Later among the works it cites.
2022
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
C. Jia, Y. Yang, Y. Xia, Y.-T. Chen, Z. Parekh, H. Pham, Q. Le, Y.-H. Sung, Z. Li, and T. Duerig, “Scaling up visual and vision-language representation learning with noisy text supervision,” in International Conference on Machine Learning . PMLR, 2021, pp. 4904–4916
2021
Cited alongside, same era.
M. M. A. Rahhal, Y. Bazi, N. A. Alsharif, L. Bashmal, N. Alajlan, and F. Melgani, “Multilanguage transformer for improved text to remote sensing image retrieval,” IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing , vol. 15, pp. 9115–9126, 2022
2022
Cited alongside, same era.
L. Mi, S. Li, C. Chappuis, and D. Tuia, “Knowledge-aware cross-modal text-image retrieval for remote sensing images,” in Proceedings of the Second Workshop on Complex Data Challenges in Earth Observation (CDCEO 2022) , 2022
2022
Cited alongside, same era.
Z. Yuan, W. Zhang, C. Tian, X. Rong, Z. Zhang, H. Wang, K. Fu, and X. Sun, “Remote sensing cross-modal text-image retrieval based on global and local information,” IEEE Transactions on Geoscience and Remote Sensing , vol. 60, pp. 1–16, 2022
2022
Cited alongside, same era.
Q. Cheng, H. Huang, Y. Xu, Y. Zhou, H. Li, and Z. Wang, “Nwpu-captions dataset and mlca-net for remote sensing image captioning,” IEEE Transactions on Geoscience and Remote Sensing , vol. 60, pp. 1–19, 2022
2022
Cited alongside, same era.
R. Ramos and B. Martins, “Using Neural Encoder-Decoder Models With Continuous Outputs for Remote Sensing Image Captioning,” IEEE Access , vol. 10, 2022
2022
Cited alongside, same era.
J. D. Silva, J. Magalhães, D. Tuia, and B. Martins, “Remote sensing visual question answering with a self-attention multi-modal encoder,” in Proceedings of the 5th ACM SIGSPATIAL International Workshop on AI for Geographic Knowledge Discovery . Association for Computing Machinery, 2022, pp. 40–49
2022
Cited alongside, same era.
Y. Bazi, M. M. Al Rahhal, M. L. Mekhalfi, M. A. Al Zuair, and F. Melgani, “Bi-modal transformer-based approach for visual question answering in remote sensing imagery,” IEEE Transactions on Geoscience and Remote Sensing , vol. 60, pp. 1–11, 2022
2022
Cited alongside, same era.
B. Martins and J. D. Silva, “Towards natural language interfaces for interacting with remote sensing data,” UC Santa Barbara: Center for Spatial Studies , 2022
2022
Cited alongside, same era.
2023
Later among the works it cites.
T. Wei, W. Yuan, J. Luo, W. Zhang, and L. Lu, “VLCA: Vision-language aligning model with cross-modal attention for bilingual remote sensing image captioning,” Journal of Systems Engineering and Electronics , vol. 34, no. 1, pp. 9–18, 2023
2023
Later among the works it cites.
Z. Zhang, L. Jiao, L. Li, X. Liu, P. Chen, F. Liu, Y. Li, and Z. Guo, “A spatial hierarchical reasoning network for remote sensing visual question answering,” IEEE Transactions on Geoscience and Remote Sensing , vol. 61, pp. 1–15, 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
J. Y. Koh, R. Salakhutdinov, and D. Fried, “Grounding language models to images for multimodal inputs and outputs,” in International Conference on Machine Learning . PMLR, 2023, pp. 17 283–17 300
2023
Later among the works it cites.
R. Ramos, B. Martins, D. Elliott, and Y. Kementchedjhieva, “Smallcap: lightweight image captioning prompted with retrieval augmentation,” pp. 2840–2849, 2023
2023
Later among the works it cites.
J. Li, D. Li, S. Savarese, and S. Hoi, “Blip-2: bootstrapping language-image pre-training with frozen image encoders and large language models,” in Proceedings of the 40th International Conference on Machine Learning , ser. ICML’23. JMLR.org, 2023
2023
Later among the works it cites.
V. Zermatten, J. C. Navarro, L. Hughes, T. Kellenberger, and D. Tuia, “Text as a richer source of supervision in semantic segmentation tasks,” in IGARSS 2023 - 2023 IEEE International Geoscience and Remote Sensing Symposium , 2023, pp. 2219–2222
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
A. Lacoste, N. Lehmann, P. Rodriguez, E. D. Sherwin, H. Kerner, B. Lütjens, J. A. Irvin, D. Dao, H. Alemohammad, A. Drouin, M. Gunturkun, G. Huang, D. Vazquez, D. Newman, Y. Bengio, S. Ermon, and X. X. Zhu, “GEO-bench: Toward foundation models for earth monitoring,” in Thirty-seventh Conference on Neural Information Processing Systems Datasets and Benchmarks Track , 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
H. Liu, C. Li, Q. Wu, and Y. J. Lee, “Visual instruction tuning,” arXiv:2304.08485 , 2023
2023
Later among the works it cites.
Y. Fang, W. Wang, B. Xie, Q. Sun, L. Wu, X. Wang, T. Huang, X. Wang, and Y. Cao, “Eva: Exploring the limits of masked visual representation learning at scale,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 19 358–19 369
2023
Later among the works it cites.
W.-L. Chiang, Z. Li, Z. Lin, Y. Sheng, Z. Wu, H. Zhang, L. Zheng, S. Zhuang, Y. Zhuang, J. E. Gonzalez, I. Stoica, and E. P. Xing, “Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality,” March 2023. [Online]. Available: https://lmsys.org/blog/2023-03-30-vicuna/
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.