Fetching the paper…
Reading the bibliography…
Ultra High Resolution (UHR) remote sensing imagery (RSI) (e.g.
1908
Earlier work this paper cites.
M. Ester, H.-P. Kriegel, J. Sander, and X. Xu, “A density-based algorithm for discovering clusters in large spatial databases with noise,” in Proceedings of the Second International Conference on Knowledge Discovery and Data Mining , ser. KDD’96. AAAI Press, 1996, p. 226–231
1996
Earlier work this paper cites.
2006
Earlier work this paper cites.
2010
Earlier work this paper cites.
V. Ordonez, G. Kulkarni, and T. L. Berg, “Im2text: Describing images using 1 million captioned photographs,” in Neural Information Processing Systems (NIPS) , 2011
2011
Earlier work this paper cites.
B. Thomee, D. A. Shamma, G. Friedland, B. Elizalde, K. Ni, D. Poland, D. Borth, and L.-J. Li, “YFCC100m,” Communications of the ACM , vol. 59, no. 2, pp. 64–73, jan 2016. [Online]. Available: https://doi.org/10.1145%2F2812802
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
S. Ren, K. He, R. Girshick, and J. Sun, “Faster r-cnn: Towards real-time object detection with region proposal networks,” IEEE transactions on pattern analysis and machine intelligence , vol. 39, no. 6, pp. 1137–1149, 2016
2016
Earlier work this paper cites.
2017
Earlier work this paper cites.
G.-S. Xia, X. Bai, J. Ding, Z. Zhu, S. Belongie, J. Luo, M. Datcu, M. Pelillo, and L. Zhang, “Dota: A large-scale dataset for object detection in aerial images,” in The IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , June 2018
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
P. Sharma, N. Ding, S. Goodman, and R. Soricut, “Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning,” in Proceedings of ACL , 2018
2018
Earlier work this paper cites.
S. Waqas Zamir, A. Arora, A. Gupta, S. Khan, G. Sun, F. Shahbaz Khan, F. Zhu, L. Shao, G.-S. Xia, and X. Bai, “isaid: A large-scale dataset for instance segmentation in aerial images,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops , 2019, pp. 28–37
2019
Earlier work this paper cites.
D. Hou, Z. Miao, H. Xing, and H. Wu, “V-rsir: An open access web-based image annotation tool for remote sensing image retrieval,” IEEE Access , vol. 7, pp. 83 852–83 862, 2019
2019
Earlier work this paper cites.
J. Deng, J. Guo, N. Xue, and S. Zafeiriou, “Arcface: Additive angular margin loss for deep face recognition,” in 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2019, pp. 4685–4694
2019
Earlier work this paper cites.
M. Grootendorst, “Keybert: Minimal keyword extraction with bert.” 2020. [Online]. Available: https://doi.org/10.5281/zenodo.4461265
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
Y.-C. Chen, L. Li, L. Yu, A. El Kholy, F. Ahmed, Z. Gan, Y. Cheng, and J. Liu, “Uniter: Universal image-text representation learning,” in Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XXX . Springer, 2020, pp. 104–120
2020
Earlier work this paper cites.
P. Lewis, E. Perez, A. Piktus, F. Petroni, V. Karpukhin, N. Goyal, H. Küttler, M. Lewis, W.-t. Yih, T. Rocktäschel et al. , “Retrieval-augmented generation for knowledge-intensive nlp tasks,” Advances in neural information processing systems , vol. 33, pp. 9459–9474, 2020
2020
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever, “Learning transferable visual models from natural language supervision,” 2021
2021
Earlier work this paper cites.
J. Ding, N. Xue, G.-S. Xia, X. Bai, W. Yang, M. Y. Yang, S. Belongie, J. Luo, M. Datcu, M. Pelillo, and L. Zhang, “Object detection in aerial images: A large-scale benchmark and challenges,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 44, no. 11, p. 7778–7796, Nov. 2022. [Online]. Available: http://dx.doi.org/10.1109/TPAMI.2021.3117983
2021
Earlier work this paper cites.
J. Wang, Z. Zheng, A. Ma, X. Lu, and Y. Zhong, “Loveda: A remote sensing land-cover dataset for domain adaptive semantic segmentation,” in Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks , J. Vanschoren and S. Yeung, Eds., vol. 1. Curran Associates, Inc., 2021. [Online]. Available: https://datasets-benchmarks-proceedings.neurips.cc/paper_files/paper/2021/file/4e732ced3463d06de0ca9a15b6153677-Paper-round2.pdf
2021
Earlier work this paper cites.
C. Schuhmann, R. Vencu, R. Beaumont, R. Kaczmarczyk, C. Mullis, A. Katta, T. Coombes, J. Jitsev, and A. Komatsuzaki, “Laion-400m: Open dataset of clip-filtered 400 million image-text pairs,” 2021
2021
Earlier work this paper cites.
S. Changpinyo, P. Sharma, N. Ding, and R. Soricut, “Conceptual 12M: Pushing web-scale image-text pre-training to recognize long-tail visual concepts,” in CVPR , 2021
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
K. Desai, G. Kaul, Z. Aysola, and J. Johnson, “RedCaps: Web-curated image-text data created by the people, for the people,” in NeurIPS Datasets and Benchmarks , 2021
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
P. Zhang, X. Li, X. Hu, J. Yang, L. Zhang, L. Wang, Y. Choi, and J. Gao, “Vinvl: Revisiting visual representations in vision-language models,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2021, pp. 5579–5588
2021
Earlier work this paper cites.
C. Schuhmann, R. Beaumont, R. Vencu, C. Gordon, R. Wightman, M. Cherti, T. Coombes, A. Katta, C. Mullis, M. Wortsman, P. Schramowski, S. Kundurthy, K. Crowson, L. Schmidt, R. Kaczmarczyk, and J. Jitsev, “Laion-5b: An open large-scale dataset for training next generation image-text models,” 2022
2022
Earlier work this paper cites.
M. Byeon, B. Park, H. Kim, S. Lee, W. Baek, and S. Kim, “Coyo-700m: Image-text pair dataset,” https://github.com/kakaobrain/coyo-dataset , 2022
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
J.-B. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, Y. Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynolds, R. Ring, E. Rutherford, S. Cabi, T. Han, Z. Gong, S. Samangooei, M. Monteiro, J. Menick, S. Borgeaud, A. Brock, A. Nematzadeh, S. Sharifzadeh, M. Binkowski, R. Barreira, O. Vinyals, A. Zisserman, and K. Simonyan, “Flamingo: a visual language model for few-shot learning,” 2022
2022
Earlier work this paper cites.
X. Zhou, R. Girdhar, A. Joulin, P. Krähenbühl, and I. Misra, “Detecting twenty-thousand classes using image-level supervision,” in Computer Vision–ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23–27, 2022, Proceedings, Part IX . Springer, 2022, pp. 350–368
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
H. Liu, C. Li, Y. Li, and Y. J. Lee, “Improved baselines with visual instruction tuning,” 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Cited alongside, same era.
G. Cheng, X. Yuan, X. Yao, K. Yan, Q. Zeng, X. Xie, and J. Han, “Towards large-scale small object detection: Survey and benchmarks,” IEEE Transactions on Pattern Analysis and Machine Intelligence , p. 1–20, 2023. [Online]. Available: http://dx.doi.org/10.1109/TPAMI.2023.3290594
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Z. Wang, R. Prabha, T. Huang, J. Wu, and R. Rajagopal, “Skyscript: A large and semantically diverse vision-language dataset for remote sensing,” Proceedings of the AAAI Conference on Artificial Intelligence , vol. 38, no. 6, pp. 5805–5813, Mar. 2024. [Online]. Available: https://ojs.aaai.org/index.php/AAAI/article/view/28393
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
P. Wu and S. Xie, “V?: Guided visual search as a core mechanism in multimodal llms,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 13 084–13 094
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
X. Zhao, D. Li, Y. Zhong, B. Hu, Y. Chen, B. Hu, and M. Zhang, “SEER: Self-aligned evidence extraction for retrieval-augmented generation,” in Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing , Y. Al-Onaizan, M. Bansal, and Y.-N. Chen, Eds. Miami, Florida, USA: Association for Computational Linguistics, Nov. 2024, pp. 3027–3041. [Online]. Available: https://aclanthology.org/2024.emnlp-main.178/
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
M. A. Arefeen, B. Debnath, M. Y. S. Uddin, and S. Chakradhar, “irag: Advancing rag for videos with an incremental approach,” in Proceedings of the 33rd ACM International Conference on Information and Knowledge Management , 2024, pp. 4341–4348
2024
Closest in time.
2024
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
K. Singhal, T. Tu, J. Gottweis, R. Sayres, E. Wulczyn, M. Amin, L. Hou, K. Clark, S. R. Pfohl, H. Cole-Lewis, D. Neal, Q. M. Rashid, M. Schaekermann, A. Wang, D. Dash, J. H. Chen, N. H. Shah, S. Lachgar, P. A. Mansfield, S. Prakash, B. Green, E. Dominowska, B. Agüera Y Arcas, N. Tomašev, Y. Liu, R. Wong, C. Semturs, S. S. Mahdavi, J. K. Barral, D. R. Webster, G. S. Corrado, Y. Matias, S. Azizi, A. Karthikesalingam, and V. Natarajan, “Toward expert-level medical question answering with large language models,” Nat. Med. , vol. 31, no. 3, pp. 943–950, Mar. 2025
2025
Closest in time.
2025
Closest in time.
Y. Zhan, Z. Xiong, and Y. Yuan, “Skyeyegpt: Unifying remote sensing vision-language tasks via instruction tuning with large language model,” ISPRS Journal of Photogrammetry and Remote Sensing , vol. 221, pp. 64–77, 2025. [Online]. Available: https://www.sciencedirect.com/science/article/pii/S0924271625000206
2025
Closest in time.
D. Muhtar, Z. Li, F. Gu, X. Zhang, and P. Xiao, “Lhrs-bot: Empowering remote sensing with vgi-enhanced large multimodal language model,” in Computer Vision – ECCV 2024 , A. Leonardis, E. Ricci, S. Roth, O. Russakovsky, T. Sattler, and G. Varol, Eds. Cham: Springer Nature Switzerland, 2025, pp. 440–457
2025
Closest in time.
Z. Zheng, S. Ermon, D. Kim, L. Zhang, and Y. Zhong, “Changen2: Multi-temporal remote sensing generative change foundation model,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 47, no. 2, pp. 725–741, 2025
2025
Closest in time.
X. Zhao, Y. Zhong, Z. Sun, X. Hu, Z. Liu, D. Li, B. Hu, and M. Zhang, “FunnelRAG: A coarse-to-fine progressive retrieval paradigm for RAG,” in Findings of the Association for Computational Linguistics: NAACL 2025 , L. Chiruzzo, A. Ritter, and L. Wang, Eds. Albuquerque, New Mexico: Association for Computational Linguistics, Apr. 2025, pp. 3029–3046. [Online]. Available: https://aclanthology.org/2025.findings-naacl.165/
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.