Fetching the paper…
Reading the bibliography…
Image-text retrieval (ITR) plays a significant role in making informed decisions for various remote sensing (RS) applications.
R. Baeza-Yates, “Modern information retrieval,” Addison Wesley google schola , vol. 2, pp. 127–136, 1999
1999
Earlier work this paper cites.
——, “Cumulated gain-based evaluation of ir techniques,” ACM Transactions on Information Systems (TOIS) , vol. 20, no. 4, pp. 422–446, 2002
2002
Earlier work this paper cites.
C. Milesi, C. D. Elvidge, R. R. Nemani, and S. W. Running, “Assessing the impact of urban land development on net primary productivity in the southeastern united states,” Remote Sensing of Environment , vol. 86, no. 3, pp. 401–410, 2003
2003
Earlier work this paper cites.
S. Rivest, Y. Bédard, M.-J. Proulx, M. Nadeau, F. Hubert, and J. Pastor, “Solap technology: Merging business intelligence with geospatial technology for interactive spatio-temporal exploration and analysis of data,” ISPRS journal of photogrammetry and remote sensing , vol. 60, no. 1, pp. 17–33, 2005
2005
Earlier work this paper cites.
V. Jovanovic, C. Moroney, and D. Nelson, “Multi-angle geometric processing for globally geo-located and co-registered misr image data,” Remote Sensing of Environment , vol. 107, no. 1-2, pp. 22–32, 2007
2007
Earlier work this paper cites.
Y. Yang and S. Newsam, “Bag-of-visual-words and spatial extensions for land-use classification,” in Proceedings of the 18th SIGSPATIAL international conference on advances in geographic information systems , 2010, pp. 270–279
2010
Earlier work this paper cites.
A. Farhadi, M. Hejrati, M. A. Sadeghi, P. Young, C. Rashtchian, J. Hockenmaier, and D. Forsyth, “Every picture tells a story: Generating sentences from images,” in Computer Vision–ECCV 2010: 11th European Conference on Computer Vision, Heraklion, Crete, Greece, September 5-11, 2010, Proceedings, Part IV 11 . Springer, 2010, pp. 15–29
2010
Earlier work this paper cites.
O. J. Reichman, M. B. Jones, and M. P. Schildhauer, “Challenges and opportunities of open data in ecology,” Science , vol. 331, no. 6018, pp. 703–705, 2011
2011
Earlier work this paper cites.
M. N. Kamel Boulos, B. Resch, D. N. Crowley, J. G. Breslin, G. Sohn, R. Burtner, W. A. Pike, E. Jezierski, and K.-Y. S. Chuang, “Crowdsourcing, citizen sensing and sensor web technologies for public and environmental health surveillance and crisis management: trends, ogc standards and application examples,” International journal of health geographics , vol. 10, pp. 1–29, 2011
2011
Earlier work this paper cites.
V. Ordonez, G. Kulkarni, and T. Berg, “Im2text: Describing images using 1 million captioned photographs,” Advances in neural information processing systems , vol. 24, 2011
2011
Earlier work this paper cites.
M. Hodosh, P. Young, and J. Hockenmaier, “Framing image description as a ranking task: Data, models and evaluation metrics,” Journal of Artificial Intelligence Research , vol. 47, pp. 853–899, 2013
2013
Earlier work this paper cites.
A. Frome, G. S. Corrado, J. Shlens, S. Bengio, J. Dean, M. Ranzato, and T. Mikolov, “Devise: A deep visual-semantic embedding model,” Advances in neural information processing systems , vol. 26, 2013
2013
Earlier work this paper cites.
S. Mills, S. Weiss, and C. Liang, “Viirs day/night band (dnb) stray light characterization and correction,” in Earth observing systems XVIII , vol. 8866. SPIE, 2013, pp. 549–566
2013
Earlier work this paper cites.
Q. Yu, Y. Shang, X. Liu, Z. Lei, X. Li, X. Zhu, X. Liu, X. Yang, A. Su, X. Zhang et al. , “Full-parameter vision navigation based on scene matching for aircrafts,” Science China Information Sciences , vol. 57, pp. 1–10, 2014
2014
Earlier work this paper cites.
J. De Leeuw, A. Vrieling, A. Shee, C. Atzberger, K. M. Hadgu, C. M. Biradar, H. Keah, and C. Turvey, “The potential and uptake of remote sensing in insurance: A review,” Remote Sensing , vol. 6, no. 11, pp. 10 888–10 912, 2014
2014
Earlier work this paper cites.
Y. Gong, L. Wang, M. Hodosh, J. Hockenmaier, and S. Lazebnik, “Improving image-sentence embeddings using large weakly annotated photo collections,” in Computer Vision–ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part IV 13 . Springer, 2014, pp. 529–545
2014
Earlier work this paper cites.
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick, “Microsoft coco: Common objects in context,” in Computer Vision–ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part V 13 . Springer, 2014, pp. 740–755
2014
Earlier work this paper cites.
2015
Earlier work this paper cites.
F. Zhao, Y. Huang, L. Wang, and T. Tan, “Deep semantic ranking based hashing for multi-label image retrieval,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2015, pp. 1556–1564
2015
Earlier work this paper cites.
B. Qu, X. Li, D. Tao, and X. Lu, “Deep semantic understanding of high resolution remote sensing image,” in 2016 International conference on computer, information and telecommunication systems (Cits) . IEEE, 2016, pp. 1–5
2016
Earlier work this paper cites.
B. Thomee, D. A. Shamma, G. Friedland, B. Elizalde, K. Ni, D. Poland, D. Borth, and L.-J. Li, “Yfcc100m: The new data in multimedia research,” Communications of the ACM , vol. 59, no. 2, pp. 64–73, 2016
2016
Earlier work this paper cites.
F. Chen, J. Chen, H. Wu, D. Hou, W. Zhang, J. Zhang, X. Zhou, and L. Chen, “A landscape shape index-based sampling approach for land cover accuracy assessment,” Science China Earth Sciences , vol. 59, pp. 2263–2274, 2016
2016
Earlier work this paper cites.
X. Lu, B. Wang, X. Zheng, and X. Li, “Exploring models and data for remote sensing image caption generation,” IEEE Transactions on Geoscience and Remote Sensing , vol. 56, no. 4, pp. 2183–2195, 2017
2017
Earlier work this paper cites.
G. Cheng, J. Han, and X. Lu, “Remote sensing image scene classification: Benchmark and state of the art,” Proceedings of the IEEE , vol. 105, no. 10, pp. 1865–1883, 2017
2017
Earlier work this paper cites.
J. Mun, M. Cho, and B. Han, “Text-guided attention model for image captioning,” in Proceedings of the AAAI conference on artificial intelligence , vol. 31, no. 1, 2017
2017
Cited alongside, same era.
T.-Y. Lin, P. Dollár, R. Girshick, K. He, B. Hariharan, and S. Belongie, “Feature pyramid networks for object detection,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2017, pp. 2117–2125
2017
Cited alongside, same era.
K. Järvelin and J. Kekäläinen, “Ir evaluation methods for retrieving highly relevant documents,” in ACM SIGIR Forum , vol. 51, no. 2. ACM New York, NY, USA, 2017, pp. 243–250
2017
Cited alongside, same era.
G. Panteras and G. Cervone, “Enhancing the temporal resolution of satellite-based flood extent generation using crowdsourced data for disaster monitoring,” International journal of remote sensing , vol. 39, no. 5, pp. 1459–1474, 2018
2018
Cited alongside, same era.
T. T. Lê, J.-L. Froger, and D. H. T. Minh, “Multiscale framework for rapid change analysis from sar image time series: Case study of flood monitoring in the central coast regions of vietnam,” Remote Sensing of Environment , vol. 269, p. 112837, 2022
2022
Later among the works it cites.
Q. Cheng, H. Huang, Y. Xu, Y. Zhou, H. Li, and Z. Wang, “Nwpu-captions dataset and mlca-net for remote sensing image captioning,” IEEE Transactions on Geoscience and Remote Sensing , vol. 60, pp. 1–19, 2022
2022
Later among the works it cites.
2022
Later among the works it cites.
C. Schuhmann, R. Beaumont, R. Vencu, C. Gordon, R. Wightman, M. Cherti, T. Coombes, A. Katta, C. Mullis, M. Wortsman et al. , “Laion-5b: An open large-scale dataset for training next generation image-text models,” Advances in Neural Information Processing Systems , vol. 35, pp. 25 278–25 294, 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
P. Sharma, N. Ding, S. Goodman, and R. Soricut, “Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning,” in Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , 2018, pp. 2556–2565
2018
Cited alongside, same era.
G. Christie, N. Fendley, J. Wilson, and R. Mukherjee, “Functional map of the world,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2018, pp. 6172–6180
2018
Cited alongside, same era.
L. Wang, Y. Li, J. Huang, and S. Lazebnik, “Learning two-branch neural networks for image-text matching tasks,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 41, no. 2, pp. 394–407, 2018
2018
Cited alongside, same era.
K.-H. Lee, X. Chen, G. Hua, H. Hu, and X. He, “Stacked cross attention for image-text matching,” in Proceedings of the European conference on computer vision (ECCV) , 2018, pp. 201–216
2018
Cited alongside, same era.
G. Sumbul, M. Charfuelan, B. Demir, and V. Markl, “Bigearthnet: A large-scale benchmark archive for remote sensing image understanding,” in IGARSS 2019-2019 IEEE International Geoscience and Remote Sensing Symposium . IEEE, 2019, pp. 5901–5904
2019
Cited alongside, same era.
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever et al. , “Language models are unsupervised multitask learners,” OpenAI blog , vol. 1, no. 8, p. 9, 2019
2019
Cited alongside, same era.
P. Ge, H. Gokon, and K. Meguro, “A review on synthetic aperture radar-based building damage assessment in disasters,” Remote Sensing of Environment , vol. 240, p. 111693, 2020
2020
Cited alongside, same era.
M. Yang, J. Liu, Y. Shen, Z. Zhao, X. Chen, Q. Wu, and C. Li, “An ensemble of generation-and retrieval-based image captioning with dual generator generative adversarial network,” IEEE Transactions on Image Processing , vol. 29, pp. 9627–9640, 2020
2020
Cited alongside, same era.
2022
Later among the works it cites.
M. Byeon, B. Park, H. Kim, S. Lee, W. Baek, and S. Kim, “Coyo-700m: Image-text pair dataset,” https://github.com/kakaobrain/coyo-dataset , 2022
2022
Later among the works it cites.
J.-B. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, Y. Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynolds et al. , “Flamingo: a visual language model for few-shot learning,” Advances in Neural Information Processing Systems , vol. 35, pp. 23 716–23 736, 2022
2022
Later among the works it cites.
2022
Later among the works it cites.
J. Li, D. Li, C. Xiong, and S. Hoi, “Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation,” in International Conference on Machine Learning . PMLR, 2022, pp. 12 888–12 900
2022
Later among the works it cites.
Z. Yuan, W. Zhang, C. Tian, X. Rong, Z. Zhang, H. Wang, K. Fu, and X. Sun, “Remote sensing cross-modal text-image retrieval based on global and local information,” IEEE Transactions on Geoscience and Remote Sensing , vol. 60, pp. 1–16, 2022
2022
Later among the works it cites.
Z. Wang, Y. Lu, Q. Li, X. Tao, Y. Guo, M. Gong, and T. Liu, “Cris: Clip-driven referring image segmentation,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2022, pp. 11 686–11 695
2022
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
W. Dai, J. Li, D. Li, A. M. H. Tiong, J. Zhao, W. Wang, B. Li, P. Fung, and S. Hoi, “Instructblip: Towards general-purpose vision-language models with instruction tuning,” 2023
2023
Later among the works it cites.
H. Liu, C. Li, Q. Wu, and Y. J. Lee, “Visual instruction tuning,” 2023
2023
Later among the works it cites.
Z. Cao, L. Jiang, P. Yue, J. Gong, X. Hu, S. Liu, H. Tan, C. Liu, B. Shangguan, and D. Yu, “A large scale training sample database system for intelligent interpretation of remote sensing imagery,” Geo-Spatial Information Science , pp. 1–20, 2023
2023
Later among the works it cites.
P. Gao, S. Geng, R. Zhang, T. Ma, R. Fang, Y. Zhang, H. Li, and Y. Qiao, “Clip-adapter: Better vision-language models with feature adapters,” International Journal of Computer Vision , pp. 1–15, 2023
2023
Later among the works it cites.
Q. Zheng, K. C. Seto, Y. Zhou, S. You, and Q. Weng, “Nighttime light remote sensing for urban applications: Progress, challenges, and prospects,” ISPRS Journal of Photogrammetry and Remote Sensing , vol. 202, pp. 125–141, 2023
2023
Later among the works it cites.
2024
Closest in time.
K. Xu, J. Ba, R. Kiros, K. Cho, A. Courville, R. Salakhudinov, R. Zemel, and Y. Bengio, “Show, attend and tell: Neural image caption generation with visual attention,” in International conference on machine learning . PMLR, 2015, pp. 2048–2057
2057
Closest in time.
Y. Wu, S. Wang, G. Song, and Q. Huang, “Learning fragment self-attention embeddings for image-text matching,” in Proceedings of the 27th ACM international conference on multimedia , 2019, pp. 2088–2096
2096
Closest in time.