Fetching the paper…
Reading the bibliography…
The rapid advancement of remote sensing foundation models, particularly vision and multimodal models, has significantly enhanced the capabilities of intelligent geospatial data interpretation.
D. H. Hubel and T. N. Wiesel, “Receptive fields, binocular interaction and functional architecture in the cat’s visual cortex,” The Journal of physiology , vol. 160, no. 1, p. 106, 1962
1962
Earlier work this paper cites.
D. Lu and Q. Weng, “A survey of image classification methods and techniques for improving classification performance,” International journal of Remote sensing , vol. 28, no. 5, pp. 823–870, 2007
2007
Earlier work this paper cites.
Y. Yang and S. Newsam, “Bag-of-visual-words and spatial extensions for land-use classification,” in Proceedings of the 18th SIGSPATIAL international conference on advances in geographic information systems , 2010, pp. 270–279
2010
Earlier work this paper cites.
D. Dai and W. Yang, “Satellite image classification via two-layer sparse coding with biased image representation,” IEEE Geoscience and remote sensing letters , vol. 8, no. 1, pp. 173–176, 2010
2010
Earlier work this paper cites.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,” Advances in neural information processing systems , vol. 25, 2012
2012
Earlier work this paper cites.
C. Szegedy, A. Toshev, and D. Erhan, “Deep neural networks for object detection,” Advances in neural information processing systems , vol. 26, 2013
2013
Earlier work this paper cites.
F. Zhang, B. Du, and L. Zhang, “Saliency-guided unsupervised feature learning for scene classification,” IEEE transactions on Geoscience and Remote Sensing , vol. 53, no. 4, pp. 2175–2184, 2014
2014
Earlier work this paper cites.
G. Cheng, J. Han, P. Zhou, and L. Guo, “Multi-class geospatial object detection and geographic image classification based on collection of part detectors,” ISPRS Journal of Photogrammetry and Remote Sensing , vol. 98, pp. 119–132, 2014
2014
Earlier work this paper cites.
B. Zhao, Y. Zhong, G.-S. Xia, and L. Zhang, “Dirichlet-derived multiple topic scene classification model for high spatial resolution remote sensing imagery,” IEEE Transactions on Geoscience and Remote Sensing , vol. 54, no. 4, pp. 2108–2123, 2015
2015
Earlier work this paper cites.
S. Song, S. P. Lichtenberg, and J. Xiao, “Sun rgb-d: A rgb-d scene understanding benchmark suite,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2015, pp. 567–576
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
A. Kusk, A. Abulaitijiang, and J. Dall, “Synthetic sar image generation using sensor, terrain and target models,” in Proceedings of EUSAR 2016: 11th European Conference on Synthetic Aperture Radar . VDE, 2016, pp. 1–5
2016
Earlier work this paper cites.
B. Qu, X. Li, D. Tao, and X. Lu, “Deep semantic understanding of high resolution remote sensing image,” in 2016 International conference on computer, information and telecommunication systems . IEEE, 2016, pp. 1–5
2016
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 770–778
2016
Earlier work this paper cites.
E. Sharifi, R. Steinacker, and B. Saghafian, “Assessment of gpm-imerg and other precipitation products against gauge data under different topographic and climatic conditions in iran: Preliminary results,” Remote Sensing , vol. 8, no. 2, p. 135, 2016
2016
Earlier work this paper cites.
G. Cheng, J. Han, and X. Lu, “Remote sensing image scene classification: Benchmark and state of the art,” Proceedings of the IEEE , vol. 105, no. 10, pp. 1865–1883, 2017
2017
Earlier work this paper cites.
G.-S. Xia, J. Hu, F. Hu, B. Shi, X. Bai, Y. Zhong, L. Zhang, and X. Lu, “Aid: A benchmark data set for performance evaluation of aerial scene classification,” IEEE Transactions on Geoscience and Remote Sensing , vol. 55, no. 7, pp. 3965–3981, 2017
2017
Earlier work this paper cites.
X. Lu, B. Wang, X. Zheng, and X. Li, “Exploring models and data for remote sensing image caption generation,” IEEE Transactions on Geoscience and Remote Sensing , vol. 56, no. 4, pp. 2183–2195, 2017
2017
Earlier work this paper cites.
W. Rawat and Z. Wang, “Deep convolutional neural networks for image classification: A comprehensive review,” Neural computation , vol. 29, no. 9, pp. 2352–2449, 2017
2017
Earlier work this paper cites.
A. Vaswani, “Attention is all you need,” Advances in Neural Information Processing Systems , 2017
2017
Earlier work this paper cites.
M. Larson, M. Soleymani, G. Gravier, B. Ionescu, and G. J. Jones, “The benchmarking initiative for multimedia evaluation: Mediaeval 2016,” IEEE MultiMedia , vol. 24, no. 1, pp. 93–96, 2017
2017
Earlier work this paper cites.
G.-S. Xia, X. Bai, J. Ding, Z. Zhu, S. Belongie, J. Luo, M. Datcu, M. Pelillo, and L. Zhang, “Dota: A large-scale dataset for object detection in aerial images,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2018, pp. 3974–3983
2018
Earlier work this paper cites.
R. C. Daudt, B. Le Saux, A. Boulch, and Y. Gousseau, “Urban change detection for multispectral earth observation using convolutional neural networks,” in IGARSS 2018-2018 IEEE International Geoscience and Remote Sensing Symposium . Ieee, 2018, pp. 2115–2118
2018
Earlier work this paper cites.
A. Radford, “Improving language understanding by generative pre-training,” 2018
2018
Earlier work this paper cites.
G. Christie, N. Fendley, J. Wilson, and R. Mukherjee, “Functional map of the world,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2018, pp. 6172–6180
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
Y. Guo, Y. Liu, T. Georgiou, and M. S. Lew, “A review of semantic segmentation using deep neural networks,” International journal of multimedia information retrieval , vol. 7, pp. 87–93, 2018
2018
Earlier work this paper cites.
J. Gu, Z. Wang, J. Kuen, L. Ma, A. Shahroudy, B. Shuai, T. Liu, X. Wang, G. Wang, J. Cai et al. , “Recent advances in convolutional neural networks,” Pattern recognition , vol. 77, pp. 354–377, 2018
2018
Earlier work this paper cites.
S. Ji, S. Wei, and M. Lu, “Fully convolutional networks for multisource building extraction from an open aerial and satellite imagery data set,” IEEE Transactions on geoscience and remote sensing , vol. 57, no. 1, pp. 574–586, 2018
2018
Earlier work this paper cites.
W. Zhou, S. Newsam, C. Li, and Z. Shao, “Patternnet: A benchmark dataset for performance evaluation of remote sensing image retrieval,” ISPRS journal of photogrammetry and remote sensing , vol. 145, pp. 197–209, 2018
2018
Earlier work this paper cites.
P. Helber, B. Bischke, A. Dengel, and D. Borth, “Eurosat: A novel dataset and deep learning benchmark for land use and land cover classification,” IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing , vol. 12, no. 7, pp. 2217–2226, 2019
2019
Earlier work this paper cites.
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever et al. , “Language models are unsupervised multitask learners,” OpenAI blog , vol. 1, no. 8, p. 9, 2019
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
G. Sumbul, M. Charfuelan, B. Demir, and V. Markl, “Bigearthnet: A large-scale benchmark archive for remote sensing image understanding,” in IGARSS 2019-2019 IEEE International Geoscience and Remote Sensing Symposium . IEEE, 2019, pp. 5901–5904
2019
Earlier work this paper cites.
K. Wang, G. Zhang, and H. Leung, “Sar target recognition based on cross-domain and cross-task transfer learning,” IEEE Access , vol. 7, pp. 153 391–153 399, 2019
2019
Earlier work this paper cites.
B. Lewis, T. Scarnati, E. Sudkamp, J. Nehrbass, S. Rosencrantz, and E. Zelnio, “A sar dataset for atr development: the synthetic and measured paired labeled experiment (sample),” in Algorithms for Synthetic Aperture Radar Imagery XXVI , vol. 10987. SPIE, 2019, pp. 39–54
2019
Earlier work this paper cites.
S. Waqas Zamir, A. Arora, A. Gupta, S. Khan, G. Sun, F. Shahbaz Khan, F. Zhu, L. Shao, G.-S. Xia, and X. Bai, “isaid: A large-scale dataset for instance segmentation in aerial images,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops , 2019, pp. 28–37
2019
Earlier work this paper cites.
Z.-Q. Zhao, P. Zheng, S.-t. Xu, and X. Wu, “Object detection with deep learning: A review,” IEEE transactions on neural networks and learning systems , vol. 30, no. 11, pp. 3212–3232, 2019
2019
Earlier work this paper cites.
C. Choy, J. Gwak, and S. Savarese, “4d spatio-temporal convnets: Minkowski convolutional neural networks,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2019, pp. 3075–3084
2019
Earlier work this paper cites.
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “Bert: Pre-training of deep bidirectional transformers for language understanding,” in Proceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human language technologies, volume 1 (long and short papers) , 2019, pp. 4171–4186
2019
Earlier work this paper cites.
K. Li, G. Wan, G. Cheng, L. Meng, and J. Han, “Object detection in optical remote sensing images: A survey and a new benchmark,” ISPRS journal of photogrammetry and remote sensing , vol. 159, pp. 296–307, 2020
2020
Earlier work this paper cites.
H. Chen and Z. Shi, “A spatial-temporal attention-based method and a new dataset for remote sensing image change detection,” Remote Sensing , vol. 12, no. 10, p. 1662, 2020
2020
Earlier work this paper cites.
S. Wang, W. Chen, S. M. Xie, G. Azzari, and D. B. Lobell, “Weakly supervised deep learning for segmentation of remote sensing imagery,” Remote Sensing , vol. 12, no. 2, p. 207, 2020
2020
Earlier work this paper cites.
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell et al. , “Language models are few-shot learners,” Advances in neural information processing systems , vol. 33, pp. 1877–1901, 2020
2020
Earlier work this paper cites.
S. Lobry, D. Marcos, J. Murray, and D. Tuia, “Rsvqa: Visual question answering for remote sensing data,” IEEE Transactions on Geoscience and Remote Sensing , vol. 58, no. 12, pp. 8555–8566, 2020
2020
Earlier work this paper cites.
S. Hao, Y. Zhou, and Y. Guo, “A brief survey on semantic segmentation with deep learning,” Neurocomputing , vol. 406, pp. 302–321, 2020
2020
Earlier work this paper cites.
V. S. F. Garnot and L. Landrieu, “Lightweight temporal self-attention for classifying satellite images time series,” in Advanced Analytics and Learning on Temporal Data: 5th ECML PKDD Workshop, AALTD 2020, Ghent, Belgium, September 18, 2020, Revised Selected Papers 6 . Springer, 2020, pp. 171–181
2020
Earlier work this paper cites.
K. He, H. Fan, Y. Wu, S. Xie, and R. Girshick, “Momentum contrast for unsupervised visual representation learning,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2020, pp. 9729–9738
2020
Earlier work this paper cites.
T. Chen, S. Kornblith, M. Norouzi, and G. Hinton, “A simple framework for contrastive learning of visual representations,” in International conference on machine learning . PMLR, 2020, pp. 1597–1607
2020
Earlier work this paper cites.
J.-B. Grill, F. Strub, F. Altché, C. Tallec, P. Richemond, E. Buchatskaya, C. Doersch, B. Avila Pires, Z. Guo, M. Gheshlaghi Azar et al. , “Bootstrap your own latent-a new approach to self-supervised learning,” Advances in neural information processing systems , vol. 33, pp. 21 271–21 284, 2020
2020
Earlier work this paper cites.
G. Baier, A. Deschemps, M. Schmitt, and N. Yokoya, “Geonrw (2020).”
2020
Earlier work this paper cites.
T. Zhang, X. Zhang, J. Li, X. Xu, B. Wang, X. Zhan, Y. Xu, X. Ke, T. Zeng, H. Su et al. , “Sar ship detection dataset (ssdd): Official release and comprehensive data analysis,” Remote Sensing , vol. 13, no. 18, p. 3690, 2021
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
X. Yang, J. Yan, Q. Ming, W. Wang, X. Zhang, and Q. Tian, “Rethinking rotated object detection with gaussian wasserstein distance loss,” in International conference on machine learning . PMLR, 2021, pp. 11 830–11 841
2021
Earlier work this paper cites.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark et al. , “Learning transferable visual models from natural language supervision,” in International conference on machine learning . PMLR, 2021, pp. 8748–8763
2021
Earlier work this paper cites.
V. Stojnic and V. Risojevic, “Self-supervised learning of remote sensing scene representations using contrastive multiview coding,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops , June 2021, pp. 1182–1191
2021
Earlier work this paper cites.
K. Ayush, B. Uzkent, C. Meng, K. Tanmay, M. Burke, D. Lobell, and S. Ermon, “Geography-aware self-supervised learning,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , October 2021, pp. 10 181–10 190
2021
Earlier work this paper cites.
O. Mañas, A. Lacoste, X. Giró-i Nieto, D. Vazquez, and P. Rodríguez, “Seasonal contrast: Unsupervised pre-training from uncurated remote sensing data,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , October 2021, pp. 9414–9423
2021
Earlier work this paper cites.
Y. Long, G.-S. Xia, S. Li, W. Yang, M. Y. Yang, X. X. Zhu, L. Zhang, and D. Li, “On creating benchmark dataset for aerial image interpretation: Reviews, guidances, and million-aid,” IEEE Journal of selected topics in applied earth observations and remote sensing , vol. 14, pp. 4205–4230, 2021
2021
Earlier work this paper cites.
G. Sumbul, A. De Wall, T. Kreuziger, F. Marcelino, H. Costa, P. Benevides, M. Caetano, B. Demir, and V. Markl, “Bigearthnet-mm: A large-scale, multimodal, multilabel benchmark archive for remote sensing image classification and retrieval [software and data sets],” IEEE Geoscience and Remote Sensing Magazine , vol. 9, no. 3, pp. 174–180, 2021
2021
Earlier work this paper cites.
X. Zheng, B. Wang, X. Du, and X. Lu, “Mutual attention inception network for remote sensing visual question answering,” IEEE Transactions on Geoscience and Remote Sensing , vol. 60, pp. 1–14, 2021
2021
Earlier work this paper cites.
M. Rahnemoonfar, T. Chowdhury, A. Sarkar, D. Varshney, M. Yari, and R. R. Murphy, “Floodnet: A high resolution aerial imagery dataset for post flood scene understanding,” IEEE Access , vol. 9, pp. 89 644–89 654, 2021
2021
Earlier work this paper cites.
T. Kattenborn, J. Leitloff, F. Schiefer, and S. Hinz, “Review on convolutional neural networks (cnn) in vegetation remote sensing,” ISPRS journal of photogrammetry and remote sensing , vol. 173, pp. 24–49, 2021
2021
Earlier work this paper cites.
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, “An image is worth 16x16 words: Transformers for image recognition at scale,” in International Conference on Learning Representations , 2021
2021
Earlier work this paper cites.
W. Wang, E. Xie, X. Li, D.-P. Fan, K. Song, D. Liang, T. Lu, P. Luo, and L. Shao, “Pyramid vision transformer: A versatile backbone for dense prediction without convolutions,” in Proceedings of the IEEE/CVF international conference on computer vision , 2021, pp. 568–578
2021
Earlier work this paper cites.
Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo, “Swin transformer: Hierarchical vision transformer using shifted windows,” in Proceedings of the IEEE/CVF international conference on computer vision , 2021, pp. 10 012–10 022
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
G. Hoxha and F. Melgani, “A novel svm-based decoder for remote sensing image captioning,” IEEE Transactions on Geoscience and Remote Sensing , vol. 60, pp. 1–14, 2021
2021
Earlier work this paper cites.
Z. Jiang, T. Chen, B. J. Mortazavi, and Z. Wang, “Self-damaging contrastive learning,” in International Conference on Machine Learning . PMLR, 2021, pp. 4927–4939
2021
Earlier work this paper cites.
L. P. Osco, J. M. Junior, A. P. M. Ramos, L. A. de Castro Jorge, S. N. Fatholahi, J. de Andrade Silva, E. T. Matsubara, H. Pistori, W. N. Gonçalves, and J. Li, “A review on deep learning in uav remote sensing,” International Journal of Applied Earth Observation and Geoinformation , vol. 102, p. 102456, 2021
2021
Earlier work this paper cites.
G. Cheng, J. Wang, K. Li, X. Xie, C. Lang, Y. Yao, and J. Han, “Anchor-free oriented proposal generator for object detection,” IEEE Transactions on Geoscience and Remote Sensing , vol. 60, pp. 1–11, 2022
2022
Earlier work this paper cites.
L. Scheibenreif, J. Hanna, M. Mommert, and D. Borth, “Self-supervised vision transformers for land-cover segmentation and classification,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops , June 2022, pp. 1422–1431
2022
Earlier work this paper cites.
P. Akiva, M. Purri, and M. Leotta, “Self-supervised material and texture representation learning for remote sensing tasks,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , June 2022, pp. 8203–8215
2022
Earlier work this paper cites.
Y. Wang, C. M. Albrecht, and X. X. Zhu, “Self-supervised vision transformers for joint sar-optical representation learning,” in IGARSS 2022-2022 IEEE International Geoscience and Remote Sensing Symposium . IEEE, 2022, pp. 139–142
2022
Earlier work this paper cites.
P. Jain, B. Schoen-Phelan, and R. Ross, “Self-supervised learning for invariant representations from multi-spectral and sar images,” IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing , vol. 15, pp. 7797–7808, 2022
2022
Earlier work this paper cites.
Y. Cong, S. Khanna, C. Meng, P. Liu, E. Rozi, Y. He, M. Burke, D. B. Lobell, and S. Ermon, “SatMAE: Pre-training transformers for temporal and multi-spectral satellite imagery,” in Advances in Neural Information Processing Systems , A. H. Oh, A. Agarwal, D. Belgrave, and K. Cho, Eds., 2022
2022
Earlier work this paper cites.
W. Li, K. Chen, H. Chen, and Z. Shi, “Geographical knowledge-driven representation learning for remote sensing images,” IEEE Transactions on Geoscience and Remote Sensing , vol. 60, pp. 1–16, 2022
2022
Earlier work this paper cites.
W. Li, K. Chen, and Z. Shi, “Geographical supervision correction for remote sensing representation learning,” IEEE Transactions on Geoscience and Remote Sensing , vol. 60, pp. 1–20, 2022
2022
Earlier work this paper cites.
T. Zhang, P. Gao, H. Dong, Y. Zhuang, G. Wang, W. Zhang, and H. Chen, “Consecutive pre-training: A knowledge transfer learning strategy with relevant unlabeled data for remote sensing domain,” Remote Sensing , vol. 14, no. 22, p. 5675, 2022
2022
Cited alongside, same era.
J. Chen, Z. Huang, R. Xia, B. Wu, L. Sheng, L. Sun, and B. Yao, “Large-scale multi-class sar image target detection dataset-1.0,” Journal of Radars , no. 1, 2022
2022
Cited alongside, same era.
Z. Yuan, W. Zhang, K. Fu, X. Li, C. Deng, H. Wang, and X. Sun, “Exploring a fine-grained multiscale method for cross-modal remote sensing image retrieval,” IEEE Transactions on Geoscience and Remote Sensing , vol. 60, pp. 1–19, 2022
2022
Cited alongside, same era.
Q. Cheng, H. Huang, Y. Xu, Y. Zhou, H. Li, and Z. Wang, “Nwpu-captions dataset and mlca-net for remote sensing image captioning,” IEEE Transactions on Geoscience and Remote Sensing , vol. 60, pp. 1–19, 2022
2022
Cited alongside, same era.
2023
Later among the works it cites.
J. Hu, R. Liu, D. Hong, A. Camero, J. Yao, M. Schneider, F. Kurz, K. Segl, and X. X. Zhu, “Mdas: A new multimodal benchmark dataset for remote sensing,” Earth System Science Data , vol. 15, no. 1, pp. 113–131, 2023
2023
Later among the works it cites.
X. Zhang, Y. Li, X. Wang, F. Liu, Z. Wu, X. Cheng, and L. Jiao, “Multi-source interactive stair attention for remote sensing image captioning,” Remote Sensing , vol. 15, no. 3, p. 579, 2023
2023
Later among the works it cites.
Y. Yuan, Y. Zhan, and Z. Xiong, “Parameter-efficient transfer learning for remote sensing image-text retrieval,” IEEE Transactions on Geoscience and Remote Sensing , 2023
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
X. Sun, P. Wang, Z. Yan, F. Xu, R. Wang, W. Diao, J. Chen, J. Li, Y. Feng, T. Xu et al. , “Fair1m: A benchmark dataset for fine-grained object recognition in high-resolution remote sensing imagery,” ISPRS Journal of Photogrammetry and Remote Sensing , vol. 184, pp. 116–130, 2022
2022
Cited alongside, same era.
Y. Sun, S. Feng, X. Li, Y. Ye, J. Kang, and X. Huang, “Visual grounding in remote sensing images,” in Proceedings of the 30th ACM International Conference on Multimedia , 2022, pp. 404–412
2022
Cited alongside, same era.
A. Toker, L. Kondmann, M. Weber, M. Eisenberger, A. Camero, J. Hu, A. P. Hoderlein, Ç. Şenaras, T. Davis, D. Cremers et al. , “Dynamicearthnet: Daily multi-spectral satellite dataset for semantic change segmentation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 21 158–21 167
2022
Cited alongside, same era.
C. Tao, J. Qi, W. Lu, H. Wang, and H. Li, “Remote sensing image scene classification with self-supervised paradigm under limited labeled samples,” IEEE Geoscience and Remote Sensing Letters , vol. 19, pp. 1–5, 2022
2022
Cited alongside, same era.
P. Jiang, D. Ergu, F. Liu, Y. Cai, and B. Ma, “A review of yolo algorithm developments,” Procedia computer science , vol. 199, pp. 1066–1073, 2022
2022
Cited alongside, same era.
Z. Liu, H. Mao, C.-Y. Wu, C. Feichtenhofer, T. Darrell, and S. Xie, “A convnet for the 2020s,” in 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 11 966–11 976
2022
Cited alongside, same era.
K. He, X. Chen, S. Xie, Y. Li, P. Dollár, and R. Girshick, “Masked autoencoders are scalable vision learners,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2022, pp. 16 000–16 009
2022
Cited alongside, same era.
J. Yu, Z. Wang, V. Vasudevan, L. Yeung, M. Seyedhosseini, and Y. Wu, “Coca: Contrastive captioners are image-text foundation models,” Transactions on Machine Learning Research , 2022. [Online]. Available: https://openreview.net/forum?id=Ee277P3AYC
2022
Cited alongside, same era.
Y. Zhang, B. Kang, B. Hooi, S. Yan, and J. Feng, “Deep long-tailed learning: A survey,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 45, no. 9, pp. 10 795–10 816, 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
C. Liu, K. Chen, H. Zhang, Z. Qi, Z. Zou, and Z. Shi, “Change-agent: Towards interactive comprehensive remote sensing change interpretation and analysis,” IEEE Transactions on Geoscience and Remote Sensing , 2024
2024
Later among the works it cites.
V. Vivanco Cepeda, G. K. Nayak, and M. Shah, “Geoclip: Clip-inspired alignment between locations and images for effective worldwide geo-localization,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Later among the works it cites.
S. Xu, C. Zhang, L. Fan, G. Meng, S. Xiang, and J. Ye, “Addressclip: Empowering vision-language models for city-wide image address localization,” in European Conference on Computer Vision . Springer, 2024, pp. 76–92
2024
Later among the works it cites.
G. Cheng, Y. Huang, X. Li, S. Lyu, Z. Xu, H. Zhao, Q. Zhao, and S. Xiang, “Change detection methods for remote sensing in the last decade: A comprehensive review,” Remote Sensing , vol. 16, no. 13, p. 2355, 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
M. Oquab, T. Darcet, T. Moutakanni, H. V. Vo, M. Szafraniec, V. Khalidov, P. Fernandez, D. HAZIZA, F. Massa, A. El-Nouby, M. Assran, N. Ballas, W. Galuba, R. Howes, P.-Y. Huang, S.-W. Li, I. Misra, M. Rabbat, V. Sharma, G. Synnaeve, H. Xu, H. Jegou, J. Mairal, P. Labatut, A. Joulin, and P. Bojanowski, “DINOv2: Learning robust visual features without supervision,” Transactions on Machine Learning Research , 2024, featured Certification. [Online]. Available: https://openreview.net/forum?id=a68SUt6zFt
2024
Later among the works it cites.
Q. Lin, S. Wang, X. Ye, R. Wang, R. Yang, and L. Jiao, “Clip-based grid features and masking for remote sensing image captioning,” IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing , 2024
2024
Later among the works it cites.
K. Chen, C. Liu, H. Chen, H. Zhang, W. Li, Z. Zou, and Z. Shi, “Rsprompter: Learning to prompt for remote sensing instance segmentation based on visual foundation model,” IEEE Transactions on Geoscience and Remote Sensing , 2024
2024
Later among the works it cites.
F. Liu, D. Chen, Z. Guan, X. Zhou, J. Zhu, Q. Ye, L. Fu, and J. Zhou, “Remoteclip: A vision language foundation model for remote sensing,” IEEE Transactions on Geoscience and Remote Sensing , 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
Z. Dong, Y. Gu, and T. Liu, “Generative convnet foundation model with sparse modeling and low-frequency reconstruction for remote sensing image interpretation,” IEEE Transactions on Geoscience and Remote Sensing , vol. 62, pp. 1–16, 2024
2024
Later among the works it cites.
V. Nedungadi, A. Kariryaa, S. Oehmcke, S. Belongie, C. Igel, and N. Lang, “Mmearth: Exploring multi-modal pretext tasks for geospatial representation learning,” 2024
2024
Later among the works it cites.
X. Guo, J. Lao, B. Dang, Y. Zhang, L. Yu, L. Ru, L. Zhong, Z. Huang, K. Wu, D. Hu et al. , “Skysense: A multi-modal remote sensing foundation model towards universal interpretation for earth observation imagery,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 27 672–27 683
2024
Later among the works it cites.
Z. Xiong, Y. Wang, F. Zhang, and X. X. Zhu, “One for all: Toward unified foundation models for earth vision,” in IGARSS 2024-2024 IEEE International Geoscience and Remote Sensing Symposium . IEEE, 2024, pp. 2734–2738
2024
Later among the works it cites.
Z. Xiong, Y. Wang, F. Zhang, A. J. Stewart, J. Hanna, D. Borth, I. Papoutsis, B. Le Saux, G. Camps-Valls, and X. X. Zhu, “Neural plasticity-inspired foundation model for observing the earth crossing modalities,” arXiv e-prints , pp. arXiv–2403, 2024
2024
Later among the works it cites.
K. Chen, C. Liu, H. Chen, H. Zhang, W. Li, Z. Zou, and Z. Shi, “Rsprompter: Learning to prompt for remote sensing instance segmentation based on visual foundation model,” IEEE Transactions on Geoscience and Remote Sensing , vol. 62, pp. 1–17, 2024
2024
Later among the works it cites.
W. Jiang, J. Zhang, D. Wang, Q. Zhang, Z. Wang, and B. Du, “Lemevit: efficient vision transformer with learnable meta tokens for remote sensing image interpretation,” in Proceedings of the Thirty-Third International Joint Conference on Artificial Intelligence , ser. IJCAI ’24, 2024. [Online]. Available: https://doi.org/10.24963/ijcai.2024/103
2024
Later among the works it cites.
2024
Later among the works it cites.
I. Dumeur, S. Valero, and J. Inglada, “Self-supervised spatio-temporal representation learning of satellite image time series,” IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing , vol. 17, pp. 4350–4367, 2024
2024
Later among the works it cites.
Z. Wang, R. Prabha, T. Huang, J. Wu, and R. Rajagopal, “Skyscript: A large and semantically diverse vision-language dataset for remote sensing,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 38, no. 6, 2024, pp. 5805–5813
2024
Later among the works it cites.
Z. Zhang, T. Zhao, Y. Guo, and J. Yin, “Rs5m and georsclip: A large scale vision-language dataset and a large vision-language model for remote sensing,” IEEE Transactions on Geoscience and Remote Sensing , 2024
2024
Later among the works it cites.
S. Mo, M. Kim, K. Lee, and J. Shin, “S-clip: Semi-supervised vision-language learning using few specialist captions,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Later among the works it cites.
U. Mall, C. P. Phoo, M. K. Liu, C. Vondrick, B. Hariharan, and K. Bala, “Remote sensing vision-language foundation models without annotations via ground remote alignment,” in The Twelfth International Conference on Learning Representations , 2024. [Online]. Available: https://openreview.net/forum?id=w9tc699w3Z
2024
Later among the works it cites.
C. Yang, Z. Li, and L. Zhang, “Bootstrapping interactive image-text alignment for remote sensing image captioning,” IEEE Transactions on Geoscience and Remote Sensing , 2024
2024
Later among the works it cites.
A. Sebaq and M. ElHelw, “Rsdiff: Remote sensing image generation from text using diffusion model,” Neural Computing and Applications , vol. 36, no. 36, pp. 23 103–23 111, 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
Z. Yu, C. Liu, L. Liu, Z. Shi, and Z. Zou, “Metaearth: A generative foundation model for global-scale remote sensing image generation,” IEEE Transactions on Pattern Analysis and Machine Intelligence , 2024
2024
Later among the works it cites.
D. Muhtar, Z. Li, F. Gu, X. Zhang, and P. Xiao, “Lhrs-bot: Empowering remote sensing with vgi-enhanced large multimodal language model,” in European Conference on Computer Vision . Springer, 2024, pp. 440–457
2024
Later among the works it cites.
2024
Later among the works it cites.
K. Kuckreja, M. S. Danish, M. Naseer, A. Das, S. Khan, and F. S. Khan, “Geochat: Grounded large vision-language model for remote sensing,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 27 831–27 840
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
W. Zhang, M. Cai, T. Zhang, Y. Zhuang, and X. Mao, “Earthgpt: A universal multi-modal large language model for multi-sensor image comprehension in remote sensing domain,” IEEE Transactions on Geoscience and Remote Sensing , 2024
2024
Later among the works it cites.
W. Zhang, M. Cai, T. Zhang, G. Lei, Y. Zhuang, and X. Mao, “Popeye: A unified visual-language model for multi-source ship detection from remote sensing imagery,” IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing , 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
Y. Bazi, L. Bashmal, M. M. Al Rahhal, R. Ricci, and F. Melgani, “Rs-llava: A large vision-language model for joint captioning and question answering in remote sensing imagery,” Remote Sensing , vol. 16, no. 9, p. 1477, 2024
2024
Later among the works it cites.
H. Guo, X. Su, C. Wu, B. Du, L. Zhang, and D. Li, “Remote sensing chatgpt: Solving remote sensing tasks with chatgpt and visual models,” in IGARSS 2024-2024 IEEE International Geoscience and Remote Sensing Symposium . IEEE, 2024, pp. 11 474–11 478
2024
Later among the works it cites.
2024
Later among the works it cites.
X. Wanyan, S. Seneviratne, S. Shen, and M. Kirley, “Extending global-local view alignment for self-supervised learning with remote sensing imagery,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops , June 2024, pp. 2443–2453
2024
Later among the works it cites.
B. Han, S. Zhang, X. Shi, and M. Reichstein, “Bridging remote sensors with multisensor geospatial foundation models,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 27 852–27 862
2024
Later among the works it cites.
M. Noman, M. Naseer, H. Cholakkal, R. M. Anwer, S. Khan, and F. S. Khan, “Rethinking transformers pre-training for multi-spectral satellite imagery,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 27 811–27 819
2024
Later among the works it cites.
J. Tian, J. Lei, J. Zhang, W. Xie, and Y. Li, “Swimdiff: Scene-wide matching contrastive learning with diffusion constraint for remote sensing image,” IEEE Transactions on Geoscience and Remote Sensing , vol. 62, pp. 1–13, 2024
2024
Later among the works it cites.
W. Li, W. Yang, T. Liu, Y. Hou, Y. Li, Z. Liu, Y. Liu, and L. Liu, “Predicting gradient is better: Exploring self-supervised learning for sar atr with a joint-embedding predictive architecture,” ISPRS Journal of Photogrammetry and Remote Sensing , vol. 218, pp. 326–338, 2024
2024
Later among the works it cites.
X. Li, D. Hong, and J. Chanussot, “S2mae: A spatial-spectral pretraining foundation model for spectral remote sensing data,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , June 2024, pp. 24 088–24 097
2024
Later among the works it cites.
D. Hong, B. Zhang, X. Li, Y. Li, C. Li, J. Yao, N. Yokoya, H. Li, P. Ghamisi, X. Jia, A. Plaza, P. Gamba, J. A. Benediktsson, and J. Chanussot, “Spectralgpt: Spectral remote sensing foundation model,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 46, no. 8, pp. 5227–5244, 2024
2024
Later among the works it cites.
Y. Wang, H. H. Hernández, C. M. Albrecht, and X. X. Zhu, “Feature guided masked autoencoder for self-supervised learning in remote sensing,” IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing , 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
M. Tang, A. Cozma, K. Georgiou, and H. Qi, “Cross-scale mae: A tale of multiscale exploitation in remote sensing,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Later among the works it cites.
A. Fuller, K. Millard, and J. R. Green, “Croma: remote sensing representations with contrastive radar-optical masked autoencoders,” in Proceedings of the 37th International Conference on Neural Information Processing Systems , ser. NIPS ’23. Red Hook, NY, USA: Curran Associates Inc., 2024
2024
Later among the works it cites.
D. Wang, J. Zhang, M. Xu, L. Liu, D. Wang, E. Gao, C. Han, H. Guo, B. Du, D. Tao et al. , “Mtp: Advancing remote sensing foundation model via multi-task pretraining,” IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing , 2024
2024
Later among the works it cites.
Z. Huang, M. Zhang, Y. Gong, Q. Liu, and Y. Wang, “Generic knowledge boosted pretraining for remote sensing images,” IEEE Transactions on Geoscience and Remote Sensing , vol. 62, pp. 1–13, 2024
2024
Later among the works it cites.
Y. Wang, C. M. Albrecht, and X. X. Zhu, “Multi-label guided soft contrastive learning for efficient earth observation pretraining,” IEEE Transactions on Geoscience and Remote Sensing , 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
D. Wang, J. Zhang, B. Du, M. Xu, L. Liu, D. Tao, and L. Zhang, “Samrs: Scaling-up remote sensing segmentation dataset with segment anything model,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
K. Cha, J. Seo, and T. Lee, “A billion-scale foundation model for remote sensing images,” IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing , pp. 1–17, 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
Z. Li, B. Hou, S. Ma, Z. Wu, X. Guo, B. Ren, and L. Jiao, “Masked angle-aware autoencoder for remote sensing images,” in ECCV . Berlin, Heidelberg: Springer-Verlag, 2024
2024
Later among the works it cites.
G. Astruc, N. Gonthier, C. Mallet, and L. Landrieu, “OmniSat: Self-supervised modality fusion for Earth observation,” ECCV , 2024
2024
Later among the works it cites.
B. Zhang, P. Zhang, X. Dong, Y. Zang, and J. Wang, “Long-clip: Unlocking the long-text capability of clip,” in European Conference on Computer Vision . Springer, 2024, pp. 310–325
2024
Later among the works it cites.
S. Moon, A. Madotto, Z. Lin, T. Nagarajan, M. Smith, S. Jain, C.-F. Yeh, P. Murugesan, P. Heidari, Y. Liu, K. Srinet, B. Damavandi, and A. Kumar, “Anymal: An efficient and scalable any-modality augmented language model,” in Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: Industry Track , F. Dernoncourt, D. Preoţiuc-Pietro, and A. Shimorina, Eds. Miami, Florida, US: Association for Computational Linguistics, Nov. 2024, pp. 1314–1332
2024
Later among the works it cites.
Y. Yuan, W. Li, J. Liu, D. Tang, X. Luo, C. Qin, L. Zhang, and J. Zhu, “Osprey: Pixel understanding with visual instruction tuning,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 28 202–28 211
2024
Later among the works it cites.
J. Hu, Y. Yao, C. Wang, S. WANG, Y. Pan, Q. Chen, T. Yu, H. Wu, Y. Zhao, H. Zhang, X. Han, Y. Lin, J. Xue, dahai li, Z. Liu, and M. Sun, “Large multilingual models pivot zero-shot multimodal learning across languages,” in The Twelfth International Conference on Learning Representations , 2024. [Online]. Available: https://openreview.net/forum?id=Kuh5qgCGCp
2024
Later among the works it cites.
M. Rußwurm, K. Klemmer, E. Rolf, R. Zbinden, and D. Tuia, “Geographic location encoding with spherical harmonics and sinusoidal representation networks,” in The Twelfth International Conference on Learning Representations , 2024
2024
Later among the works it cites.
W. Hong, W. Wang, Q. Lv, J. Xu, W. Yu, J. Ji, Y. Wang, Z. Wang, Y. Dong, M. Ding et al. , “Cogagent: A visual language model for gui agents,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 14 281–14 290
2024
Later among the works it cites.
Z. Li, B. Yang, Q. Liu, Z. Ma, S. Zhang, J. Yang, Y. Sun, Y. Liu, and X. Bai, “Monkey: Image resolution and text label are important things for large multi-modal models,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 26 763–26 773
2024
Later among the works it cites.
Z. Lin, D. Liu, R. Zhang, P. Gao, L. Qiu, H. Xiao, H. Qiu, W. Shao, K. Chen, J. Han, S. Huang, Y. Zhang, X. He, Y. Qiao, and H. Li, “Sphinx: A mixer of weights, visual embeddings and image scales for multi-modal large language models,” in ECCV . Berlin, Heidelberg: Springer-Verlag, 2024, p. 36–55
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
C. Yang, Z. Li, and L. Zhang, “Bootstrapping interactive image-text alignment for remote sensing image captioning,” IEEE Transactions on Geoscience and Remote Sensing , 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
S. Liu, Z. Zeng, T. Ren, F. Li, H. Zhang, J. Yang, Q. Jiang, C. Li, J. Yang, H. Su et al. , “Grounding dino: Marrying dino with grounded pre-training for open-set object detection,” in European Conference on Computer Vision . Springer, 2025, pp. 38–55
2025
Closest in time.
Z. Xi, W. Chen, X. Guo, W. He, Y. Ding, B. Hong, M. Zhang, J. Wang, S. Jin, E. Zhou et al. , “The rise and potential of large language model based agents: A survey,” Science China Information Sciences , vol. 68, no. 2, p. 121101, 2025
2025
Closest in time.