Fetching the paper…
Reading the bibliography…
Remote Sensing (RS) is a crucial technology for observing, monitoring, and interpreting our planet, with broad applications across geoscience, economics, humanitarian fields, etc.
J. W. Rouse, R. H. Haas, J. A. Schell, D. W. Deering
1974
Earlier work this paper cites.
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” in
2009
Earlier work this paper cites.
Y. Yang and S. Newsam, “Bag-of-visual-words and spatial extensions for land-use classification,” in
2010
Earlier work this paper cites.
J. Sohl-Dickstein, E. Weiss, N. Maheswaranathan, and S. Ganguli, “Deep unsupervised learning using nonequilibrium thermodynamics,” in
2015
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in
2016
Earlier work this paper cites.
B. Qu, X. Li, D. Tao, and X. Lu, “Deep semantic understanding of high resolution remote sensing image,” in
2016
Earlier work this paper cites.
A. Vaswani, “Attention is all you need,”
2017
Earlier work this paper cites.
X. Lu, B. Wang, X. Zheng, and X. Li, “Exploring models and data for remote sensing image caption generation,”
2017
Earlier work this paper cites.
G. Cheng, J. Han, and X. Lu, “Remote sensing image scene classification: Benchmark and state of the art,”
2017
Earlier work this paper cites.
G. Christie, N. Fendley, J. Wilson, and R. Mukherjee, “Functional map of the world,” in
2018
Earlier work this paper cites.
A. v. d. Oord, Y. Li, and O. Vinyals, “Representation learning with contrastive predictive coding,”
2018
Earlier work this paper cites.
G. Van Horn, O. Mac Aodha, Y. Song, Y. Cui, C. Sun, A. Shepard, H. Adam, P. Perona, and S. Belongie, “The inaturalist species classification and detection dataset,” in
2018
Earlier work this paper cites.
G.-S. Xia, X. Bai, J. Ding, Z. Zhu, S. Belongie, J. Luo, M. Datcu, M. Pelillo, and L. Zhang, “Dota: A large-scale dataset for object detection in aerial images,” in
2018
Earlier work this paper cites.
A. Radford, “Improving language understanding by generative pre-training,” 2018
2018
Earlier work this paper cites.
R. C. Daudt, B. Le Saux, A. Boulch, and Y. Gousseau, “Urban change detection for multispectral earth observation using convolutional neural networks,” in
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
G. Sumbul, M. Charfuelan, B. Demir, and V. Markl, “Bigearthnet: A large-scale benchmark archive for remote sensing image understanding,” in
2019
Earlier work this paper cites.
N. Jean, S. Wang, A. Samar, G. Azzari, D. Lobell, and S. Ermon, “Tile2vec: Unsupervised representation learning for spatially distributed data,” in
2019
Earlier work this paper cites.
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever
2019
Earlier work this paper cites.
P. Helber, B. Bischke, A. Dengel, and D. Borth, “Eurosat: A novel dataset and deep learning benchmark for land use and land cover classification,”
2019
Earlier work this paper cites.
L. Jing and Y. Tian, “Self-supervised visual feature learning with deep neural networks: A survey,”
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
K. He, H. Fan, Y. Wu, S. Xie, and R. Girshick, “Momentum contrast for unsupervised visual representation learning,” in
2020
Earlier work this paper cites.
T. Chen, S. Kornblith, M. Norouzi, and G. Hinton, “A simple framework for contrastive learning of visual representations,” in
2020
Earlier work this paper cites.
X. Chen, H. Fan, R. Girshick, and K. He, “Improved baselines with momentum contrastive learning,”
2020
Earlier work this paper cites.
J.-B. Grill, F. Strub, F. Altché, C. Tallec, P. Richemond, E. Buchatskaya, C. Doersch, B. Avila Pires, Z. Guo, M. Gheshlaghi Azar
2020
Earlier work this paper cites.
M. Caron, I. Misra, J. Mairal, P. Goyal, P. Bojanowski, and A. Joulin, “Unsupervised learning of visual features by contrasting cluster assignments,”
2020
Earlier work this paper cites.
S. Lobry, D. Marcos, J. Murray, and D. Tuia, “Rsvqa: Visual question answering for remote sensing data,”
2020
Earlier work this paper cites.
T. B. Brown, “Language models are few-shot learners,”
2020
Earlier work this paper cites.
J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models,”
2020
Earlier work this paper cites.
K. Li, G. Wan, G. Cheng, L. Meng, and J. Han, “Object detection in optical remote sensing images: A survey and a new benchmark,”
2020
Earlier work this paper cites.
2021
Earlier work this paper cites.
O. Manas, A. Lacoste, X. Giró-i Nieto, D. Vazquez, and P. Rodriguez, “Seasonal contrast: Unsupervised pre-training from uncurated remote sensing data,” in
2021
Earlier work this paper cites.
K. Ayush, B. Uzkent, C. Meng, K. Tanmay, M. Burke, D. Lobell, and S. Ermon, “Geography-aware self-supervised learning,” in
2021
Earlier work this paper cites.
W. Li, K. Chen, H. Chen, and Z. Shi, “Geographical knowledge-driven representation learning for remote sensing images,”
2021
Earlier work this paper cites.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark
2021
Earlier work this paper cites.
G. Ilharco, M. Wortsman, R. Wightman, C. Gordon, N. Carlini, R. Taori, A. Dave, V. Shankar, H. Namkoong, J. Miller, H. Hajishirzi, A. Farhadi, and L. Schmidt, “Openclip,” Jul. 2021, if you use this software, please cite it as below. [Online]. Available:
2021
Earlier work this paper cites.
S. Lobry, B. Demir, and D. Tuia, “Rsvqa meets bigearthnet: a new, large-scale, visual question answering dataset for remote sensing,” in
2021
Earlier work this paper cites.
X. Zheng, B. Wang, X. Du, and X. Lu, “Mutual attention inception network for remote sensing visual question answering,”
2021
Earlier work this paper cites.
M. Rahnemoonfar, T. Chowdhury, A. Sarkar, D. Varshney, M. Yari, and R. R. Murphy, “Floodnet: A high resolution aerial imagery dataset for post flood scene understanding,”
2021
Earlier work this paper cites.
Z. Yuan, W. Zhang, K. Fu, X. Li, C. Deng, H. Wang, and X. Sun, “Exploring a fine-grained multiscale method for cross-modal remote sensing image retrieval,”
2021
Earlier work this paper cites.
J. Song, C. Meng, and S. Ermon, “Denoising diffusion implicit models,” in
2021
Earlier work this paper cites.
Y. Wang, C. M. Albrecht, N. A. A. Braham, L. Mou, and X. X. Zhu, “Self-supervised learning in remote sensing: A review,”
2022
Earlier work this paper cites.
Y. Cong, S. Khanna, C. Meng, P. Liu, E. Rozi, Y. He, M. Burke, D. Lobell, and S. Ermon, “Satmae: Pre-training transformers for temporal and multi-spectral satellite imagery,”
2022
Earlier work this paper cites.
D. Wang, J. Zhang, B. Du, G.-S. Xia, and D. Tao, “An empirical study of remote sensing pretraining,”
2022
Earlier work this paper cites.
P. Akiva, M. Purri, and M. Leotta, “Self-supervised material and texture representation learning for remote sensing tasks,” in
2022
Earlier work this paper cites.
K. He, X. Chen, S. Xie, Y. Li, P. Dollár, and R. Girshick, “Masked autoencoders are scalable vision learners,” in
2022
Earlier work this paper cites.
H. Li, Y. Li, G. Zhang, R. Liu, H. Huang, Q. Zhu, and C. Tao, “Global and local contrastive self-supervised learning for semantic segmentation of hr remote sensing images,”
2022
Earlier work this paper cites.
X. Sun, P. Wang, W. Lu, Z. Zhu, X. Lu, Q. He, J. Li, X. Rong, Z. Yang, H. Chang
2022
Earlier work this paper cites.
D. Wang, Q. Zhang, Y. Xu, J. Zhang, B. Du, D. Tao, and L. Zhang, “Advancing plain vision transformer toward remote sensing foundation model,”
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
J. Li, D. Li, C. Xiong, and S. Hoi, “Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation,” in
2022
Earlier work this paper cites.
Z. Yuan, L. Mou, Z. Xiong, and X. X. Zhu, “Change detection meets visual question answering,”
2022
Earlier work this paper cites.
Q. Cheng, H. Huang, Y. Xu, Y. Zhou, H. Li, and Z. Wang, “Nwpu-captions dataset and mlca-net for remote sensing image captioning,”
2022
Earlier work this paper cites.
Y. Sun, S. Feng, X. Li, Y. Ye, J. Kang, and X. Huang, “Visual grounding in remote sensing images,” in
2022
Cited alongside, same era.
Y. Zhong, J. Yang, P. Zhang, C. Li, N. Codella, L. H. Li, L. Zhou, X. Dai, L. Yuan, Y. Li
2022
Cited alongside, same era.
A. Singh, R. Hu, V. Goswami, G. Couairon, W. Galuba, M. Rohrbach, and D. Kiela, “Flava: A foundational language and vision alignment model,” in
2022
Cited alongside, same era.
E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen, “LoRA: Low-rank adaptation of large language models,” in
2022
Cited alongside, same era.
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer, “High-resolution image synthesis with latent diffusion models,” in
2022
Cited alongside, same era.
2024
Closest in time.
X. Guo, J. Lao, B. Dang, Y. Zhang, L. Yu, L. Ru, L. Zhong, Z. Huang, K. Wu, D. Hu
2024
Closest in time.
2024
Closest in time.
B. Han, S. Zhang, X. Shi, and M. Reichstein, “Bridging remote sensors with multisensor geospatial foundation models,” in
2024
Closest in time.
M. Noman, M. Naseer, H. Cholakkal, R. M. Anwer, S. Khan, and F. S. Khan, “Rethinking transformers pre-training for multi-spectral satellite imagery,” in
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Xiao, J. Huang, D. Guan, F. Zhan, and S. Lu, “Transfer learning from synthetic to real lidar point cloud for semantic segmentation,” in
2022
Cited alongside, same era.
L. Wang, R. Li, C. Zhang, S. Fang, C. Duan, X. Meng, and P. M. Atkinson, “Unetformer: A unet-like transformer for efficient semantic segmentation of remote sensing urban scene imagery,”
2022
Cited alongside, same era.
G. Cheng, J. Wang, K. Li, X. Xie, C. Lang, Y. Yao, and J. Han, “Anchor-free oriented proposal generator for object detection,”
2022
Cited alongside, same era.
A. Xiao, J. Huang, D. Guan, X. Zhang, S. Lu, and L. Shao, “Unsupervised point cloud representation learning with deep neural networks: A survey,”
2023
Cited alongside, same era.
L. Jiao, Z. Huang, X. Lu, X. Liu, Y. Yang, J. Zhao, J. Zhang, B. Hou, S. Yang, F. Liu
2023
Cited alongside, same era.
M. Mendieta, B. Han, X. Shi, Y. Zhu, and C. Chen, “Towards geospatial foundation models via continual pretraining,” in
2023
Cited alongside, same era.
Y. Wang, N. A. A. Braham, Z. Xiong, C. Liu, C. M. Albrecht, and X. X. Zhu, “Ssl4eo-s12: A large-scale multimodal, multitemporal dataset for self-supervised learning in earth observation [software and data sets],”
2023
Cited alongside, same era.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
A. Fuller, K. Millard, and J. Green, “Croma: Remote sensing representations with contrastive radar-optical masked autoencoders,”
2024
Closest in time.
D. Wang, J. Zhang, M. Xu, L. Liu, D. Wang, E. Gao, C. Han, H. Guo, B. Du, D. Tao
2024
Closest in time.
Z. Huang, M. Zhang, Y. Gong, Q. Liu, and Y. Wang, “Generic knowledge boosted pre-training for remote sensing images,”
2024
Closest in time.
Z. Dong, Y. Gu, and T. Liu, “Generative convnet foundation model with sparse modeling and low-frequency reconstruction for remote sensing image interpretation,”
2024
Closest in time.
D. Hong, B. Zhang, X. Li, Y. Li, C. Li, J. Yao, N. Yokoya, H. Li, P. Ghamisi, X. Jia
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
X. Ma, Q. Wu, X. Zhao, X. Zhang, M.-O. Pun, and B. Huang, “Sam-assisted remote sensing imagery semantic segmentation with object and boundary constraints,”
2024
Closest in time.
X. Zhang, Y. Liu, Y. Lin, Q. Liao, and Y. Li, “Uv-sam: Adapting segment anything model for urban village identification,” in
2024
Closest in time.
L. Ding, K. Zhu, D. Peng, H. Tang, K. Yang, and L. Bruzzone, “Adapting segment anything model for change detection in vhr remote sensing images,”
2024
Closest in time.
Z. Zheng, Y. Zhong, L. Zhang, and S. Ermon, “Segment any change,”
2024
Closest in time.
K. Chen, C. Liu, H. Chen, H. Zhang, W. Li, Z. Zou, and Z. Shi, “Rsprompter: Learning to prompt for remote sensing instance segmentation based on visual foundation model,”
2024
Closest in time.
A. Xiao, W. Xuan, H. Qi, Y. Xing, R. Ren, X. Zhang, and S. Lu, “Cat-sam: Conditional tuning for few-shot adaptation of segmentation anything model,” in
2024
Closest in time.
L. Yang, B. Kang, Z. Huang, X. Xu, J. Feng, and H. Zhao, “Depth anything: Unleashing the power of large-scale unlabeled data,” in
2024
Closest in time.
2024
Closest in time.
J. Chen, B. Liu, A. Yu, Y. Quan, T. Li, and W. Guo, “Depth feature fusion network for building extraction in remote sensing images,”
2024
Closest in time.
A. Xiao, W. Xuan, H. Qi, Y. Xing, N. Yokoya, and S. Lu, “Segment anything with multiple modalities,”
2024
Closest in time.
Q. team, “Qwen2-vl,” 2024
2024
Closest in time.
W. Zhang, M. Cai, T. Zhang, Y. Zhuang, and X. Mao, “Earthgpt: A universal multi-modal large language model for multi-sensor image comprehension in remote sensing domain,”
2024
Closest in time.
K. Li, G. Vosselman, and M. Y. Yang, “Hrvqa: A visual question answering benchmark for high-resolution aerial images,”
2024
Closest in time.
J. Wang, Z. Zheng, Z. Chen, A. Ma, and Y. Zhong, “Earthvqa: Towards queryable earth via relational reasoning-based remote sensing visual question answering,” in
2024
Closest in time.
F. Liu, D. Chen, Z. Guan, X. Zhou, J. Zhu, Q. Ye, L. Fu, and J. Zhou, “Remoteclip: A vision language foundation model for remote sensing,”
2024
Closest in time.
Z. Wang, R. Prabha, T. Huang, J. Wu, and R. Rajagopal, “Skyscript: A large and semantically diverse vision-language dataset for remote sensing,” in
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
T. Cheng, L. Song, Y. Ge, W. Liu, X. Wang, and Y. Shan, “Yolo-world: Real-time open-vocabulary object detection,” in
2024
Closest in time.
2024
Closest in time.
H. Liu, C. Li, Y. Li, B. Li, Y. Zhang, S. Shen, and Y. J. Lee, “Llava-next: Improved reasoning, ocr, and world knowledge,” January 2024. [Online]. Available:
2024
Closest in time.
U. Mall, C. P. Phoo, M. K. Liu, C. Vondrick, B. Hariharan, and K. Bala, “Remote sensing vision-language foundation models without annotations via ground remote alignment,” in
2024
Closest in time.
V. Vivanco Cepeda, G. K. Nayak, and M. Shah, “Geoclip: Clip-inspired alignment between locations and images for effective worldwide geo-localization,”
2024
Closest in time.
J. Bourcier, G. Dashyan, K. Alahari, and J. Chanussot, “Learning representations of satellite images from metadata supervision,” in
2024
Closest in time.
K. Kuckreja, M. S. Danish, M. Naseer, A. Das, S. Khan, and F. S. Khan, “Geochat: Grounded large vision-language model for remote sensing,” in
2024
Closest in time.
J. Luo, Z. Pang, Y. Zhang, T. Wang, L. Wang, B. Dang, J. Lao, J. Wang, J. Chen, Y. Tan
2024
Closest in time.
2024
Closest in time.
R. Manvi, S. Khanna, G. Mai, M. Burke, D. B. Lobell, and S. Ermon, “Geollm: Extracting geospatial knowledge from large language models,” in
2024
Closest in time.
Z. Zheng, S. Ermon, D. Kim, L. Zhang, and Y. Zhong, “Changen2: Multi-temporal remote sensing generative change foundation model,”
2024
Closest in time.
Y. Li, X. Li, Y. Dai, Q. Hou, L. Liu, Y. Liu, M.-M. Cheng, and J. Yang, “Lsknet: A foundation lightweight backbone for remote sensing,”
2024
Closest in time.
2024
Closest in time.
D. Wang, M. Hu, Y. Jin, Y. Miao, J. Yang, Y. Xu, X. Qin, J. Ma, L. Sun, C. Li
2025
Closest in time.
J. Li, Y. Liu, X. Wang, Y. Peng, C. Sun, S. Wang, Z. Sun, T. Ke, X. Jiang, T. Lu
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
S. Soni, A. Dudhane, H. Debary, M. Fiaz, M. A. Munir, M. S. Danish, P. Fraccaro, C. D. Watson, L. J. Klein, F. S. Khan
2025
Closest in time.
F. Wang, H. Wang, M. Chen, D. Wang, Y. Wang, Z. Guo, Q. Ma, L. Lan, W. Yang, J. Zhang
2025
Closest in time.
2025
Closest in time.
T. Shinde, “Model compression meets resolution scaling for efficient remote sensing classification,” in
2025
Closest in time.
P. Ghamisi, W. Yu, A. Marinoni, C. M. Gevaert, C. Persello, S. Selvakumaran, M. Girotto, B. P. Horton, P. Rufin, P. Hostert
2025
Closest in time.