Fetching the paper…
Reading the bibliography…
General-purpose foundation models have led to recent breakthroughs in artificial intelligence.
1909
Earlier work this paper cites.
S. Suzuki and K. Abe, “Topological structural analysis of digitized binary images by border following,”
1985
Earlier work this paper cites.
Z. Cao, T. Qin, T. Liu, M. Tsai, and H. Li, “Learning to rank: from pairwise approach to listwise approach,” in
2007
Earlier work this paper cites.
Y. Yang and S. Newsam, “Bag-of-visual-words and spatial extensions for land-use classification,” in
2010
Earlier work this paper cites.
C. Zauner, “Implementation and benchmarking of perceptual image hash functions,” 2010
2010
Earlier work this paper cites.
2010
Earlier work this paper cites.
G.-S. Xia, W. Yang, J. Delon, Y. Gousseau, H. Sun, and H. Maître, “Structural high-resolution satellite image indexing,” 2010
2010
Earlier work this paper cites.
“Vaihingen dataset,”
2012
Earlier work this paper cites.
“Potsdam dataset,”
2012
Earlier work this paper cites.
Q. Zou, L. Ni, T. Zhang, and Q. Wang, “Deep learning based feature selection for remote sensing scene classification,”
2015
Earlier work this paper cites.
A. Robicquet, A. Sadeghian, A. Alahi, and S. Savarese, “Learning social etiquette: Human trajectory understanding in crowded scenes,” in
2016
Earlier work this paper cites.
L. Zhao, P. Tang, and L. Huo, “Feature significance-based multibag-of-visual-words model for remote sensing image scene classification,”
2016
Earlier work this paper cites.
G.-S. Xia, J. Hu, F. Hu, B. Shi, X. Bai, Y. Zhong, L. Zhang, and X. Lu, “Aid: A benchmark data set for performance evaluation of aerial scene classification,”
2016
Earlier work this paper cites.
B. Zhao, Y. Zhong, G.-S. Xia, and L. Zhang, “Dirichlet-derived multiple topic scene classification model for high spatial resolution remote sensing imagery,”
2016
Earlier work this paper cites.
K. Sohn, “Improved deep metric learning with multi-class n-pair loss objective,” in
2016
Earlier work this paper cites.
X. Lu, B. Wang, X. Zheng, and X. Li, “Exploring models and data for remote sensing image caption generation,”
2017
Earlier work this paper cites.
G.-S. Xia, X. Bai, J. Ding, Z. Zhu, S. J. Belongie, J. Luo, M. Datcu, M. Pelillo, and L. Zhang, “Dota: A large-scale dataset for object detection in aerial images,”
2017
Earlier work this paper cites.
Z. Liu, L. Yuan, L. Weng, and Y. Yang, “A high resolution optical satellite image dataset for ship recognition and some new baselines,” in
2017
Earlier work this paper cites.
M.-R. Hsieh, Y.-L. Lin, and W. H. Hsu, “Drone-based object counting by spatially regularized regional proposal network,”
2017
Earlier work this paper cites.
F. Faghri, D. J. Fleet, J. R. Kiros, and S. Fidler, “Vse++: Improving visual-semantic embeddings with hard negatives,” in
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
P. Helber, B. Bischke, A. R. Dengel, and D. Borth, “Eurosat: A novel dataset and deep learning benchmark for land use and land cover classification,”
2017
Earlier work this paper cites.
H. Li, C. Tao, Z. Wu, J. Chen, J. Gong, and M. Deng, “Rsi-cb: A large-scale remote sensing image classification benchmark using crowdsourced data,”
2017
Earlier work this paper cites.
G. Cheng, J. Han, and X. Lu, “Remote sensing image scene classification: Benchmark and state of the art,”
2017
Earlier work this paper cites.
P. Zhu, L. Wen, X. Bian, H. Ling, and Q. Hu, “Vision meets drones: A challenge,”
2018
Earlier work this paper cites.
K. Lee, X. Chen, G. Hua, H. Hu, and X. He, “Stacked cross attention for image-text matching,” in
2018
Earlier work this paper cites.
J. Devlin, M. Chang, K. Lee, and K. Toutanova, “BERT: pre-training of deep bidirectional transformers for language understanding,” in
2019
Earlier work this paper cites.
R. Geirhos, P. Rubisch, C. Michaelis, M. Bethge, F. A. Wichmann, and W. Brendel, “Imagenet-trained cnns are biased towards texture; increasing shape bias improves accuracy and robustness,” in
2019
Earlier work this paper cites.
N. Jean, S. Wang, A. Samar, G. Azzari, D. B. Lobell, and S. Ermon, “Tile2vec: Unsupervised representation learning for spatially distributed data,” in
2019
Earlier work this paper cites.
Y. Zhang, Y. Yuan, Y. Feng, and X. Lu, “Hierarchical and robust convolutional neural network for very high-resolution remote sensing object detection,”
2019
Earlier work this paper cites.
S. W. Zamir, A. Arora, A. Gupta, S. H. Khan, G. Sun, F. S. Khan, F. Zhu, L. Shao, G. Xia, and X. Bai, “isaid: A large-scale dataset for instance segmentation in aerial images,” in
2019
Earlier work this paper cites.
T. Wang, X. Xu, Y. Yang, A. Hanjalic, H. T. Shen, and J. Song, “Matching images and text with multi-modal tensor fusion and re-ranking,”
2019
Earlier work this paper cites.
Q. Wang, S. Liu, J. Chanussot, and X. Li, “Scene classification with recurrent attention of vhr remote sensing images,”
2019
Earlier work this paper cites.
T. Chen, S. Kornblith, M. Norouzi, and G. E. Hinton, “A simple framework for contrastive learning of visual representations,” in
2020
Earlier work this paper cites.
T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei, “Language models are few-shot learners,” in
2020
Earlier work this paper cites.
J. Kang, R. Fernández-Beltran, P. Duan, S. Liu, and A. J. Plaza, “Deep unsupervised embedding for remotely sensed images based on spatially augmented momentum contrast,”
2020
Earlier work this paper cites.
Z. Zhao, Z. Luo, J. Li, C. Chen, and Y. Piao, “When self-supervised learning meets scene classification: Remote sensing scene classification based on a multitask learning framework,”
2020
Earlier work this paper cites.
T. A. M. Ali, Y. Bazi, M. M. A. Rahhal, M. L. Mekhalfi, L. Rangarajan, and M. A. A. Zuair, “Textrs: Deep bidirectional triplet network for matching text to remote sensing images,”
2020
Cited alongside, same era.
M. M. A. Rahhal, Y. Bazi, T. Abdullah, M. L. Mekhalfi, and M. A. A. Zuair, “Deep unsupervised embedding for remote sensing image retrieval using textual cues,”
2020
Cited alongside, same era.
H. Chen and Z. Shi, “A spatial-temporal attention-based method and a new dataset for remote sensing image change detection,”
2020
Cited alongside, same era.
S. Vujasinovi’c, S. Becker, T. Breuer, S. Bullinger, N. Scherer-Negenborn, and M. Arens, “Integration of the 3d environment for uav onboard visual object tracking,”
2020
Cited alongside, same era.
G. Hoxha, F. Melgani, and B. Demir, “Toward remote sensing image retrieval under a deep image captioning perspective,”
2022
Later among the works it cites.
N. Mu, A. Kirillov, D. Wagner, and S. Xie, “Slip: Self-supervision meets language-image pre-training,” in
2022
Later among the works it cites.
Y. Li, F. Liang, L. Zhao, Y. Cui, W. Ouyang, J. Shao, F. Yu, and J. Yan, “Supervision exists everywhere: A data efficient contrastive language-image pre-training paradigm,” in
2022
Later among the works it cites.
J. Yu, Z. Wang, V. Vasudevan, L. Yeung, M. Seyedhosseini, and Y. Wu, “Coca: Contrastive captioners are image-text foundation models,”
2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2020
Cited alongside, same era.
2021
Cited alongside, same era.
K. He, X. Chen, S. Xie, Y. Li, P. Doll’ar, and R. B. Girshick, “Masked autoencoders are scalable vision learners,”
2021
Cited alongside, same era.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever, “Learning transferable visual models from natural language supervision,” in
2021
Cited alongside, same era.
Z. Xie, Z. Zhang, Y. Cao, Y. Lin, J. Bao, Z. Yao, Q. Dai, and H. Hu, “Simmim: a simple framework for masked image modeling,”
2021
Cited alongside, same era.
Z. Yuan, W. Zhang, K. Fu, X. Li, C. Deng, H. Wang, and X. Sun, “Exploring a fine-grained multiscale method for cross-modal remote sensing image retrieval,”
2021
Cited alongside, same era.
W. Li, K. Chen, H. Chen, and Z. Shi, “Geographical knowledge-driven representation learning for remote sensing images,”
2021
Cited alongside, same era.
V. Stojnic and V. Risojevic, “Self-supervised learning of remote sensing scene representations using contrastive multiview coding,” in
2021
Cited alongside, same era.
2022
Later among the works it cites.
2022
Later among the works it cites.
Y. Zhang, H. Jiang, Y. Miura, C. D. Manning, and C. P. Langlotz, “Contrastive learning of medical visual representations from paired images and text,” in
2022
Later among the works it cites.
Z. Wang, Z. Wu, D. Agarwal, and J. Sun, “Medclip: Contrastive learning from unpaired medical images and text,” in
2022
Later among the works it cites.
W. Shin, J. Park, T. Woo, Y. Cho, K. Oh, and H. Song, “e-clip: Large-scale vision-language representation learning in e-commerce,”
2022
Later among the works it cites.
H. Li, W. Xiong, Y. Cui, and Z. Xiong, “A fusion-based contrastive learning model for cross-modal remote sensing retrieval,”
2022
Later among the works it cites.
2023
Closest in time.
OpenAI, “Gpt-4 technical report,”
2023
Closest in time.
K. Cha, J. Seo, and T. Lee, “A billion-scale foundation model for remote sensing images,”
2023
Closest in time.
2023
Closest in time.
X. Kong and X. Zhang, “Understanding masked image modeling via learning occlusion invariant feature,” pp. 6241–6251, 2023. [Online]. Available:
2023
Closest in time.
S. Li, D. Wu, F. Wu, Z. Zang, and S. Z. Li, “Architecture-agnostic masked image modeling - from vit back to CNN,” in
2023
Closest in time.
L. Kong, M. Q. Ma, G. Chen, E. P. Xing, Y. Chi, L. Morency, and K. Zhang, “Understanding masked autoencoders via hierarchical latent variable models,” in
2023
Closest in time.
Z. Xie, Z. Geng, J. Hu, Z. Zhang, H. Hu, and Y. Cao, “Revealing the dark secrets of masked image modeling,” in
2023
Closest in time.
C. Tao, X. Zhu, W. Su, G. Huang, B. Li, J. Zhou, Y. Qiao, X. Wang, and J. Dai, “Siamese image modeling for self-supervised vision representation learning,” in
2023
Closest in time.
S. Tukra, F. Hoffman, and K. Chatfield, “Improving visual representation learning through perceptual understanding,” in
2023
Closest in time.
N. Park, W. Kim, B. Heo, T. Kim, and S. Yun, “What do self-supervised vision transformers learn?” in
2023
Closest in time.
2023
Closest in time.
Y. Xiao, Q. Yuan, K. Jiang, J. He, Y. Wang, and L. Zhang, “From degrade to upgrade: Learning a self-supervised degradation guided adaptive network for blind remote sensing image super-resolution,”
2023
Closest in time.
J. Zhang, J. Huang, S. Jin, and S. Lu, “Vision-language models for vision tasks: A survey,”
2023
Closest in time.
W. Zhang, J. Li, S. Li, J. Chen, W. Zhang, X. Gao, and X. Sun, “Hypersphere-based remote sensing cross-modal text-image retrieval via curriculum learning,”
2023
Closest in time.
2023
Closest in time.
F. Liu, D. Chen, X. Du, R. Gao, and F. Xu, “Mep-3m: A large-scale multi-modal e-commerce product dataset,”
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
D. Chen, J. Liu, W. Dai, and B. Wang, “Visual instruction tuning with polite flamingo,”
2023
Closest in time.
2023
Closest in time.
J. Pan, Q. Ma, and C. Bai, “A prior instruction representation framework for remote sensing image-text retrieval,” in
2023
Closest in time.
2023
Closest in time.
X. Zhai, B. Mustafa, A. Kolesnikov, and L. Beyer, “Sigmoid loss for language image pre-training,”
2023
Closest in time.
F. Liu, T. Zhang, W. Dai, W. Cai, X. Zhou, and D. Chen, “Few-shot adaptation of multi-modal foundation models: A survey,” 2024
2024
Closest in time.