Fetching the paper…
Reading the bibliography…
Object detection, particularly open-vocabulary object detection, plays a crucial role in Earth sciences, such as environmental monitoring, natural disaster assessment, and land-use planning.
Dimensionality reduction by learning an invariant mapping
Hadsell, R.; Chopra, S.; LeCun, Y.; and LeCun, Y. 2006 · 2006
Earlier work this paper cites.
Visualizing data using t-SNE
Van der Maaten, L.; and Hinton, G. 2008 · 2008
Earlier work this paper cites.
Multi-class geospatial object detection and geographic image classification based on collection of part detectors
Cheng, G.; Han, J.; Zhou, P.; and Guo, L. 2014 · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P.; and Ba, J. 2014 · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
Lin, T.-Y.; Maire, M.; Belongie, S.; Hays, J.; Perona, P.; Ramanan, D.; Dollár, P.; and Zitnick, C. L. 2014 · 2014
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
Ren, S.; He, K.; Girshick, R.; and Sun, J. 2015 · 2015
Earlier work this paper cites.
A high resolution optical satellite image dataset for ship recognition and some new baselines
Liu, Z.; Yuan, L.; Weng, L.; and Yang, Y. 2017 · 2017
Earlier work this paper cites.
Accurate object localization in remote sensing images based on convolutional neural networks
Long, Y.; Gong, Y.; Xiao, Z.; and Liu, Q. 2017 · 2017
Earlier work this paper cites.
YOLO9000: better, faster, stronger
Redmon, J.; and Farhadi, A. 2017 · 2017
Earlier work this paper cites.
AID: A benchmark data set for performance evaluation of aerial scene classification
Xia, G.-S.; Hu, J.; Hu, F.; Shi, B.; Bai, X.; Zhong, Y.; Zhang, L.; and Lu, X. 2017 · 2017
Earlier work this paper cites.
Lstd: A low-shot transfer detector for object detection
Chen, H.; Wang, Y.; Wang, G.; and Qiao, Y. 2018 · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J.; Chang, M.-W.; Lee, K.; and Toutanova, K. 2018 · 2018
Earlier work this paper cites.
Transfer Learning in Multilingual Neural Machine Translation with Dynamic Vocabulary
Lakew, S. M.; Erofeeva, A.; Negri, M.; Federico, M.; and Turchi, M. 2018 · 2018
Earlier work this paper cites.
xView: Objects in Context in Overhead Imagery
Lam, D.; Kuzma, R.; McGee, K.; Dooley, S.; Laielli, M.; Klaric, M.; Bulatov, Y.; and McCord, B. 2018 · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Radford, A.; Narasimhan, K.; Salimans, T.; Sutskever, I.; et al. 2018 · 2018
Earlier work this paper cites.
Neural response generation with dynamic vocabularies
Wu, Y.; Wu, W.; Yang, D.; Xu, C.; and Li, Z. 2018 · 2018
Earlier work this paper cites.
DOTA: A Large-Scale Dataset for Object Detection in Aerial Images
Xia, G.-S.; Bai, X.; Ding, J.; Zhu, Z.; Belongie, S.; Luo, J.; Datcu, M.; Pelillo, M.; and Zhang, L. 2018 · 2018
Earlier work this paper cites.
Generalized intersection over union: A metric and a loss for bounding box regression
Rezatofighi, H.; Tsoi, N.; Gwak, J.; Sadeghian, A.; Reid, I.; and Savarese, S. 2019 · 2019
Earlier work this paper cites.
Deep learning based fossil-fuel power plant monitoring in high resolution remote sensing images: A comparative study
Zhang, H.; and Deng, Q. 2019 · 2019
Earlier work this paper cites.
Hierarchical and Robust Convolutional Neural Network for Very High-Resolution Remote Sensing Object Detection
Zhang, Y.; Yuan, Y.; Feng, Y.; and Lu, X. 2019 · 2019
Earlier work this paper cites.
Object detection in optical remote sensing images: A survey and a new benchmark
Li, K.; Wan, G.; Cheng, G.; Meng, L.; and Han, J. 2020 · 2020
Cited alongside, same era.
Geography-aware self-supervised learning
Ayush, K.; Uzkent, B.; Meng, C.; Tanmay, K.; Burke, M.; Lobell, D.; and Ermon, S. 2021 · 2021
Cited alongside, same era.
Open-vocabulary object detection via vision and language knowledge distillation
Gu, X.; Lin, T.-Y.; Kuo, W.; and Cui, Y. 2021 · 2021
Cited alongside, same era.
NWPU-RESISC45 Dataset with 12 classes
Hichri, H. 2021 · 2021
Cited alongside, same era.
Meta-learning in neural networks: A survey
Hospedales, T.; Antoniou, A.; Micaelli, P.; and Storkey, A. 2021 · 2021
Cited alongside, same era.
Swin transformer: Hierarchical vision transformer using shifted windows
Liu, Z.; Lin, Y.; Cao, Y.; Hu, H.; Wei, Y.; Zhang, Z.; Lin, S.; and Guo, B. 2021 · 2021
Reducing Semantic Confusion: Scene-aware Aggregation Network for Remote Sensing Cross-modal Retrieval
Pan, J.; Ma, Q.; Bai, C.; and Bai, C. 2023b · 2023
Later among the works it cites.
Scale-mae: A scale-aware masked autoencoder for multiscale geospatial representation learning
Reed, C. J.; Gupta, R.; Li, S.; Brockman, S.; Funk, C.; Clipp, B.; Keutzer, K.; Candido, S.; Uyttendaele, M.; and Darrell, T. 2023 · 2023
Later among the works it cites.
TOV: The original vision model for optical remote sensing image understanding via self-supervised learning
Tao, C.; Qi, J.; Zhang, G.; Zhu, Q.; Lu, W.; and Li, H. 2023 · 2023
Later among the works it cites.
SAMRS: Scaling-up Remote Sensing Segmentation Dataset with Segment Anything Model
Wang, D.; Zhang, J.; Du, B.; Xu, M.; Liu, L.; Tao, D.; and Zhang, L. 2023 · 2023
Later among the works it cites.
DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection
Zhang, H.; Li, F.; Liu, S.; Zhang, L.; Su, H.; Zhu, J.; Ni, L.; and Shum, H.-Y. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. 2021 · 2021
Cited alongside, same era.
Open-vocabulary object detection using captions
Zareian, A.; Rosa, K. D.; Hu, D. H.; and Chang, S.-F. 2021 · 2021
Cited alongside, same era.
Deformable DETR: Deformable Transformers for End-to-End Object Detection
Zhu, X.; Su, W.; Lu, L.; Li, B.; Wang, X.; and Dai, J. 2021 · 2021
Cited alongside, same era.
Slicing Aided Hyper Inference and Fine-tuning for Small Object Detection
Akyon, F. C.; Altinuc, S. O.; and Temizel, A. 2022 · 2022
Cited alongside, same era.
Satmae: Pre-training transformers for temporal and multi-spectral satellite imagery
Cong, Y.; Khanna, S.; Meng, C.; Liu, P.; Rozi, E.; He, Y.; Burke, M.; Lobell, D.; and Ermon, S. 2022 · 2022
Cited alongside, same era.
Grounded language-image pre-training
Li, L. H.; Zhang, P.; Zhang, H.; Yang, J.; Li, C.; Zhong, Y.; Wang, L.; Yuan, L.; Zhang, L.; Hwang, J.-N.; et al. 2022 · 2022
Cited alongside, same era.
Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks
Chen, Z.; Wu, J.; Wang, W.; Su, W.; Chen, G.; Xing, S.; Zhong, M.; Zhang, Q.; Zhu, X.; Lu, L.; et al. 2024 · 2024
Closest in time.
Cross-Domain Few-Shot Object Detection via Enhanced Open-Set Object Detector
Fu, Y.; Wang, Y.; Pan, Y.; Huai, L.; Qiu, X.; Shangguan, Z.; Liu, T.; Kong, L.; Fu, Y.; Van Gool, L.; et al. 2024 · 2024
Closest in time.
Skysense: A multi-modal remote sensing foundation model towards universal interpretation for earth observation imagery
Guo, X.; Lao, J.; Dang, B.; Zhang, Y.; Yu, L.; Ru, L.; Zhong, L.; Huang, Z.; Wu, K.; Hu, D.; et al. 2024 · 2024
Closest in time.
Segment anything model for medical images?
Huang, Y.; Yang, X.; Liu, L.; Zhou, H.; Chang, A.; Zhou, X.; Chen, R.; Yu, J.; Chen, J.; Chen, C.; Liu, S.; Chi, H.; Hu, X.; Yue, K.; Li, L.; Grau, V.; Fan, D.-P.; Dong, F.; and Ni, D. 2024 · 2024
Closest in time.
T-Rex2: Towards Generic Object Detection via Text-Visual Prompt Synergy
Jiang, Q.; Li, F.; Zeng, Z.; Ren, T.; Liu, S.; and Zhang, L. 2024 · 2024
Closest in time.
Geochat: Grounded large vision-language model for remote sensing
Kuckreja, K.; Danish, M. S.; Naseer, M.; Das, A.; Khan, S.; and Khan, F. S. 2024 · 2024
Closest in time.
Direction-Oriented Visual–Semantic Embedding Model for Remote Sensing Image–Text Retrieval
Ma, Q.; Pan, J.; and Bai, C. 2024 · 2024
Closest in time.
Remote Sensing Vision-Language Foundation Models without Annotations via Ground Remote Alignment
Mall, U.; Phoo, C. P.; Liu, M. K.; Vondrick, C.; Hariharan, B.; and Bala, K. 2024 · 2024
Closest in time.
PIR: Remote Sensing Image-Text Retrieval with Prior Instruction Representation Learning
Pan, J.; Ma, M.; Ma, Q.; Bai, C.; and Chen, S. 2024 · 2024
Closest in time.
Aligning and prompting everything all at once for universal visual perception
Shen, Y.; Fu, C.; Chen, P.; Zhang, M.; Li, K.; Sun, X.; Wu, Y.; Lin, S.; and Ji, R. 2024 · 2024
Closest in time.
VideoGrounding-DINO: Towards Open-Vocabulary Spatio-Temporal Video Grounding
Wasim, S. T.; Naseer, M.; Khan, S.; Yang, M.-H.; and Khan, F. S. 2024 · 2024
Closest in time.
Multi-modal queried object detection in the wild
Xu, Y.; Zhang, M.; Fu, C.; Chen, P.; Yang, X.; Li, K.; and Xu, C. 2024 · 2024
Closest in time.
Good at captioning, bad at counting: Benchmarking GPT-4V on Earth observation data
Zhang, C.; and Wang, S. 2024 · 2024
Closest in time.
RS5M and GeoRSCLIP: A large scale vision-language dataset and a large vision-language model for remote sensing
Zhang, Z.; Zhao, T.; Guo, Y.; and Yin, J. 2024 · 2024
Closest in time.