Fetching the paper…
Reading the bibliography…
Remote sensing visual question answering (RSVQA) opens new opportunities for the use of overhead imagery by the general public, by enabling human-machine interaction with natural language.
RoBERTa: A Robustly Optimized BERT Pretraining Approach
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., Stoyanov, V., 2019 · 1907
Earlier work this paper cites.
Sanh, V., Debut, L., Chaumond, J., Wolf, T., 2020 · 1910
Earlier work this paper cites.
HuggingFace’s Transformers: State-of-the-art Natural Language Processing
Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., Davison, J., Shleifer, S., von Platen, P., Ma, C., Jernite, Y., Plu, J., Xu, C., Scao, T.L., Gugger, S., Drame, M., Lhoest, Q., Rush, A.M., 2020 · 1910
Earlier work this paper cites.
ImageNet: A large-scale hierarchical image database, in: 2009 IEEE Conference on Computer Vision and Pattern Recognition, IEEE, Miami, FL. pp. 248–255
Deng, J., Dong, W., Socher, R., Li, L.J., Kai Li, Li Fei-Fei, 2009 · 2009
Earlier work this paper cites.
Bag-of-visual-words and spatial extensions for land-use classification, in: Proceedings of the 18th SIGSPATIAL International Conference on Advances in Geographic Information Systems - GIS ’10, ACM Press, San Jose, California. p. 270
Yang, Y., Newsam, S., 2010 · 2010
Earlier work this paper cites.
Im2Text: Describing Images Using 1 Million Captioned Photographs, in: Shawe-Taylor, J., Zemel, R., Bartlett, P., Pereira, F., Weinberger, K.Q. (Eds.), Advances in Neural Information Processing Systems, Curran Associates, Inc
Ordonez, V., Kulkarni, G., Berg, T., 2011 · 2011
Earlier work this paper cites.
FloodNet: A High Resolution Aerial Imagery Dataset for Post Flood Scene Understanding
Rahnemoonfar, M., Chowdhury, T., Sarkar, A., Varshney, D., Yari, M., Murphy, R., 2020 · 2012
Earlier work this paper cites.
Efficient Estimation of Word Representations in Vector Space
Mikolov, T., Chen, K., Corrado, G., Dean, J., 2013 · 2013
Earlier work this paper cites.
Learning Phrase Representations using RNN Encoder–Decoder for Statistical Machine Translation, in: Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), Association for Computational Linguistics, Doha, Qatar. pp. 1724–1734
Cho, K., van Merrienboer, B., Gulcehre, C., Bahdanau, D., Bougares, F., Schwenk, H., Bengio, Y., 2014 · 2014
Earlier work this paper cites.
Microsoft COCO: Common Objects in Context, in: Fleet, D., Pajdla, T., Schiele, B., Tuytelaars, T. (Eds.), Computer Vision – ECCV 2014. Springer International Publishing, Cham. volume 8693, pp. 740–755
Lin, T.Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., Zitnick, C.L., 2014 · 2014
Earlier work this paper cites.
Saliency-Guided Unsupervised Feature Learning for Scene Classification
Zhang, F., Du, B., Zhang, L., 2015 · 2014
Earlier work this paper cites.
VQA: Visual Question Answering, pp. 2425–2433
Antol, S., Agrawal, A., Lu, J., Mitchell, M., Batra, D., Zitnick, C.L., Parikh, D., 2015 · 2015
Earlier work this paper cites.
Distilling the Knowledge in a Neural Network, in: The proceedings of Deep Learning Worshop, arXiv
Hinton, G., Vinyals, O., Dean, J., 2015 · 2015
Earlier work this paper cites.
Kingma, D.P., Ba, J., 2015 · 2015
Earlier work this paper cites.
Skip-Thought Vectors, in: arXiv:1506.06726 [cs]
Kiros, R., Zhu, Y., Salakhutdinov, R., Zemel, R.S., Torralba, A., Urtasun, R., Fidler, S., 2015 · 2015
Earlier work this paper cites.
Neural Module Networks, in: 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), IEEE, Las Vegas, NV, USA. pp. 39–48
Andreas, J., Rohrbach, M., Darrell, T., Klein, D., 2016 · 2016
Earlier work this paper cites.
Deep Residual Learning for Image Recognition, pp. 770–778
He, K., Zhang, X., Ren, S., Sun, J., 2016 · 2016
Earlier work this paper cites.
Google’s Neural Machine Translation System: Bridging the Gap between Human and Machine Translation
Wu, Y., Schuster, M., Chen, Z., Le, Q.V., Norouzi, M., Macherey, W., Krikun, M., Cao, Y., Gao, Q., Macherey, K., Klingner, J., Shah, A., Johnson, M., Liu, X., Kaiser, L., Gouws, S., Kato, Y., Kudo, T., Kazawa, H., Stevens, K., Kurian, G., Patil, N., Wang, W., Young, C., Smith, J., Riesa, J., Rudnick, A., Vinyals, O., Corrado, G., Hughes, M., Dean, J., 2016 · 2016
Earlier work this paper cites.
Stacked Attention Networks for Image Question Answering, in: 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), IEEE, Las Vegas, NV, USA. pp. 21–29
Yang, Z., He, X., Gao, J., Deng, L., Smola, A., 2016 · 2016
Earlier work this paper cites.
MUTAN: Multimodal Tucker Fusion for Visual Question Answering, in: 2017 IEEE International Conference on Computer Vision (ICCV), IEEE, Venice. pp. 2631–2639
Ben-younes, H., Cadene, R., Cord, M., Thome, N., 2017 · 2017
Earlier work this paper cites.
Making the v in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering, pp. 6904–6913
Goyal, Y., Khot, T., Summers-Stay, D., Batra, D., Parikh, D., 2017 · 2017
Earlier work this paper cites.
CLEVR: A Diagnostic Dataset for Compositional Language and Elementary Visual Reasoning, in: 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), IEEE, Honolulu, HI. pp. 1988–1997
Johnson, J., Hariharan, B., van der Maaten, L., Fei-Fei, L., Zitnick, C.L., Girshick, R., 2017 · 2017
Earlier work this paper cites.
Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations
Krishna, R., Zhu, Y., Groth, O., Johnson, J., Hata, K., Kravitz, J., Chen, S., Kalantidis, Y., Li, L.J., Shamma, D.A., Bernstein, M.S., Fei-Fei, L., 2017 · 2017
Cited alongside, same era.
Graph-Structured Representations for Visual Question Answering, pp. 1–9
Teney, D., Liu, L., van den Hengel, A., 2017 · 2017
Cited alongside, same era.
Attention is All you Need, in: Guyon, I., Luxburg, U.V., Bengio, S., Wallach, H., Fergus, R., Vishwanathan, S., Garnett, R. (Eds.), Advances in Neural Information Processing Systems, Curran Associates, Inc
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, L., Polosukhin, I., 2017 · 2017
Cited alongside, same era.
Image Captioning and Visual Question Answering Based on Attributes and External Knowledge
Wu, Q., Shen, C., Wang, P., Dick, A., Hengel, A.v.d., 2018 · 2017
Cited alongside, same era.
AID: A Benchmark Data Set for Performance Evaluation of Aerial Scene Classification
Xia, G.S., Hu, J., Hu, F., Shi, B., Bai, X., Zhong, Y., Zhang, L., Lu, X., 2017 · 2017
Cross-Modal Visual Question Answering for Remote Sensing Data, in: 2021 Digital Image Computing: Techniques and Applications (DICTA), IEEE, Gold Coast, Australia. pp. 1–9
Felix, R., Repasky, B., Hodge, S., Zolfaghari, R., Abbasnejad, E., Sherrah, J., 2021 · 2021
Later among the works it cites.
Roses are Red, Violets are Blue… But Should VQA expect Them To?, in: 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), IEEE, Nashville, TN, USA. pp. 2775–2784
Kervadec, C., Antipov, G., Baccouche, M., Wolf, C., 2021 · 2021
Later among the works it cites.
RSVQA meets BigEarthNet: a new, large-scale, visual question answering dataset for remote sensing, IEEE
Lobry, S., Demir, B., Tuia, D., 2021 · 2021
Later among the works it cites.
Counterfactual VQA: A Cause-Effect Look at Language Bias, in: 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), IEEE, Nashville, TN, USA. pp. 12695–12705
Niu, Y., Tang, K., Zhang, H., Lu, Z., Hua, X.S., Wen, J.R., 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Don’t Just Assume; Look and Answer: Overcoming Priors for Visual Question Answering, in: 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, IEEE, Salt Lake City, UT. pp. 4971–4980
Agrawal, A., Batra, D., Parikh, D., Kembhavi, A., 2018 · 2018
Cited alongside, same era.
Bottom-Up and Top-Down Attention for Image Captioning and Visual Question Answering, in: 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, IEEE, Salt Lake City, UT. pp. 6077–6086
Anderson, P., He, X., Buehler, C., Teney, D., Johnson, M., Gould, S., Zhang, L., 2018 · 2018
Cited alongside, same era.
Challenges and Solutions for Utilizing Earth Observations in the "Big Data" era URL: https://zenodo.org/record/2391936 , doi: 10.5281/ZENODO.2391936
Lachezar, F., Lyubka, P., Vasil, K., Stuart, F., 2018 · 2018
Cited alongside, same era.
Conceptual Captions: A Cleaned, Hypernymed, Image Alt-text Dataset For Automatic Image Captioning, in: Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Association for Computational Linguistics, Melbourne, Australia. pp. 2556–2565
Sharma, P., Ding, N., Goodman, S., Soricut, R., 2018 · 2018
Cited alongside, same era.
DOTA: A Large-Scale Dataset for Object Detection in Aerial Images, in: 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, IEEE, Salt Lake City, UT. pp. 3974–3983
Xia, G.S., Bai, X., Ding, J., Zhu, Z., Belongie, S., Luo, J., Datcu, M., Pelillo, M., Zhang, L., 2018 · 2018
Cited alongside, same era.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding, in: Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), Association for Computational Linguistics, Minneapolis, Minnesota. pp. 4171–4186
Devlin, J., Chang, M.W., Lee, K., Toutanova, K., 2019 · 2019
Cited alongside, same era.
GQA: A New Dataset for Real-World Visual Reasoning and Compositional Question Answering, pp. 6700–6709
Hudson, D.A., Manning, C.D., 2019 · 2019
Cited alongside, same era.
Multimodal Few-Shot Learning with Frozen Language Models
Tsimpoukelli, M., Menick, J., Cabi, S., Eslami, S.M.A., Vinyals, O., Hill, F., 2021 · 2021
Later among the works it cites.
UFO: A UniFied TransfOrmer for Vision-Language Representation Learning
Wang, J., Hu, X., Gan, Z., Yang, Z., Dai, X., Liu, Z., Lu, Y., Wang, L., 2021 · 2021
Later among the works it cites.
Florence: A New Foundation Model for Computer Vision
Yuan, L., Chen, D., Chen, Y.L., Codella, N., Dai, X., Gao, J., Hu, H., Huang, X., Li, B., Li, C., Liu, C., Liu, M., Liu, Z., Lu, Y., Shi, Y., Wang, L., Wang, J., Xiao, B., Xiao, Z., Yang, J., Zeng, M., Zhou, L., Zhang, P., 2021 · 2021
Later among the works it cites.
MERLOT: Multimodal Neural Script Knowledge Models
Zellers, R., Lu, X., Hessel, J., Yu, Y., Park, J.S., Cao, J., Farhadi, A., Choi, Y., 2021 · 2021
Later among the works it cites.
Mutual Attention Inception Network for Remote Sensing Visual Question Answering
Zheng, X., Wang, B., Du, X., Lu, X., 2021 · 2021
Later among the works it cites.
Bi-Modal Transformer-Based Approach for Visual Question Answering in Remote Sensing Imagery
Bazi, Y., Rahhal, M.M.A., Mekhalfi, M.L., Zuair, M.A.A., Melgani, F., 2022 · 2022
Later among the works it cites.
Commonsense Knowledge Reasoning and Generation with Pre-trained Language Models: A Survey
Bhargava, P., Ng, V., 2022 · 2022
Later among the works it cites.
Image-text Retrieval: A Survey on Recent Research and Development, in: Proceedings of the Thirty-First International Joint Conference on Artificial Intelligence, International Joint Conferences on Artificial Intelligence Organization, Vienna, Austria. pp. 5410–5417
Cao, M., Li, S., Li, J., Nie, L., Zhang, M., 2022 · 2022
Later among the works it cites.
Language Transformers for Remote Sensing Visual Question Answering, in: IGARSS 2022 - 2022 IEEE International Geoscience and Remote Sensing Symposium, IEEE, Kuala Lumpur, Malaysia. pp. 4855–4858
Chappuis, C., Mendez, V., Walt, E., Lobry, S., Le Saux, B., Tuia, D., 2022a · 2022
Later among the works it cites.
Embedding Spatial Relations in Visual Question Answering for Remote Sensing, in: 2022 26th International Conference on Pattern Recognition (ICPR), IEEE, Montreal, QC, Canada. pp. 310–316
Faure, M., Lobry, S., Kurtz, C., Wendling, L., 2022 · 2022
Later among the works it cites.
SwapMix: Diagnosing and Regularizing the Over-Reliance on Visual Context in Visual Question Answering, in: 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), IEEE, New Orleans, LA, USA. pp. 5068–5078
Gupta, V., Li, Z., Kortylewski, A., Zhang, C., Li, Y., Yuille, A., 2022 · 2022
Later among the works it cites.
Jarrahi, M.H., Memariani, A., Guha, S., 2022 · 2022
Later among the works it cites.
Multi-modal fusion transformer for visual question answering in remote sensing, in: Pierdicca, N., Bruzzone, L., Bovolo, F. (Eds.), Image and Signal Processing for Remote Sensing XXVIII, SPIE, Berlin, Germany. p. 21
Siebert, T., Clasen, K.N., Ravanbakhsh, M., Demir, B., 2022 · 2022
Later among the works it cites.
CoCa: Contrastive Captioners are Image-Text Foundation Models
Yu, J., Wang, Z., Vasudevan, V., Yeung, L., Seyedhosseini, M., Wu, Y., 2022 · 2022
Later among the works it cites.
Yuan, Z., Mou, L., Wang, Q., Zhu, X.X., 2022a · 2022
Later among the works it cites.
Change Detection Meets Visual Question Answering
Yuan, Z., Mou, L., Xiong, Z., Zhu, X.X., 2022b · 2022
Later among the works it cites.
Multi-task prompt-rsvqa to explicitly count objects on aerial images, in: Workshop on Machine Vision for Earth Observation (MVEO) at the 34th British Machine Vision Conference (BMVC)
Chappuis, C., Sertic, C., Santacroce, N., Castillo Navarro, J., Lobry, S., Le Saux, B., Tuia, D., 2023 · 2023
Closest in time.