Fetching the paper…
Reading the bibliography…
Visual question answering (VQA) for remote sensing scene has great potential in intelligent human-computer interaction system.
“Curriculum learning,”
Y. Bengio, Jérme Louradour, R. Collobert, and J. Weston, · 2009
Earlier work this paper cites.
“Self-paced learning for latent variable models,”
M Pawan Kumar, Benjamin Packer, and Daphne Koller, · 2010
Earlier work this paper cites.
“Bag-of-visual-words and spatial extensions for land-use classification,”
Yi Yang and Shawn Newsam, · 2010
Earlier work this paper cites.
“A multi-world approach to question answering about real-world scenes based on uncertain input,”
Mateusz Malinowski and Mario Fritz, · 2014
Earlier work this paper cites.
“Saliency-guided unsupervised feature learning for scene classification,”
Fan Zhang, Bo Du, and Liangpei Zhang, · 2014
Earlier work this paper cites.
“Land use classification in remote sensing images by convolutional neural networks,”
Marco Castelluccio, Giovanni Poggi, Carlo Sansone, and Luisa Verdoliva, · 2015
Earlier work this paper cites.
“VQA: Visual question answering,”
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh, · 2015
Earlier work this paper cites.
“Self-paced curriculum learning,”
Lu Jiang, Deyu Meng, Qian Zhao, Shiguang Shan, and Alexander G. Hauptmann, · 2015
Earlier work this paper cites.
“ABC-CNN: An attention based convolutional neural network for visual question answering,”
Kan Chen, Jiang Wang, Liang-Chieh Chen, Haoyuan Gao, Wei Xu, and Ram Nevatia, · 2015
Earlier work this paper cites.
“Faster R-CNN: Towards real-time object detection with region proposal networks,”
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun, · 2015
Earlier work this paper cites.
“A survey on object detection in optical remote sensing images,”
Gong Cheng and Junwei Han, · 2016
Earlier work this paper cites.
“Where to look: Focus regions for visual question answering,”
Kevin J Shih, Saurabh Singh, and Derek Hoiem, · 2016
Earlier work this paper cites.
“Visual7W: Grounded question answering in images,”
Yuke Zhu, Oliver Groth, Michael Bernstein, and Li Fei-Fei, · 2016
Earlier work this paper cites.
“Easy questions first? A case study on curriculum learning for question answering,”
Mrinmaya Sachan and Eric Xing, · 2016
Earlier work this paper cites.
“Multimodal compact bilinear pooling for visual question answering and visual grounding,”
Akira Fukui, Dong Huk Park, Daylen Yang, Anna Rohrbach, Trevor Darrell, and Marcus Rohrbach, · 2016
Earlier work this paper cites.
“Hadamard product for low-rank bilinear pooling,”
Jin-Hwa Kim, Kyoung-Woon On, Woosang Lim, Jeonghee Kim, Jung-Woo Ha, and Byoung-Tak Zhang, · 2016
Earlier work this paper cites.
“Deep learning in remote sensing: A comprehensive review and list of resources,”
Xiao Xiang Zhu, Devis Tuia, Lichao Mou, Gui-Song Xia, Liangpei Zhang, Feng Xu, and Friedrich Fraundorfer, · 2017
Earlier work this paper cites.
“Automatic road detection and centerline extraction via cascaded end-to-end convolutional neural network,”
Guangliang Cheng, Ying Wang, Shibiao Xu, Hongzhen Wang, Shiming Xiang, and Chunhong Pan, · 2017
Earlier work this paper cites.
“Deep learning applied to NLP,”
Marc Moreno Lopez and Jugal Kalita, · 2017
Earlier work this paper cites.
“Visual question answering: A survey of methods and datasets,”
Qi Wu, Damien Teney, Peng Wang, Chunhua Shen, Anthony Dick, and Anton van den Hengel, · 2017
Cited alongside, same era.
“Multi-modal factorized bilinear pooling with co-attention learning for visual question answering,”
Zhou Yu, Jun Yu, Jianping Fan, and Dacheng Tao, · 2017
Cited alongside, same era.
“MUTAN: Multimodal tucker fusion for visual question answering,”
Hedi Ben-Younes, Rémi Cadene, Matthieu Cord, and Nicolas Thome, · 2017
Cited alongside, same era.
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin, · 2017
Cited alongside, same era.
“Learning multi-attention convolutional neural network for fine-grained image recognition,”
Heliang Zheng, Jianlong Fu, Tao Mei, and Jiebo Luo, · 2017
Cited alongside, same era.
“Land-use land-cover classification by machine learning classifiers for satellite observations—a review,”
Swapan Talukdar, Pankaj Singha, Susanta Mahato, Swades Pal, Yuei-An Liou, Atiqur Rahman, et al., · 2020
Later among the works it cites.
“Object detection in optical remote sensing images: A survey and a new benchmark,”
Ke Li, Gang Wan, Gong Cheng, Liqiu Meng, and Junwei Han, · 2020
Later among the works it cites.
“Say as you wish: Fine-grained control of image caption generation with abstract scene graphs,”
Shizhe Chen, Qin Jin, Peng Wang, and Qi Wu, · 2020
Later among the works it cites.
“Multiple interaction learning with question-type prior knowledge for constraining answer search space in visual question answering,”
Tuong Do, Binh X Nguyen, Huy Tran, Erman Tjiputra, Quang D Tran, and Thanh-Toan Do, · 2020
Later among the works it cites.
“Multimodal feature fusion by relational reasoning and attention for visual question answering,”
Weifeng Zhang, Jing Yu, Hua Hu, Haiyang Hu, and Zengchang Qin, · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Aid: A benchmark data set for performance evaluation of aerial scene classification,”
Gui-Song Xia, Jingwen Hu, Fan Hu, Baoguang Shi, Xiang Bai, Yanfei Zhong, Liangpei Zhang, and Xiaoqiang Lu, · 2017
Cited alongside, same era.
“Robust pcanet for hyperspectral image change detection,”
Zhenghang Yuan, Qi Wang, and Xuelong Li, · 2018
Cited alongside, same era.
“GETNET: A general end-to-end 2-D CNN framework for hyperspectral image change detection,”
Qi Wang, Zhenghang Yuan, Qian Du, and Xuelong Li, · 2018
Cited alongside, same era.
“Learning spectral-spatial-temporal features via a recurrent convolutional neural network for change detection in multispectral imagery,”
Lichao Mou, Lorenzo Bruzzone, and Xiao Xiang Zhu, · 2018
Cited alongside, same era.
“Learning to count objects in natural images for visual question answering,”
Yan Zhang, Jonathon Hare, and Adam Prügel-Bennett, · 2018
Cited alongside, same era.
“Beyond bilinear: Generalized multimodal factorized high-order pooling for visual question answering,”
Zhou Yu, Jun Yu, Chenchao Xiang, Jianping Fan, and Dacheng Tao, · 2018
Cited alongside, same era.
“Tips and tricks for visual question answering: Learnings from the 2017 challenge,”
Damien Teney, Peter Anderson, Xiaodong He, and Anton Van Den Hengel, · 2018
Cited alongside, same era.
Later among the works it cites.
“RSVQA: Visual question answering for remote sensing data,”
Sylvain Lobry, Diego Marcos, Jesse Murray, and Devis Tuia, · 2020
Later among the works it cites.
“Unified vision-language pre-training for image captioning and VQA,”
Luowei Zhou, Hamid Palangi, Lei Zhang, Houdong Hu, Jason Corso, and Jianfeng Gao, · 2020
Later among the works it cites.
“Answer again: Improving VQA with cascaded-answering model,”
Liang Peng, Yang Yang, Xiaopeng Zhang, Yanli Ji, Huimin Lu, and Heng Tao Shen, · 2020
Later among the works it cites.
“Object-centric diagnosis of visual reasoning,”
Jianwei Yang, Jiayuan Mao, Jiajun Wu, Devi Parikh, David D Cox, Joshua B Tenenbaum, and Chuang Gan, · 2020
Later among the works it cites.
“MSN: Modality separation networks for RGB-D scene recognition,”
Zhitong Xiong, Yuan Yuan, and Qi Wang, · 2020
Later among the works it cites.
“In defense of grid features for visual question answering,”
Huaizu Jiang, Ishan Misra, Marcus Rohrbach, Erik Learned-Miller, and Xinlei Chen, · 2020
Later among the works it cites.
“A competence-aware curriculum for visual concepts learning via question answering,”
Qing Li, Siyuan Huang, Yining Hong, and Song-Chun Zhu, · 2020
Later among the works it cites.
“Self-paced curriculum learning for visual question answering on remote sensing data,”
Zhenghang Yuan, Lichao Mou, and Xiao Xiang Zhu, · 2021
Later among the works it cites.
“ASK: Adaptively selecting key local features for RGB-D scene recognition,”
Zhitong Xiong, Yuan Yuan, and Qi Wang, · 2021
Later among the works it cites.
“Mutual attention inception network for remote sensing visual question answering,”
Xiangtao Zheng, Binqiang Wang, Xingqian Du, and Xiaoqiang Lu, · 2021
Later among the works it cites.
“Spatial transformer networks,”
Max Jaderberg, Karen Simonyan, Andrew Zisserman, et al., · 2025
Closest in time.
“Show, attend and tell: Neural image caption generation with visual attention,”
Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron Courville, Ruslan Salakhudinov, Rich Zemel, and Yoshua Bengio, · 2057
Closest in time.
“Multi-modal remote sensing image matching method based on deep learning technology,”
Hao Han, Canhai Li, and Xiaofeng Qiu, · 2083
Closest in time.