Fetching the paper…
Reading the bibliography…
Visual question answering (VQA) is a task of answering a visual question that is a pair of question and image.
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C.L. Zitnick, and D. Parikh, “VQA: Visual Question Answering,” International Conference on Computer Vision (ICCV), 2015
2015
Earlier work this paper cites.
N. Mostafazadeh, I. Misra, J. Devlin, M. Mitchell, X. He, and L. Vanderwende, “Generating natural questions about an image,” Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Berlin, Germany, pp.1802–1813, Association for Computational Linguistics, Aug. 2016
2016
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), June 2016
2016
Earlier work this paper cites.
Language in Vision
Q. Wu, D. Teney, P. Wang, C. Shen, A. Dick, and A. van den Hengel, “Visual question answering: A survey of methods and datasets,” Computer Vision and Image Understanding, vol.163, pp.21 – 40, 2017 · 2017
Earlier work this paper cites.
A. Das, S. Kottur, K. Gupta, A. Singh, D. Yadav, J.M. Moura, D. Parikh, and D. Batra, “Visual Dialog,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017
2017
Earlier work this paper cites.
Y. Goyal, T. Khot, D. Summers-Stay, D. Batra, and D. Parikh, “Making the V in VQA matter: Elevating the role of image understanding in Visual Question Answering,” Conference on Computer Vision and Pattern Recognition (CVPR), 2017
2017
Earlier work this paper cites.
S. Zhang, L. Qu, S. You, Z. Yang, and J. Zhang, “Automatic generation of grounded visual questions,” Proceedings of the Twenty-Sixth International Joint Conference on Artificial Intelligence, IJCAI 2017, Melbourne, Australia, August 19-25, 2017, pp.4235–4243, 2017
2017
Earlier work this paper cites.
H. Ben-younes, R. Cadene, M. Cord, and N. Thome, “Mutan: Multimodal tucker fusion for visual question answering,” The IEEE International Conference on Computer Vision (ICCV), Oct 2017
2017
Earlier work this paper cites.
E. Jang, S. Gu, and B. Poole, “Categorical reparameterization with gumbel-softmax,” 5th International Conference on Learning Representations, ICLR 2017, Toulon, France, April 24-26, 2017, Conference Track Proceedings, 2017
2017
Cited alongside, same era.
2017
Cited alongside, same era.
D. Gurari, Q. Li, A.J. Stangl, A. Guo, C. Lin, K. Grauman, J. Luo, and J.P. Bigham, “Vizwiz grand challenge: Answering visual questions from blind people,” The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), June 2018
2018
Cited alongside, same era.
A. Das, S. Datta, G. Gkioxari, S. Lee, D. Parikh, and D. Batra, “Embodied Question Answering,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018
2018
Cited alongside, same era.
A. Singh, V. Natarajan, Y. Jiang, X. Chen, M. Shah, M. Rohrbach, D. Batra, and D. Parikh, “Pythia-a platform for vision & language research,” SysML Workshop, NeurIPS, 2018
2018
Later among the works it cites.
D. Teney, P. Anderson, X. He, and A. van den Hengel, “Tips and tricks for visual question answering: Learnings from the 2017 challenge,” The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), June 2018
2018
Later among the works it cites.
N. Bhattacharya, Q. Li, and D. Gurari, “Why does a visual question have different answers?,” The IEEE International Conference on Computer Vision (ICCV), October 2019
2019
Later among the works it cites.
M. Shah, X. Chen, M. Rohrbach, and D. Parikh, “Cycle-consistency for robust visual question answering,” The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), June 2019
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y. Li, N. Duan, B. Zhou, X. Chu, W. Ouyang, X. Wang, and M. Zhou, “Visual question generation as dual task of visual question answering,” The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), June 2018
2018
Cited alongside, same era.
Z. Fan, Z. Wei, P. Li, Y. Lan, and X. Huang, “A question type driven framework to diversify visual question generation,” Proceedings of the Twenty-Seventh International Joint Conference on Artificial Intelligence, IJCAI 2018, July 13-19, 2018, Stockholm, Sweden, pp.4048–4054, 2018
2018
Cited alongside, same era.
P. Anderson, X. He, C. Buehler, D. Teney, M. Johnson, S. Gould, and L. Zhang, “Bottom-up and top-down attention for image captioning and visual question answering,” The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), June 2018
2018
Cited alongside, same era.
2018
Cited alongside, same era.
R. Krishna, M. Bernstein, and L. Fei-Fei, “Information maximizing visual question generation,” IEEE Conference on Computer Vision and Pattern Recognition, 2019
2019
Later among the works it cites.
L. Liu, J. Tang, X. Wan, and Z. Guo, “Generating diverse and descriptive image captions using visual paraphrases,” The IEEE International Conference on Computer Vision (ICCV), October 2019
2019
Later among the works it cites.
2019
Later among the works it cites.
Z. Yu, J. Yu, Y. Cui, D. Tao, and Q. Tian, “Deep modular co-attention networks for visual question answering,” The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), June 2019
2019
Later among the works it cites.