Fetching the paper…
Reading the bibliography…
In visual question answering (VQA), an algorithm must answer text-based questions about images.
One-shot learning of object categories
L. Fei-Fei, R. Fergus, and P. Perona · 2006
Earlier work this paper cites.
Microsoft COCO: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Earlier work this paper cites.
A multi-world approach to question answering about real-world scenes based on uncertain input
M. Malinowski and M. Fritz · 2014
Earlier work this paper cites.
VQA: Visual question answering
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. L. Zitnick, and D. Parikh · 2015
Earlier work this paper cites.
Are you talking to a machine? Dataset and methods for multilingual image question answering
H. Gao, J. Mao, J. Zhou, Z. Huang, L. Wang, and W. Xu · 2015
Earlier work this paper cites.
Compositional memory for visual question answering
A. Jiang, F. Wang, F. Porikli, and Y. Li · 2015
Earlier work this paper cites.
Skip-thought vectors
R. Kiros, Y. Zhu, R. Salakhutdinov, R. S. Zemel, A. Torralba, R. Urtasun, and S. Fidler · 2015
Earlier work this paper cites.
Ask your neurons: A neural-based approach to answering questions about images
M. Malinowski, M. Rohrbach, and M. Fritz · 2015
Earlier work this paper cites.
Exploring models and data for image question answering
M. Ren, R. Kiros, and R. Zemel · 2015
Earlier work this paper cites.
Simple baseline for visual question answering
B. Zhou, Y. Tian, S. Sukhbaatar, A. Szlam, and R. Fergus · 2015
Earlier work this paper cites.
Deep compositional question answering with neural module networks
J. Andreas, M. Rohrbach, T. Darrell, and D. Klein · 2016
Earlier work this paper cites.
Learning to compose neural networks for question answering
J. Andreas, M. Rohrbach, T. Darrell, and D. Klein · 2016
Earlier work this paper cites.
Multimodal compact bilinear pooling for visual question answering and visual grounding
A. Fukui, D. H. Park, D. Yang, A. Rohrbach, T. Darrell, and M. Rohrbach · 2016
Cited alongside, same era.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Cited alongside, same era.
A focused dynamic attention model for visual question answering
I. Ilievski, S. Yan, and J. Feng · 2016
Cited alongside, same era.
Revisiting visual question answering baselines
A. Jabri, A. Joulin, and L. van der Maaten · 2016
Cited alongside, same era.
Answer-type prediction for visual question answering
K. Kafle and C. Kanan · 2016
Cited alongside, same era.
Multimodal residual learning for visual QA
J.-H. Kim, S.-W. Lee, D.-H. Kwak, M.-O. Heo, J. Kim, J.-W. Ha, and B.-T. Zhang · 2016
Ask me anything: Free-form visual question answering based on knowledge from external sources
Q. Wu, P. Wang, C. Shen, A. van den Hengel, and A. R. Dick · 2016
Later among the works it cites.
Dynamic memory networks for visual and textual question answering
C. Xiong, S. Merity, and R. Socher · 2016
Later among the works it cites.
Ask, attend and answer: Exploring question-guided spatial attention for visual question answering
H. Xu and K. Saenko · 2016
Later among the works it cites.
Stacked attention networks for image question answering
Z. Yang, X. He, J. Gao, L. Deng, and A. J. Smola · 2016
Later among the works it cites.
Yin and yang: Balancing and answering binary visual questions
P. Zhang, Y. Goyal, D. Summers-Stay, D. Batra, and D. Parikh · 2016
Later among the works it cites.
Visual7w: Grounded question answering in images
Y. Zhu, O. Groth, M. Bernstein, and L. Fei-Fei · 2016
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Visual Genome: Connecting language and vision using crowdsourced dense image annotations
R. Krishna, Y. Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y. Kalantidis, L.-J. Li, D. A. Shamma, M. Bernstein, and L. Fei-Fei · 2016
Cited alongside, same era.
Hierarchical question-image co-attention for visual question answering
J. Lu, J. Yang, D. Batra, and D. Parikh · 2016
Cited alongside, same era.
Training recurrent answering units with joint loss minimization for VQA
H. Noh and B. Han · 2016
Cited alongside, same era.
Image question answering using convolutional neural network with dynamic parameter prediction
H. Noh, P. H. Seo, and B. Han · 2016
Cited alongside, same era.
Dualnet: Domain-invariant network for visual question answering
K. Saito, A. Shin, Y. Ushiku, and T. Harada · 2016
Cited alongside, same era.
Where to look: Focus regions for visual question answering
K. J. Shih, S. Singh, and D. Hoiem · 2016
Cited alongside, same era.
Later among the works it cites.
Making the V in VQA matter: Elevating the role of image understanding in Visual Question Answering
Y. Goyal, T. Khot, D. Summers-Stay, D. Batra, and D. Parikh · 2017
Closest in time.
CLEVR: A diagnostic dataset for compositional language and elementary visual reasoning
J. Johnson, B. Hariharan, L. van der Maaten, L. Fei-Fei, C. L. Zitnick, and R. Girshick · 2017
Closest in time.
Visual question answering: Datasets, algorithms, and future challenges
K. Kafle and C. Kanan · 2017
Closest in time.
Data augmentation for visual question answering
K. Kafle, M. Yousefhussien, and C. Kanan · 2017
Closest in time.
Visual question answering: A survey of methods and datasets
Q. Wu, D. Teney, P. Wang, C. Shen, A. Dick, and A. v. d. Hengel · 2017
Closest in time.