Fetching the paper…
Reading the bibliography…
Recently, studies of visual question answering have explored various architectures of end-to-end networks and achieved promising results on both natural and synthetic datasets, which require explicitly compositional reasoning.
P. F. Felzenszwalb, R. B. Girshick, D. McAllester, and D. Ramanan, “Object detection with discriminatively trained part-based models,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 32, no. 9, pp. 1627–1645, Sept 2010
2010
Earlier work this paper cites.
D. Chen and C. D. Manning, “A fast and accurate dependency parser using neural networks,” in EMNLP , 2014
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
M. Malinowski, M. Rohrbach, and M. Fritz, “Ask your neurons: A neural-based approach to answering questions about images,” in ICCV , 2015, pp. 1–9
2015
Earlier work this paper cites.
B. M. Lake, R. Salakhutdinov, and J. B. Tenenbaum, “Human-level concept learning through probabilistic program induction,” Science , vol. 350, no. 6266, pp. 1332–1338, 2015
2015
Earlier work this paper cites.
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. L. Zitnick, and D. Parikh, “VQA: Visual Question Answering,” in ICCV , 2015
2015
Earlier work this paper cites.
J. Lu, J. Yang, D. Batra, and D. Parikh, “Hierarchical question-image co-attention for visual question answering,” in NIPS , 2016, pp. 289–297
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
A. Jabri, A. Joulin, and L. van der Maaten, “Revisiting visual question answering baselines,” in ECCV , ser. Lecture Notes in Computer Science, vol. 9912. Springer, 2016, pp. 727–739
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
H. Xu and K. Saenko, Ask, Attend and Answer: Exploring Question-Guided Spatial Attention for Visual Question Answering . Cham: Springer International Publishing, 2016, pp. 451–466
2016
Earlier work this paper cites.
K. J. Shih, S. Singh, and D. Hoiem, “Where to look: Focus regions for visual question answering,” in CVPR , 2016
2016
Earlier work this paper cites.
Y. Zhu, O. Groth, M. S. Bernstein, and L. Fei-Fei, “Visual7w: Grounded question answering in images,” in CVPR , 2016, pp. 4995–5004
2016
Earlier work this paper cites.
Z. Yang, X. He, J. Gao, L. Deng, and A. Smola, “Stacked attention networks for image question answering,” in CVPR , June 2016, pp. 21–29
2016
Earlier work this paper cites.
A. Fukui, D. H. Park, D. Yang, A. Rohrbach, T. Darrell, and M. Rohrbach, “Multimodal compact bilinear pooling for visual question answering and visual grounding,” in EMNLP , 2016, pp. 457–468
2016
Earlier work this paper cites.
A. Kumar, O. Irsoy, P. Ondruska, M. Iyyer, J. Bradbury, I. Gulrajani, V. Zhong, R. Paulus, and R. Socher, “Ask me anything: Dynamic memory networks for natural language processing,” in ICML , 2016, pp. 1378–1387
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
J. Andreas, M. Rohrbach, T. Darrell, and D. Klein, “Neural module networks,” in CVPR . IEEE Computer Society, 2016, pp. 39–48
2016
Cited alongside, same era.
——, “Learning to compose neural networks for question answering,” in HLT-NAACL . The Association for Computational Linguistics, 2016, pp. 1545–1554
2016
Cited alongside, same era.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in CVPR . IEEE Computer Society, 2016, pp. 770–778
2016
Cited alongside, same era.
Y. Goyal, T. Khot, D. Summers-Stay, D. Batra, and D. Parikh, “Making the V in VQA matter: Elevating the role of image understanding in Visual Question Answering,” in CVPR , 2017
2017
Cited alongside, same era.
J. Johnson, B. Hariharan, L. van der Maaten, L. Fei-Fei, C. L. Zitnick, and R. Girshick, “CLEVR: A diagnostic dataset for compositional language and elementary visual reasoning,” in CVPR , 2017
E. Perez, F. Strub, H. de Vries, V. Dumoulin, and A. C. Courville, “Film: Visual reasoning with a general conditioning layer,” in AAAI . AAAI Press, 2018
2018
Later among the works it cites.
G. E. Hinton, S. Sabour, and N. Frosst, “Matrix capsules with EM routing,” in ICLR , 2018
2018
Later among the works it cites.
A. Agrawal, D. Batra, D. Parikh, and A. Kembhavi, “Don’t just assume; look and answer: Overcoming priors for visual question answering,” in CVPR , June 2018
2018
Later among the works it cites.
D. Gordon, A. Kembhavi, M. Rastegari, J. Redmon, D. Fox, and A. Farhadi, “Iqa: Visual question answering in interactive environments,” in CVPR , June 2018
2018
Later among the works it cites.
K. Kafle, B. Price, S. Cohen, and C. Kanan, “Dvqa: Understanding data visualizations via question answering,” in CVPR , June 2018
2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2017
Cited alongside, same era.
D. Teney, L. Liu, and A. van den Hengel, “Graph-structured representations for visual question answering,” in CVPR , July 2017
2017
Cited alongside, same era.
C. Zhu, Y. Zhao, S. Huang, K. Tu, and Y. Ma, “Structured attentions for visual question answering,” in ICCV , Oct 2017
2017
Cited alongside, same era.
R. Hu, J. Andreas, M. Rohrbach, T. Darrell, and K. Saenko, “Learning to reason: End-to-end module networks for visual question answering,” in ICCV , 2017
2017
Cited alongside, same era.
J. Johnson, B. Hariharan, L. van der Maaten, J. Hoffman, F. Li, C. L. Zitnick, and R. B. Girshick, “Inferring and executing programs for visual reasoning,” in ICCV , 2017
2017
Cited alongside, same era.
S. Sabour, N. Frosst, and G. E. Hinton, “Dynamic routing between capsules,” in NIPS , I. Guyon, U. V. Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett, Eds. Curran Associates, Inc., 2017, pp. 3856–3866
2017
Cited alongside, same era.
Z. Yu, J. Yu, J. Fan, and D. Tao, “Multi-modal factorized bilinear pooling with co-attention learning for visual question answering,” ICCV , pp. 1839–1848, 2017
2017
Cited alongside, same era.
H. Ben-younes, R. Cadene, M. Cord, and N. Thome, “Mutan: Multimodal tucker fusion for visual question answering,” in ICCV , Oct 2017
2017
Cited alongside, same era.
M. Narasimhan and A. G. Schwing, “Straight to the facts: Learning knowledge base retrieval for factual visual question answering,” in ECCV , September 2018
2018
Later among the works it cites.
I. Misra, R. Girshick, R. Fergus, M. Hebert, A. Gupta, and L. van der Maaten, “Learning by asking questions,” in CVPR , June 2018
2018
Later among the works it cites.
Y. Li, N. Duan, B. Zhou, X. Chu, W. Ouyang, X. Wang, and M. Zhou, “Visual question generation as dual task of visual question answering,” in CVPR , June 2018
2018
Later among the works it cites.
D. Teney and A. van den Hengel, “Visual question answering as a meta learning task,” in ECCV , September 2018
2018
Later among the works it cites.
W. Norcliffe-Brown, E. Vafeias, and S. Parisot, “Learning conditioned graph structures for interpretable visual question answering,” in NIPS , 2018
2018
Later among the works it cites.
M. Narasimhan, S. Lazebnik, and A. G. Schwing, “Out of the box: Reasoning with graph convolution nets for factual visual question answering,” in NIPS , 2018
2018
Later among the works it cites.
K. Yi, J. Wu, C. Gan, A. Torralba, P. Kohli, and J. B. Tenenbaum, “Neural-symbolic vqa: Disentangling reasoning from vision and language understanding,” in NIPS , 2018
2018
Later among the works it cites.
D. Mascharka, P. Tran, R. Soklaski, and A. Majumdar, “Transparency by design: Closing the gap between performance and interpretability in visual reasoning,” in CVPR , June 2018
2018
Later among the works it cites.
Q. Cao, X. Liang, B. Li, G. Li, and L. Lin, “Visual question reasoning on general dependency tree,” in CVPR , June 2018
2018
Later among the works it cites.
D. A. Hudson and C. D. Manning, “Compositional attention networks for machine reasoning,” 2018
2018
Later among the works it cites.
D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” in ICLR , 2015
2019
Later among the works it cites.