Fetching the paper…
Reading the bibliography…
Accurately answering a question about a given image requires combining observations with general knowledge.
Gradient-based learning applied to document recognition
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
Learning context-dependent mappings from sentences to logical form
L. S. Zettlemoyer and M. Collins · 2005
Earlier work this paper cites.
Learning to map sentences to logical form: Structured classification with probabilistic categorial grammars
L. S. Zettlemoyer and M. Collins · 2005
Earlier work this paper cites.
Dbpedia: A nucleus for a web of open data
S. Auer, C. Bizer, G. Kobilarov, J. Lehmann, R. Cyganiak, and Z. Ives · 2007
Earlier work this paper cites.
A survey on question answering technology from an information retrieval perspective
O. Kolomiyets and M.-F. Moens · 2011
Earlier work this paper cites.
Template-based question answering over RDF data
C. Unger, L. Bühmann, J. Lehmann, A.-C. N. Ngomo, D. Gerber, and P. Cimiano · 2012
Earlier work this paper cites.
Semantic Parsing on Freebase from Question-Answer Pairs
J. Berant, A. Chou, R. Frostig, and P. Liang · 2013
Earlier work this paper cites.
Large-scale Semantic Parsing via Schema Matching and Lexicon Extension
Q. Cai and A. Yates · 2013
Earlier work this paper cites.
Jointly learning to parse and perceive: Connecting natural language to the physical world
J. Krishnamurthy and T. Kollar · 2013
Earlier work this paper cites.
Scaling semantic parsers with on-the-fly ontology matching
T. Kwiatkowski, E. Choi, Y. Artzi, and L. Zettlemoyer · 2013
Earlier work this paper cites.
Learning dependency-based compositional semantics
P. Liang, M. I. Jordan, and D. Klein · 2013
Earlier work this paper cites.
Semantic parsing via paraphrasing
J. Berant and P. Liang · 2014
Earlier work this paper cites.
Question answering with sub-graph embeddings
A. Bordes, S. Chopra, and J. Weston · 2014
Earlier work this paper cites.
Open question answering with weakly supervised embedding models
A. Bordes, J. Weston, and N. Usunier · 2014
Earlier work this paper cites.
Open question answering over curated and extracted knowledge bases
A. Fader, L. Zettlemoyer, and O. Etzioni · 2014
Earlier work this paper cites.
A Multi-World Approach to Question Answering about Real-World Scenes based on Uncertain Input
M. Malinowski and M. Fritz · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
J. Pennington, R. Socher, and C. D. Manning · 2014
Earlier work this paper cites.
Webchild: Harvesting and organizing commonsense knowledge from the web
N. Tandon, G. de Melo, F. Suchanek, and G. Weikum · 2014
Earlier work this paper cites.
Information extraction over structured data: Question answering with Freebase
X.Yao and B. V. Durme · 2014
Earlier work this paper cites.
VQA: Visual Question Answering
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. L. Zitnick, and D. Parikh · 2015
Earlier work this paper cites.
Large-scale simple question answering with memory networks
A. Bordes, N. Usunier, S. Chopra, and J. Weston · 2015
Earlier work this paper cites.
Question answering over freebase with multi-column convolutional neural networks
L. Dong, F. Wei, M. Zhou, and K. Xu · 2015
Earlier work this paper cites.
Are you talking to a machine? Dataset and Methods for Multilingual Image Question Answering
H. Gao, J. Mao, J. Zhou, Z. Huang, L. Wang, and W. Xu · 2015
Earlier work this paper cites.
Compositional memory for visual question answering
A. Jiang, F. Wang, F. Porikli, and Y. Li · 2015
Cited alongside, same era.
Ask your neurons: A neural-based approach to answering questions about images
M. Malinowski, M. Rohrbach, and M. Fritz · 2015
Cited alongside, same era.
Exploring models and data for image question answering
M. Ren, R. Kiros, and R. Zemel · 2015
Cited alongside, same era.
Semantic parsing via staged query graph generation: Question answering with knowledge base
S. W. t. Yih, M.-W. Chang, X. He, and J. Gao · 2015
Cited alongside, same era.
Visual madlibs: Fill in the blank image generation and question answering
L. Yu, E. Park, A. Berg, and T. Berg · 2015
Cited alongside, same era.
Yin and yang: Balancing and answering binary visual questions
Sequence-based structured prediction for semantic parsing
C. Xiao, M. Dymetman, and C. Gardent · 2016
Later among the works it cites.
Dynamic memory networks for visual and textual question answering
C. Xiong, S. Merity, and R. Socher · 2016
Later among the works it cites.
Ask, attend and answer: Exploring question-guided spatial attention for visual question answering
H. Xu and K. Saenko · 2016
Later among the works it cites.
Stacked attention networks for image question answering
Z. Yang, X. He, J. Gao, L. Deng, and A. Smola · 2016
Later among the works it cites.
Question answering over knowledge base with neural attention combining global knowledge information
Y. Zhang, K. Liu, S. He, G. Ji, Z. Liu, H. Wu, and J. Zhao · 2016
Later among the works it cites.
Visual7W: Grounded Question Answering in Images
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
P. Zhang, Y. Goyal, D. Summers-Stay, D. Batra, and D. Parikh · 2015
Cited alongside, same era.
Simple baseline for visual question answering
B. Zhou, Y. Tian, S. Sukhbataar, A. Szlam, and R. Fergus · 2015
Cited alongside, same era.
Building a large-scale multimodal Knowledge Base for Visual Question Answering
Y. Zhu, C. Zhang, C. Ré, and L. Fei-Fei · 2015
Cited alongside, same era.
Deep compositional question answering with neural module networks
J. Andreas, M. Rohrbach, T. Darrell, and D. Klein · 2016
Cited alongside, same era.
Human attention in visual question answering: Do humans and deep networks look at the same regions?
A. Das, H. Agrawal, C. L. Zitnick, D. Parikh, and D. Batra · 2016
Cited alongside, same era.
Multimodal compact bilinear pooling for visual question answering and visual grounding
A. Fukui, D. H. Park, D. Yang, A. Rohrbach, T. Darrell, and M. Rohrbach · 2016
Cited alongside, same era.
Revisiting Visual Question Answering Baselines
A. Jabri, A. Joulin, and L. van der Maaten · 2016
Cited alongside, same era.
Y. Zhu, O. Groth, M. Bernstein, and L. Fei-Fei · 2016
Later among the works it cites.
Measuring machine intelligence through visual question answering
C. L. Zitnick, A. Agrawal, S. Antol, M. Mitchell, D. Batra, and D. Parikh · 2016
Later among the works it cites.
Mutan: Multimodal tucker fusion for visual question answering
H. Ben-younes, R. Cadene, M. Cord, and N. Thome · 2017
Later among the works it cites.
Visual Dialog
A. Das, S. Kottur, K. Gupta, A. Singh, D. Yadav, J. M. Moura, D. Parikh, and D. Batra · 2017
Later among the works it cites.
Learning cooperative visual dialog agents with deep reinforcement learning
A. Das, S. Kottur, J. M. Moura, S. Lee, and D. Batra · 2017
Later among the works it cites.
Creativity: Generating Diverse Questions using Variational Autoencoders
U. Jain, Z. Zhang, and A. G. Schwing · 2017
Later among the works it cites.
Clevr: A diagnostic dataset for compositional language and elementary visual reasoning
J. Johnson, B. Hariharan, L. van der Maaten, L. Fei-Fei, C. L. Zitnick, and R. Girshick · 2017
Later among the works it cites.
Visual question generation as dual task of visual question answering
Y. Li, N. Duan, B. Zhou, X. Chu, W. Ouyang, and X. Wang · 2017
Later among the works it cites.
High-Order Attention Models for Visual Question Answering
I. Schwartz, A. G. Schwing, and T. Hazan · 2017
Later among the works it cites.
Conceptnet 5.5: An open multilingual graph of general knowledge
R. Speer, J. Chin, and C. Havasi · 2017
Later among the works it cites.
Explicit knowledge-based reasoning for visual question answering
P. Wang, Q. Wu, C. Shen, A. Dick, and A. Van Den Henge · 2017
Later among the works it cites.
Iqa: Visual question answering in interactive environments
D. Gordon, A. Kembhavi, M. Rastegari, J. Redmon, D. Fox, and A. Farhadi · 2018
Closest in time.
Two can play this Game: Visual Dialog with Discriminative Question Generation and Answering
U. Jain, S. Lazebnik, and A. G. Schwing · 2018
Closest in time.
Straight to the facts: Learning knowledge base retrieval for factual visual question answering
M. Narasimhan and A. G. Schwing · 2018
Closest in time.
Modeling relational data with graph convolutional networks
M. Schlichtkrull, T. N. Kipf, P. Bloem, R. v. d. Berg, I. Titov, and M. Welling · 2018
Closest in time.
Fvqa: Fact-based visual question answering
P. Wang, Q. Wu, C. Shen, A. Dick, and A. v. d. Hengel · 2018
Closest in time.
Zero-shot recognition via semantic embeddings and knowledge graphs
X. Wang, Y. Ye, and A. Gupta · 2018
Closest in time.