Fetching the paper…
Reading the bibliography…
Visual Question Answering (VQA) is a challenging task that has received increasing attention from both the computer vision and the natural language processing communities.
Understanding natural language
T. Winograd · 1972
Earlier work this paper cites.
Verbs semantics and lexical selection
Z. Wu and M. Palmer · 1994
Earlier work this paper cites.
Question answering
D. Jurafsky and J. H. Martin · 2000
Earlier work this paper cites.
Conversational robots: building blocks for grounding word meaning
D. Roy, K.-Y. Hsiao, and N. Mavridis · 2003
Earlier work this paper cites.
ConceptNet - A practical commonsense reasoning toolkit
H. Liu and P. Singh · 2004
Earlier work this paper cites.
DBpedia: A nucleus for a web of open data
S. Auer, C. Bizer, G. Kobilarov, J. Lehmann, R. Cyganiak, and Z. Ives · 2007
Earlier work this paper cites.
Open information extraction for the web
M. Banko, M. J. Cafarella, S. Soderland, M. Broadhead, and O. Etzioni · 2007
Earlier work this paper cites.
Online learning of relaxed ccg grammars for parsing to logical form
L. S. Zettlemoyer and M. Collins · 2007
Earlier work this paper cites.
Freebase: a collaboratively created graph database for structuring human knowledge
K. Bollacker, C. Evans, P. Paritosh, T. Sturge, and J. Taylor · 2008
Earlier work this paper cites.
The stanford typed dependencies representation
M.-C. de Marneffe and C. D. Manning · 2008
Earlier work this paper cites.
SPARQL query language for RDF
E. Prud’Hommeaux, A. Seaborne, et al · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Robust spoken instruction understanding for hri
R. Cantrell, M. Scheutz, P. Schermerhorn, and X. Wu · 2010
Earlier work this paper cites.
Toward an Architecture for Never-Ending Language Learning
A. Carlson, J. Betteridge, B. Kisiel, and B. Settles · 2010
Earlier work this paper cites.
Open Information Extraction: The Second Generation
O. Etzioni, A. Fader, J. Christensen, S. Soderland, and M. Mausam · 2011
Earlier work this paper cites.
Identifying relations for open information extraction
A. Fader, S. Soderland, and O. Etzioni · 2011
Earlier work this paper cites.
A survey on question answering technology from an information retrieval perspective
O. Kolomiyets and M.-F. Moens · 2011
Earlier work this paper cites.
Composing simple image descriptions using web-scale n-grams
S. Li, G. Kulkarni, T. L. Berg, A. C. Berg, and Y. Choi · 2011
Earlier work this paper cites.
A joint model of language and perception for grounded attribute learning
C. Matuszek, N. FitzGerald, L. Zettlemoyer, L. Bo, and D. Fox · 2012
Earlier work this paper cites.
Indoor segmentation and support inference from rgbd images
N. Silberman, D. Hoiem, P. Kohli, and R. Fergus · 2012
Earlier work this paper cites.
Framing image description as a ranking task: Data, models and evaluation metrics
M. Hodosh, P. Young, and J. Hockenmaier · 2013
Earlier work this paper cites.
YAGO2: A spatially and temporally enhanced knowledge base from Wikipedia
J. Hoffart, F. M. Suchanek, K. Berberich, and G. Weikum · 2013
Earlier work this paper cites.
Toward interactive grounded language acqusition
T. Kollar, J. Krishnamurthy, and G. P. Strimel · 2013
Earlier work this paper cites.
Learning dependency-based compositional semantics
P. Liang, M. I. Jordan, and D. Klein · 2013
Earlier work this paper cites.
Efficient estimation of word representations in vector space
T. Mikolov, K. Chen, G. Corrado, and J. Dean · 2013
Earlier work this paper cites.
Bringing semantics into focus using visual abstraction
C. L. Zitnick and D. Parikh · 2013
Earlier work this paper cites.
Learning the visual interpretation of sentences
C. L. Zitnick, D. Parikh, and L. Vanderwende · 2013
Earlier work this paper cites.
Zero-shot learning via visual abstraction
S. Antol, C. L. Zitnick, and D. Parikh · 2014
Earlier work this paper cites.
Predicting object dynamics in scenes
D. F. Fouhey and C. Zitnick · 2014
Earlier work this paper cites.
Resource description framework, 2014
R. W. Group et al · 2014
Earlier work this paper cites.
A neural network for factoid question answering over paragraphs
M. Iyyer, J. Boyd-Graber, L. Claudino, R. Socher, and H. Daumé III · 2014
Earlier work this paper cites.
Deep fragment embeddings for bidirectional image sentence mapping
A. Karpathy, A. Joulin, and F. F. Li · 2014
Earlier work this paper cites.
Referitgame: Referring to objects in photographs of natural scenes
S. Kazemzadeh, V. Ordonez, M. Matten, and T. L. Berg · 2014
Earlier work this paper cites.
Microsoft COCO: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Earlier work this paper cites.
A multi-world approach to question answering about real-world scenes based on uncertain input
M. Malinowski and M. Fritz · 2014
Earlier work this paper cites.
Glove: Global Vectors for Word Representation
J. Pennington, R. Socher, and C. Manning · 2014
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2014
Cited alongside, same era.
Question answering: A survey of research, techniques and issues
V. Singh and S. K. Dwivedi · 2014
Cited alongside, same era.
Webchild: Harvesting and organizing commonsense knowledge from the web
N. Tandon, G. de Melo, F. Suchanek, and G. Weikum · 2014
Cited alongside, same era.
Acquiring Comparative Commonsense Knowledge from the Web
N. Tandon, G. De Melo, and G. Weikum · 2014
Cited alongside, same era.
Joint video and text parsing for understanding events and answering queries
K. Tu, M. Meng, M. W. Lee, T. E. Choe, and S.-C. Zhu · 2014
Describing videos by exploiting temporal structure
L. Yao, A. Torabi, K. Cho, N. Ballas, C. Pal, H. Larochelle, and A. Courville · 2015
Later among the works it cites.
Visual madlibs: Fill in the blank image generation and question answering
L. Yu, E. Park, A. C. Berg, and T. L. Berg · 2015
Later among the works it cites.
Simple baseline for visual question answering
B. Zhou, Y. Tian, S. Sukhbaatar, A. Szlam, and R. Fergus · 2015
Later among the works it cites.
Building a Large-scale Multimodal Knowledge Base System for Answering Visual Queries
Y. Zhu, C. Zhang, C. Ré, and L. Fei-Fei · 2015
Later among the works it cites.
Learning to compose neural networks for question answering
J. Andreas, M. Rohrbach, T. Darrell, and D. Klein · 2016
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Show and tell: A neural image caption generator
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan · 2014
Cited alongside, same era.
J. Weston, S. Chopra, and A. Bordes · 2014
Cited alongside, same era.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
P. Young, A. Lai, M. Hodosh, and J. Hockenmaier · 2014
Cited alongside, same era.
Visualizing and understanding convolutional networks
M. D. Zeiler and R. Fergus · 2014
Cited alongside, same era.
VQA: Visual Question Answering
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. L. Zitnick, and D. Parikh · 2015
Cited alongside, same era.
Large-scale simple question answering with memory networks
A. Bordes, N. Usunier, S. Chopra, and J. Weston · 2015
Cited alongside, same era.
Neural Module Networks
J. Andreas, M. Rohrbach, T. Darrell, and D. Klein · 2016
Closest in time.
Language to logical form with neural attention
L. Dong and M. Lapata · 2016
Closest in time.
Multimodal compact bilinear pooling for visual question answering and visual grounding
A. Fukui, D. H. Park, D. Yang, A. Rohrbach, T. Darrell, and M. Rohrbach · 2016
Closest in time.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Closest in time.
Natural language object retrieval
R. Hu, H. Xu, M. Rohrbach, J. Feng, K. Saenko, and T. Darrell · 2016
Closest in time.
A focused dynamic attention model for visual question answering
I. Ilievski, S. Yan, and J. Feng · 2016
Closest in time.
Revisiting visual question answering baselines
A. Jabri, A. Joulin, and L. van der Maaten · 2016
Closest in time.
Answer-type prediction for visual question answering
K. Kafle and C. Kanan · 2016
Closest in time.
A diagram is worth a dozen images
A. Kembhavi, M. Salvato, E. Kolve, M. J. Seo, H. Hajishirzi, and A. Farhadi · 2016
Closest in time.
Multimodal residual learning for visual qa
J.-H. Kim, S.-W. Lee, D.-H. Kwak, M.-O. Heo, J. Kim, J.-W. Ha, and B.-T. Zhang · 2016
Closest in time.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
R. Krishna, Y. Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y. Kalantidis, L.-J. Li, D. A. Shamma, M. Bernstein, and L. Fei-Fei · 2016
Closest in time.
Ask me anything: Dynamic memory networks for natural language processing
A. Kumar, O. Irsoy, J. Su, J. Bradbury, R. English, B. Pierce, P. Ondruska, I. Gulrajani, and R. Socher · 2016
Closest in time.
Hierarchical question-image co-attention for visual question answering
J. Lu, J. Yang, D. Batra, and D. Parikh · 2016
Closest in time.
Learning to Answer Questions From Image using Convolutional Neural Network
L. Ma, Z. Lu, and H. Li · 2016
Closest in time.
Generation and comprehension of unambiguous object descriptions
J. Mao, H. Jonathan, A. Toshev, O. Camburu, A. Yuille, and K. Murphy · 2016
Closest in time.
Training recurrent answering units with joint loss minimization for vqa
H. Noh and B. Han · 2016
Closest in time.
Image Question Answering using Convolutional Neural Network with Dynamic Parameter Prediction
H. Noh, P. H. Seo, and B. Han · 2016
Closest in time.
Dualnet: Domain-invariant network for visual question answering
K. Saito, A. Shin, Y. Ushiku, and T. Harada · 2016
Closest in time.
Where to look: Focus regions for visual question answering
K. J. Shih, S. Singh, and D. Hoiem · 2016
Closest in time.
Fvqa: Fact-based visual question answering
P. Wang, Q. Wu, C. Shen, A. v. d. Hengel, and A. Dick · 2016
Closest in time.
FVQA: Fact-based visual question answering
P. Wang, Q. Wu, C. Shen, A. van den Hengel, and A. Dick · 2016
Closest in time.
What Value Do Explicit High Level Concepts Have in Vision to Language Problems?
Q. Wu, C. Shen, A. v. d. Hengel, L. Liu, and A. Dick · 2016
Closest in time.
Q. Wu, C. Shen, A. v. d. Hengel, P. Wang, and A. Dick · 2016
Closest in time.
Ask Me Anything: Free-form Visual Question Answering Based on Knowledge from External Sources
Q. Wu, P. Wang, C. Shen, A. Dick, and A. v. d. Hengel · 2016
Closest in time.
Dynamic memory networks for visual and textual question answering
C. Xiong, S. Merity, and R. Socher · 2016
Closest in time.
Stacked Attention Networks for Image Question Answering
Z. Yang, X. He, J. Gao, L. Deng, and A. Smola · 2016
Closest in time.
Yin and yang: Balancing and answering binary visual questions
P. Zhang, Y. Goyal, D. Summers-Stay, D. Batra, and D. Parikh · 2016
Closest in time.
Visual7W: Grounded Question Answering in Images
Y. Zhu, O. Groth, M. Bernstein, and L. Fei-Fei · 2016
Closest in time.
Adopting abstract images for semantic scene understanding
C. L. Zitnick, R. Vedantam, and D. Parikh · 2016
Closest in time.