Fetching the paper…
Reading the bibliography…
We propose the inverse problem of Visual question answering (iVQA), and explore its suitability as a benchmark for visuo-linguistic understanding.
Clever Hans (The Horse of Mr. Von Osten): A Contribution to Experimental Animal and Human Psychology
O. Pfungst · 1911
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Describing objects by their attributes
A. Farhadi, I. Endres, D. Hoiem, and D. Forsyth · 2009
Earlier work this paper cites.
Recognizing human-object interactions in still images by modeling the mutual context of objects and human poses
B. Yao and L. Fei-Fei · 2012
Earlier work this paper cites.
Counterfactual reasoning and learning systems: The example of computational advertising
L. Bottou, J. Peters, J. Quiñonero Candela, D. X. Charles, D. M. Chickering, E. Portugaly, D. Ray, P. Simard, and E. Snelson · 2013
Earlier work this paper cites.
Speech recognition with deep recurrent neural networks
A. Graves, A. Mohamed, and G. Hinton · 2013
Earlier work this paper cites.
Referit game: Referring to objects in photographs of natural scenes
S. Kazemzadeh, V. Ordonez, M. Matten, and T. L. Berg · 2014
Earlier work this paper cites.
A simple method to determine if a music information retrieval system is a horse
B. L. Sturm · 2014
Earlier work this paper cites.
Microsoft COCO captions: Data collection and evaluation server
X. Chen, H. Fang, T. Lin, R. Vedantam, S. Gupta, P. Dollár, and C. L. Zitnick · 2015
Earlier work this paper cites.
Exploring nearest neighbor approaches for image captioning
J. Devlin, S. Gupta, R. Girshick, M. Mitchell, and C. L. Zitnick · 2015
Earlier work this paper cites.
Visual turing test for computer vision systems
D. Geman, S. Geman, N. Hallonquist, and L. Younes · 2015
Earlier work this paper cites.
Explaining and harnessing adversarial examples
I. J. Goodfellow, J. Shlens, and C. Szegedy · 2015
Earlier work this paper cites.
Deep visual-semantic alignments for generating image descriptions
A. Karpathy and L. Fei-Fei · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. Kingma and J. Ba · 2015
Cited alongside, same era.
Ask your neurons: A neural-based approach to answering questions about images
M. Malinowski, M. Rohrbach, and M. Fritz · 2015
Cited alongside, same era.
Show and tell: A neural image caption generator
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan · 2015
Cited alongside, same era.
Show, attend and tell: Neural image caption generation with visual attention
K. Xu, J. Ba, R. Kiros, K. Cho, A. Courville, R. Salakhudinov, R. Zemel, and Y. Bengio · 2015
Cited alongside, same era.
Analyzing the behavior of visual question answering models
A. Agrawal, D. Batra, and D. Parikh · 2016
Cited alongside, same era.
Vqa: Visual question answering
A. Agrawal, J. Lu, S. Antol, M. Mitchell, C. L. Zitnick, D. Parikh, and D. Batra · 2016
Cited alongside, same era.
Visual relationship detection with language priors
C. Lu, R. Krishna, M. Bernstein, and L. Fei-Fei · 2016
Later among the works it cites.
Generating natural questions about an image
N. Mostafazadeh, I. Misra, J. Devlin, M. Mitchell, X. He, and L. Vanderwende · 2016
Later among the works it cites.
Ask, attend and answer: Exploring question-guided spatial attention for visual question answering
H. Xu and K. Saenko · 2016
Later among the works it cites.
Image caption generation with text-conditional semantic attention
L. Zhou, C. Xu, P. Koch, and J. J. Corso · 2016
Later among the works it cites.
Visual dialog
A. Das, S. Kottur, K. Gupta, A. Singh, D. Yadav, J. M. Moura, D. Parikh, and D. Batra · 2017
Closest in time.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Neural module networks
J. Andreas, M. Rohrbach, T. Darrell, and D. Klein · 2016
Cited alongside, same era.
Human attention in visual question answering: Do humans and deep networks look at the same regions?
A. Das, H. Agrawal, C. L. Zitnick, D. Parikh, and D. Batra · 2016
Cited alongside, same era.
Multimodal compact bilinear pooling for visual question answering and visual grounding
A. Fukui, D. H. Park, D. Yang, A. Rohrbach, T. Darrell, and M. Rohrbach · 2016
Cited alongside, same era.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Cited alongside, same era.
Revisiting visual question answering baselines
A. Jabri, A. Joulin, and L. van der Maaten · 2016
Cited alongside, same era.
Densecap: Fully convolutional localization networks for dense captioning
J. Johnson, A. Karpathy, and L. Fei-Fei · 2016
Cited alongside, same era.
Y. Goyal, T. Khot, D. Summers-Stay, D. Batra, and D. Parikh · 2017
Closest in time.
Hadamard product for low-rank bilinear pooling
J.-H. Kim, K. W. On, W. Lim, J. Kim, J.-W. Ha, and B.-T. Zhang · 2017
Closest in time.
Semantic regularisation for recurrent image annotation
F. Liu, T. Xiang, T. M. Hospedales, W. Yang, and C. Sun · 2017
Closest in time.
Knowing when to look: Adaptive attention via a visual sentinel for image captioning
J. Lu, C. Xiong, D. Parikh, and R. Socher · 2017
Closest in time.
Image-grounded conversations: Multimodal context for natural question and response generation
N. Mostafazadeh, C. Brockett, B. Dolan, M. Galley, J. Gao, G. P. Spithourakis, and L. Vanderwende · 2017
Closest in time.
Fvqa: Fact-based visual question answering
P. Wang, Q. Wu, C. Shen, A. Dick, and A. van den Hengel · 2017
Closest in time.
Automatic generation of grounded visual questions
S. Zhang, L. Qu, S. You, Z. Yang, and J. Zhang · 2017
Closest in time.