Fetching the paper…
Reading the bibliography…
When building artificial intelligence systems that can reason and answer questions about visual data, we need diagnostic tests to analyze our progress and discover shortcomings.
Clever Hans (The horse of Mr. von Osten): A contribution to experimental animal and human psychology
O. Pfungst · 1911
Earlier work this paper cites.
Understanding Natural Language
T. Winograd · 1972
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Every picture tells a story: Generating sentences for images
A. Farhadi, M. Hejrati, A. Sadeghi, P. Young, C. Rashtchian, J. Hockenmaier, and D. Forsyth · 2010
Earlier work this paper cites.
The Winograd schema challenge
H. J. Levesque, E. Davis, and L. Morgenstern · 2011
Earlier work this paper cites.
Efficient estimation of word representations in vector space
T. Mikolov, K. Chen, G. Corrado, and J. Dean · 2013
Earlier work this paper cites.
Bringing semantics into focus using visual abstraction
C. Zitnick and D. Parikh · 2013
Earlier work this paper cites.
Representation learning: A review and new perspectives
Y. Bengio, A. Courville, and P. Vincent · 2014
Earlier work this paper cites.
Referitgame: Referring to objects in photographs of natural scenes
S. Kazemzadeh, V. Ordonez, M. Matten, and T. Berg · 2014
Earlier work this paper cites.
Microsoft COCO: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollar, and C. Zitnick · 2014
Earlier work this paper cites.
A multi-world approach to question answering about real-world scenes based on uncertain input
M. Malinowski and M. Fritz · 2014
Earlier work this paper cites.
Towards a visual Turing challenge
M. Malinowski and M. Fritz · 2014
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
N. Srivastava, G. E. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov · 2014
Earlier work this paper cites.
A simple method to determine if a music information retrieval system is a horse
B. Sturm · 2014
Earlier work this paper cites.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
P. Young, A. Lai, M. Hodosh, and J. Hockenmaier · 2014
Earlier work this paper cites.
VQA: Visual question answering
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. Zitnick, and D. Parikh · 2015
Earlier work this paper cites.
Deja image-captions: A corpus of expressive image descriptions in repetition
J. Chen, P. Kuznetsova, D. Warren, and Y. Choi · 2015
Earlier work this paper cites.
Are you talking to a machine? Dataset and methods for multilingual image question answering
H. Gao, J. Mao, J. Zhou, Z. Huang, L. Wang, and W. Xu · 2015
Cited alongside, same era.
Visual Turing test for computer vision systems
D. Geman, S. Geman, N. Hallonquist, and L. Younes · 2015
Cited alongside, same era.
Image retrieval using scene graphs
J. Johnson, R. Krishna, M. Stark, L.-J. Li, D. A. Shamma, M. S. Bernstein, and L. Fei-Fei · 2015
Cited alongside, same era.
Inferring algorithmic patterns with stack-augmented recurrent nets
A. Joulin and T. Mikolov · 2015
Cited alongside, same era.
Adam: A method for stochastic optimization
D. Kingma and J. Ba · 2015
Cited alongside, same era.
Ask your neurons: A neural-based approach to answering questions about images
M. Malinowski, M. Rohrbach, and M. Fritz · 2015
Cited alongside, same era.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Closest in time.
Revisiting visual question answering baselines
A. Jabri, A. Joulin, and L. van der Maaten · 2016
Closest in time.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
R. Krishna, Y. Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y. Kalantidis, L. Jia-Li, D. Shamma, M. Bernstein, and L. Fei-Fei · 2016
Closest in time.
Hierarchical question-image co-attention for visual question answering
J. Lu, J. Yang, D. Batra, and D. Parikh · 2016
Closest in time.
Learning to answer questions from image using convolutional neural network
L. Ma, Z. Lu, and H. Li · 2016
Closest in time.
Question relevance in vqa: Identifying non-visual and false-premise questions
A. Ray, G. Christie, M. Bansal, D. Batra, and D. Parikh · 2016
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Exploring models and data for image question answering
M. Ren, R. Kiros, and R. Zemel · 2015
Cited alongside, same era.
Memory networks
J. Weston, S. Chopra, and A. Bordes · 2015
Cited alongside, same era.
Visual madlibs: Fill in the blank image generation and question answering
L. Yu, E. Park, A. Berg, and T. Berg · 2015
Cited alongside, same era.
Simple baseline for visual question answering
B. Zhou, Y. Tian, S. Sukhbataar, A. Szlam, and R. Fergus · 2015
Cited alongside, same era.
Analyzing the behavior of visual question answering models
A. Agrawal, D. Batra, and D. Parikh · 2016
Cited alongside, same era.
Learning to compose neural networks for question answering
J. Andreas, M. Rohrbach, T. Darrell, and D. Klein · 2016
Cited alongside, same era.
Closest in time.
Where to look: Focus regions for visual question answering
K. Shih, S. Singh, and D. Hoiem · 2016
Closest in time.
Horse taxonomy and taxidermy
B. Sturm · 2016
Closest in time.
Movieqa: Understanding stories in movies through question-answering
M. Tapaswi, Y. Zhu, R. Stiefelhagen, A. Torralba, R. Urtasun, and S. Fidler · 2016
Closest in time.
Towards ai-complete question answering: A set of prerequisite toy tasks
J. Weston, A. Bordes, S. Chopra, A. Rush, B. van Merriënboer, A. Joulin, and T. Mikolov · 2016
Closest in time.
Image captioning and visual question answering based on attributes and their related external knowledge
Q. Wu, C. Shen, A. van den Hengel, P. Wang, and A. Dick · 2016
Closest in time.
Dynamic memory networks for visual and textual question answering
C. Xiong, S. Merity, and R. Socher · 2016
Closest in time.
Ask, attend, and answer: Exploring question-guided spatial attention for visual question answering
H. Xu and K. Saenko · 2016
Closest in time.
Stacked attention networks for image question answering
Z. Yang, X. He, J. Gao, L. Deng, and A. Smola · 2016
Closest in time.
Yin and yang: Balancing and answering binary visual questions
P. Zhang, Y. Goyal, D. Summers-Stay, D. Batra, and D. Parikh · 2016
Closest in time.
Visual7w: Grounded question answering in images
Y. Zhu, O. Groth, M. Bernstein, and L. Fei-Fei · 2016
Closest in time.