Fetching the paper…
Reading the bibliography…
We address a question answering task on real-world images that is set up as a Visual Turing Test.
A coefficient of agreement for nominal scales
J. Cohen et al · 1960
Earlier work this paper cites.
The equivalence of weighted kappa and the intraclass correlation coefficient as measures of reliability
J. L. Fleiss and J. Cohen · 1973
Earlier work this paper cites.
Verbs semantics and lexical selection
Z. Wu and M. Palmer · 1994
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Earlier work this paper cites.
A joint model of language and perception for grounded attribute learning
C. Matuszek, N. Fitzgerald, L. Zettlemoyer, L. Bo, and D. Fox · 2012
Earlier work this paper cites.
Indoor segmentation and support inference from rgbd images
N. Silberman, D. Hoiem, P. Kohli, and R. Fergus · 2012
Earlier work this paper cites.
Jointly learning to parse and perceive: Connecting natural language to the physical world
J. Krishnamurthy and T. Kollar · 2013
Earlier work this paper cites.
Learning dependency-based compositional semantics
P. Liang, M. I. Jordan, and D. Klein · 2013
Earlier work this paper cites.
Fine-grained semantic typing of emerging entities
N. Nakashole, T. Tylenda, and G. Weikum · 2013
Earlier work this paper cites.
Learning the visual interpretation of sentences
C. L. Zitnick, D. Parikh, and L. Vanderwende · 2013
Earlier work this paper cites.
Semantic parsing via paraphrasing
J. Berant and P. Liang · 2014
Earlier work this paper cites.
Learning phrase representations using rnn encoder-decoder for statistical machine translation
K. Cho, B. van Merrienboer, C. Gulcehre, F. Bougares, H. Schwenk, D. Bahdanau, and Y. Bengio · 2014
Cited alongside, same era.
A neural network for factoid question answering over paragraphs
M. Iyyer, J. Boyd-Graber, L. Claudino, R. Socher, and H. D. III · 2014
Cited alongside, same era.
Caffe: Convolutional architecture for fast feature embedding
Y. Jia, E. Shelhamer, J. Donahue, S. Karayev, J. Long, R. Girshick, S. Guadarrama, and T. Darrell · 2014
Cited alongside, same era.
Deep fragment embeddings for bidirectional image sentence mapping
A. Karpathy, A. Joulin, and L. Fei-Fei · 2014
Cited alongside, same era.
What are you talking about? text-to-image coreference
C. Kong, D. Lin, M. Bansal, R. Urtasun, and S. Fidler · 2014
Cited alongside, same era.
J. Weston, S. Chopra, and A. Bordes · 2014
Later among the works it cites.
W. Zaremba and I. Sutskever · 2014
Later among the works it cites.
Vqa: Visual question answering
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. L. Zitnick, and D. Parikh · 2015
Closest in time.
Long-term recurrent convolutional networks for visual recognition and description
J. Donahue, L. A. Hendricks, S. Guadarrama, M. Rohrbach, S. Venugopalan, K. Saenko, and T. Darrell · 2015
Closest in time.
Are you talking to a machine? dataset and methods for multilingual image question answering
H. Gao, J. Mao, J. Zhou, Z. Huang, L. Wang, and W. Xu · 2015
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A multi-world approach to question answering about real-world scenes based on uncertain input
M. Malinowski and M. Fritz · 2014
Cited alongside, same era.
Towards a visual turing challenge
M. Malinowski and M. Fritz · 2014
Cited alongside, same era.
Imagenet large scale visual recognition challenge
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, A. C. Berg, and L. Fei-Fei · 2014
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2014
Cited alongside, same era.
Sequence to sequence learning with neural networks
I. Sutskever, O. Vinyals, and Q. V. V. Le · 2014
Cited alongside, same era.
Going deeper with convolutions
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich · 2014
Cited alongside, same era.
Show and tell: A neural image caption generator
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan · 2014
Cited alongside, same era.
Visual turing test for computer vision systems
D. Geman, S. Geman, N. Hallonquist, and L. Younes · 2015
Closest in time.
Deep visual-semantic alignments for generating image descriptions
A. Karpathy and L. Fei-Fei · 2015
Closest in time.
Learning to answer questions from image using convolutional neural network
L. Ma, Z. Lu, and H. Li · 2015
Closest in time.
Image question answering: A visual semantic embedding model and a new dataset
M. Ren, R. Kiros, and R. Zemel · 2015
Closest in time.
Sequence to sequence – video to text
S. Venugopalan, M. Rohrbach, J. Donahue, R. Mooney, T. Darrell, and K. Saenko · 2015
Closest in time.
Translating videos to natural language using deep recurrent neural networks
S. Venugopalan, H. Xu, J. Donahue, M. Rohrbach, R. Mooney, and K. Saenko · 2015
Closest in time.
Visual madlibs: Fill in the blank image generation and question answering
L. Yu, E. Park, A. C. Berg, and T. L. Berg · 2015
Closest in time.