Fetching the paper…
Reading the bibliography…
We address a question answering task on real-world images that is set up as a Visual Turing Test.
A coefficient of agreement for nominal scales
Jacob Cohen et al · 1960
Earlier work this paper cites.
The equivalence of weighted kappa and the intraclass correlation coefficient as measures of reliability
Joseph L Fleiss and Jacob Cohen · 1973
Earlier work this paper cites.
Verbs semantics and lexical selection
Zhibiao Wu and Martha Palmer · 1994
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner · 1998
Earlier work this paper cites.
Foundations of statistical natural language processing , volume 999
Christopher D Manning and Hinrich Schütze · 1999
Earlier work this paper cites.
Accurate unlexicalized parsing
Dan Klein and Christopher D Manning · 2003
Earlier work this paper cites.
Theano: new features and speed improvements
Frédéric Bastien, Pascal Lamblin, Razvan Pascanu, James Bergstra, Ian J. Goodfellow, Arnaud Bergeron, Nicolas Bouchard, and Yoshua Bengio · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
A joint model of language and perception for grounded attribute learning
Cynthia Matuszek, Nicholas Fitzgerald, Luke Zettlemoyer, Liefeng Bo, and Dieter Fox · 2012
Earlier work this paper cites.
Indoor segmentation and support inference from rgbd images
Nathan Silberman, Derek Hoiem, Pushmeet Kohli, and Rob Fergus · 2012
Earlier work this paper cites.
Jointly learning to parse and perceive: Connecting natural language to the physical world
Jayant Krishnamurthy and Thomas Kollar · 2013
Earlier work this paper cites.
Learning dependency-based compositional semantics
Percy Liang, Michael I Jordan, and Dan Klein · 2013
Earlier work this paper cites.
Fine-grained semantic typing of emerging entities
Ndapandula Nakashole, Tomasz Tylenda, and Gerhard Weikum · 2013
Earlier work this paper cites.
Grounding Action Descriptions in Videos
Michaela Regneri, Marcus Rohrbach, Dominikus Wetzel, Stefan Thater, Bernt Schiele, and Manfred Pinkal · 2013
Earlier work this paper cites.
Learning the visual interpretation of sentences
C Lawrence Zitnick, Devi Parikh, and Lucy Vanderwende · 2013
Earlier work this paper cites.
Semantic parsing via paraphrasing
Jonathan Berant and Percy Liang · 2014
Earlier work this paper cites.
Learning phrase representations using rnn encoder-decoder for statistical machine translation
Kyunghyun Cho, Bart van Merrienboer, Caglar Gulcehre, Fethi Bougares, Holger Schwenk, Dzmitry Bahdanau, and Yoshua Bengio · 2014
Earlier work this paper cites.
A neural network for factoid question answering over paragraphs
Mohit Iyyer, Jordan Boyd-Graber, Leonardo Claudino, Richard Socher, and Hal Daumé III · 2014
Earlier work this paper cites.
Caffe: Convolutional architecture for fast feature embedding
Yangqing Jia, Evan Shelhamer, Jeff Donahue, Sergey Karayev, Jonathan Long, Ross Girshick, Sergio Guadarrama, and Trevor Darrell · 2014
Earlier work this paper cites.
A convolutional neural network for modelling sentences
Nal Kalchbrenner, Edward Grefenstette, and Phil Blunsom · 2014
Earlier work this paper cites.
Deep fragment embeddings for bidirectional image sentence mapping
Andrej Karpathy, Armand Joulin, and Li Fei-Fei · 2014
Earlier work this paper cites.
Referit game: Referring to objects in photographs of natural scenes
Sahar Kazemzadeh, Vicente Ordonez, Mark Matten, and Tamara L. Berg · 2014
Earlier work this paper cites.
Convolutional neural networks for sentence classification
Yoon Kim · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
What are you talking about? text-to-image coreference
Chen Kong, Dahua Lin, Mohit Bansal, Raquel Urtasun, and Sanja Fidler · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Cited alongside, same era.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher D. Manning · 2014
Cited alongside, same era.
Imagenet large scale visual recognition challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li Fei-Fei · 2014
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2014
Cited alongside, same era.
Sequence to sequence learning with neural networks
Ilya Sutskever, Oriol Vinyals, and Quoc V. V Le · 2014
Cited alongside, same era.
Stacked attention networks for image question answering
Zichao Yang, Xiaodong He, Jianfeng Gao, Li Deng, and Alex Smola · 2015
Later among the works it cites.
Visual madlibs: Fill in the blank description generation and question answering
Licheng Yu, Eunbyung Park, Alexander C Berg, and Tamara L Berg · 2015
Later among the works it cites.
Simple baseline for visual question answering
Bolei Zhou, Yuandong Tian, Sainbayar Sukhbaatar, Arthur Szlam, and Rob Fergus · 2015
Later among the works it cites.
Uncovering temporal context for video question and answering
Linchao Zhu, Zhongwen Xu, Yi Yang, and Alexander G Hauptmann · 2015
Later among the works it cites.
Multi-cue zero-shot learning with strong supervision
Zeynep Akata, Mateusz Malinowski, Mario Fritz, and Bernt Schiele · 2016
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich · 2014
Cited alongside, same era.
Trecvid med 14
Trecvid · 2014
Cited alongside, same era.
Show and tell: A neural image caption generator
Oriol Vinyals, Alexander Toshev, Samy Bengio, and Dumitru Erhan · 2014
Cited alongside, same era.
Jason Weston, Sumit Chopra, and Antoine Bordes · 2014
Cited alongside, same era.
Wojciech Zaremba and Ilya Sutskever · 2014
Cited alongside, same era.
Vqa: Visual question answering
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh · 2015
Cited alongside, same era.
Abc-cnn: An attention based convolutional neural network for visual question answering
Kan Chen, Jiang Wang, Liang-Chieh Chen, Haoyuan Gao, Wei Xu, and Ram Nevatia · 2015
Cited alongside, same era.
Closest in time.
Xplore-m-ego: Contextual media retrieval using natural language queries
Sreyasi Nag Chowdhury, Mateusz Malinowski, Andreas Bulling, and Mario Fritz · 2016
Closest in time.
Multimodal compact bilinear pooling for visual question answering and visual grounding
Akira Fukui, Dong Huk Park, Daylen Yang, Anna Rohrbach, Trevor Darrell, and Marcus Rohrbach · 2016
Closest in time.
A focused dynamic attention model for visual question answering
Ilija Ilievski, Shuicheng Yan, and Jiashi Feng · 2016
Closest in time.
Answer-type prediction for visual question answering
Kushal Kafle and Christopher Kanan · 2016
Closest in time.
Hadamard product for low-rank bilinear pooling
Jin-Hwa Kim, Kyoung Woon On, Jeonghee Kim, Jung-Woo Ha, and Byoung-Tak Zhang · 2016
Closest in time.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A Shamma, Michael Bernstein, and Li Fei-Fei · 2016
Closest in time.
Hierarchical Co-Attention for Visual Question Answering
Jiasen Lu, Jianwei Yang, Dhruv Batra, and Devi Parikh · 2016
Closest in time.
Learning to answer questions from image using convolutional neural network
Lin Ma, Zhengdong Lu, and Hang Li · 2016
Closest in time.
Tutorial on answering questions about images with deep learning
Mateusz Malinowski and Mario Fritz · 2016
Closest in time.
Generation and comprehension of unambiguous object descriptions
Junhua Mao, Jonathan Huang, Alexander Toshev, Oana Camburu, Alan Yuille, and Kevin Murphy · 2016
Closest in time.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
Bryan Plummer, Liwei Wang, Chris Cervantes, Juan Caicedo, Julia Hockenmaier, and Svetlana Lazebnik · 2016
Closest in time.
Highway networks for visual question answering
Aaditya Prakash and James Storer · 2016
Closest in time.
Dualnet: Domain-invariant network for visual question answering
Kuniaki Saito, Andrew Shin, Yoshitaka Ushiku, and Tatsuya Harada · 2016
Closest in time.
Where to look: Focus regions for visual question answering
Kevin J Shih, Saurabh Singh, and Derek Hoiem · 2016
Closest in time.
Movieqa: Understanding stories in movies through question-answering
Makarand Tapaswi, Yukun Zhu, Rainer Stiefelhagen, Antonio Torralba, Raquel Urtasun, and Sanja Fidler · 2016
Closest in time.
Learning deep structure-preserving image-text embeddings
Liwei Wang, Yin Li, and Svetlana Lazebnik · 2016
Closest in time.
Ask Me Anything: Free-form Visual Question Answering Based on Knowledge from External Sources
Qi Wu, Peng Wang, Chunhua Shen, Anton van den Hengel, and Anthony Dick · 2016
Closest in time.
Dynamic memory networks for visual and textual question answering
Caiming Xiong, Stephen Merity, and Richard Socher · 2016
Closest in time.
Modeling context in referring expressions
Licheng Yu, Patrick Poirson, Shan Yang, Alexander C Berg, and Tamara L Berg · 2016
Closest in time.
Visual7W: Grounded Question Answering in Images
Yuke Zhu, Oliver Groth, Michael Bernstein, and Li Fei-Fei · 2016
Closest in time.