Fetching the paper…
Reading the bibliography…
In this paper we consider the problem of continuously discovering image contents by actively asking image based questions and subsequently answering the questions being asked.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
HAL’s Legacy: 2001’s Computer as Dream and Reality
David G Stork · 1998
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Chin-Yew Lin · 2004
Earlier work this paper cites.
Automatic question generation for vocabulary assessment
Jonathan C Brown, Gwen A Frishkoff, and Maxine Eskenazi · 2005
Earlier work this paper cites.
The turing test: Verbal behavior as the hallmark of intelligence edited by stuart shieber
William J. Rapaport · 2005
Earlier work this paper cites.
Automation of question generation from sentences
Husam Ali, Yllias Chali, and Sadid A Hasan · 2010
Earlier work this paper cites.
Every picture tells a story: Generating sentences from images
Ali Farhadi, Mohsen Hejrati, Mohammad Amin Sadeghi, Peter Young, Cyrus Rashtchian, Julia Hockenmaier, and David Forsyth · 2010
Earlier work this paper cites.
Good question! statistical ranking for question generation
Michael Heilman and Noah A Smith · 2010
Earlier work this paper cites.
I2t: Image parsing to text description
Benjamin Z. Yao, Xiong Yang, Liang Lin, Mun Wai Lee, and Song Chun Zhu · 2010
Earlier work this paper cites.
Baby talk: Understanding and generating image descriptions
Girish Kulkarni, Visruth Premraj, Sagnik Dhar, Siming Li, Yejin Choi, Alexander C Berg, and Tamara L Berg · 2011
Earlier work this paper cites.
Im2text: Describing images using 1 million captioned photographs
Vicente Ordonez, Girish Kulkarni, and Tamara L. Berg · 2011
Earlier work this paper cites.
Corpus-guided sentence generation of natural images
Yezhou Yang, Ching Lik Teo, Hal Daumé, III, and Yiannis Aloimonos · 2011
Earlier work this paper cites.
Active scene recognition with vision and language
Xiaodong Yu, Cornelia Fermuller, Ching Lik Teo, Yezhou Yang, and Yiannis Aloimonos · 2011
Earlier work this paper cites.
Collective generation of natural image descriptions
Polina Kuznetsova, Vicente Ordonez, Alexander C. Berg, Tamara L. Berg, and Yejin Choi · 2012
Cited alongside, same era.
The uncanny valley [from the field]
Masahiro Mori, Karl F MacDorman, and Norri Kageki · 2012
Cited alongside, same era.
Image description using visual dependency representations
Desmond Elliott and Frank Keller · 2013
Cited alongside, same era.
Framing image description as a ranking task: Data, models and evaluation metrics
Micah Hodosh, Peter Young, and Julia Hockenmaier · 2013
Cited alongside, same era.
Learning a recurrent visual representation for image caption generation
Xinlei Chen and C Lawrence Zitnick · 2014
Cited alongside, same era.
Long-term recurrent convolutional networks for visual recognition and description
Grounded compositional semantics for finding and describing images with sentences
Richard Socher, Andrej Karpathy, Quoc V. Le, Christopher D. Manning, and Andrew Y. Ng · 2014
Later among the works it cites.
Cider: Consensus-based image description evaluation
Ramakrishna Vedantam, C Lawrence Zitnick, and Devi Parikh · 2014
Later among the works it cites.
Show and tell: A neural image caption generator
Oriol Vinyals, Alexander Toshev, Samy Bengio, and Dumitru Erhan · 2014
Later among the works it cites.
From images to sentences through scene description graphs using commonsense reasoning and knowledge
Somak Aditya, Yezhou Yang, Chitta Baral, Cornelia Fermuller, and Yiannis Aloimonos · 2015
Closest in time.
The cognitive dialogue: A new model for vision implementing common sense reasoning
Yiannis Aloimonos and Cornelia Fermüller · 2015
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jeff Donahue, Lisa Anne Hendricks, Sergio Guadarrama, Marcus Rohrbach, Subhashini Venugopalan, Kate Saenko, and Trevor Darrell · 2014
Cited alongside, same era.
Deep visual-semantic alignments for generating image descriptions
Andrej Karpathy and Fei-Fei Li · 2014
Cited alongside, same era.
Unifying visual-semantic embeddings with multimodal neural language models
Ryan Kiros, Ruslan Salakhutdinov, and Richard S Zemel · 2014
Cited alongside, same era.
Meteor universal: Language specific translation evaluation for any target language
Michael Denkowski Alon Lavie · 2014
Cited alongside, same era.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Cited alongside, same era.
Towards a visual turing challenge
Mateusz Malinowski and Mario Fritz · 2014
Cited alongside, same era.
Explain images with multimodal recurrent neural networks
Junhua Mao, Wei Xu, Yi Yang, Jiang Wang, and Alan L Yuille · 2014
Cited alongside, same era.
Closest in time.
Vqa: Visual question answering
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C. Lawrence Zitnick, and Devi Parikh · 2015
Closest in time.
Microsoft coco captions: Data collection and evaluation server
Xinlei Chen, Hao Fang, Tsung-Yi Lin, Ramakrishna Vedantam, Saurabh Gupta, Piotr Dollar, and C Lawrence Zitnick · 2015
Closest in time.
Are you talking to a machine? dataset and methods for multilingual image question answering
Haoyuan Gao, Junhua Mao, Jie Zhou, Zhiheng Huang, Lei Wang, and Wei Xu · 2015
Closest in time.
Image retrieval using scene graphs
Justin Johnson, Ranjay Krishna, Michael Stark, Jia Li, Michael Bernstein, and Li Fei-Fei · 2015
Closest in time.
Learning to answer questions from image using convolutional neural network
Lin Ma, Zhengdong Lu, and Hang Li · 2015
Closest in time.
Ask your neurons: A neural-based approach to answering questions about images
Mateusz Malinowski, Marcus Rohrbach, and Mario Fritz · 2015
Closest in time.
Exploring models and data for image question answering
Mengye Ren, Ryan Kiros, and Richard Zemel · 2015
Closest in time.
Generating semantically precise scene graphs from textual descriptions for improved image retrieval
Sebastian Schuster, Ranjay Krishna, Angel Chang, Li Fei-Fei, and Christopher D. Manning · 2015
Closest in time.