Fetching the paper…
Reading the bibliography…
In this paper, we propose the first model to be able to generate visually grounded questions with diverse types for a single image.
Improved backing-off for m-gram language modeling
Reinhard Kneser and Hermann Ney · 1995
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Natural image statistics and neural representation
Eero P Simoncelli and Bruno A Olshausen · 2001
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei jing Zhu · 2002
Earlier work this paper cites.
Matching words and pictures
Kobus Barnard, Pinar Duygulu, David Forsyth, Nando de Freitas, David M Blei, and Michael I Jordan · 2003
Earlier work this paper cites.
Statistical phrase-based translation
Philipp Koehn, Franz Josef Och, and Daniel Marcu · 2003
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Chin-Yew Lin · 2004
Earlier work this paper cites.
Meteor: An automatic metric for mt evaluation with improved correlation with human judgments
Satanjeev Banerjee and Alon Lavie · 2005
Earlier work this paper cites.
Dbpedia: A nucleus for a web of open data
Sóren Auer, Christian Bizer, Georgi Kobilarov, Jens Lehmann, Richard Cyganiak, and Zachary G. Ives · 2007
Earlier work this paper cites.
Yago: a core of semantic knowledge
Fabian M. Suchanek, Gjergji Kasneci, and Gerhard Weikum · 2007
Earlier work this paper cites.
Freebase: a collaboratively created graph database for structuring human knowledge
Kurt D. Bollacker, Colin Evans, Praveen Paritosh, Tim Sturge, and Jamie Taylor · 2008
Earlier work this paper cites.
Kernel density estimators
David W Scott · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Towards total scene understanding: Classification, annotation and segmentation in an automatic framework
Li-Jia Li, Richard Socher, and Li Fei-Fei · 2009
Earlier work this paper cites.
Baby talk: Understanding and generating image descriptions
Girish Kulkarni, Visruth Premraj, Sagnik Dhar, Siming Li, Yejin Choi, Alexander C Berg, and Tamara L Berg · 2011
Earlier work this paper cites.
From image annotation to image description
Ankush Gupta and Prashanth Mannem · 2012
Earlier work this paper cites.
Collective generation of natural image descriptions
Polina Kuznetsova, Vicente Ordonez, Alexander C Berg, Tamara L Berg, and Yejin Choi · 2012
Cited alongside, same era.
Framing image description as a ranking task: Data, models and evaluation metrics
Micah Hodosh, Peter Young, and Julia Hockenmaier · 2013
Cited alongside, same era.
Scaling semantic parsers with on-the-fly ontology matching
Tom Kwiatkowski, Eunsol Choi, Yoav Artzi, and Luke Zettlemoyer · 2013
Cited alongside, same era.
Learning the visual interpretation of sentences
C Lawrence Zitnick, Devi Parikh, and Lucy Vanderwende · 2013
Cited alongside, same era.
Open question answering with weakly supervised embedding models
Antoine Bordes, Jason Weston, and Nicolas Usunier · 2014
Cited alongside, same era.
Learning a recurrent visual representation for image caption generation
Gated feedback recurrent neural networks
Junyoung Chung, Caglar Gülçehre, Kyunghyun Cho, and Yoshua Bengio · 2015
Later among the works it cites.
Are you talking to a machine? dataset and methods for multilingual image question
Haoyuan Gao, Junhua Mao, Jie Zhou, Zhiheng Huang, Lei Wang, and Wei Xu · 2015
Later among the works it cites.
Visual turing test for computer vision systems
Donald Geman, Stuart Geman, Neil Hallonquist, and Laurent Younes · 2015
Later among the works it cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Later among the works it cites.
Densecap: Fully convolutional localization networks for dense captioning
Justin Johnson, Andrej Karpathy, and Li Fei-Fei · 2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Xinlei Chen and C Lawrence Zitnick · 2014
Cited alongside, same era.
Learning phrase representations using rnn encoder-decoder for statistical machine translation
Kyunghyun Cho, Bart Van Merriënboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba · 2014
Cited alongside, same era.
What are you talking about? text-to-image coreference
Chen Kong, Dahua Lin, Mohit Bansal, Raquel Urtasun, and Sanja Fidler · 2014
Cited alongside, same era.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Cited alongside, same era.
A multi-world approach to question answering about real-world scenes based on uncertain input
Mateusz Malinowski and Mario Fritz · 2014
Cited alongside, same era.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher D. Manning · 2014
Cited alongside, same era.
Deep visual-semantic alignments for generating image descriptions
Andrej Karpathy and Li Fei-Fei · 2015
Later among the works it cites.
Ask your neurons: A neural-based approach to answering questions about images
Mateusz Malinowski, Marcus Rohrbach, and Mario Fritz · 2015
Later among the works it cites.
Exploring models and data for image question answering
Mengye Ren, Ryan Kiros, and Richard Zemel · 2015
Later among the works it cites.
Going deeper with convolutions
Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich · 2015
Later among the works it cites.
Show and tell: A neural image caption generator
Oriol Vinyals, Alexander Toshev, Samy Bengio, and Dumitru Erhan · 2015
Later among the works it cites.
Towards ai-complete question answering: A set of prerequisite toy tasks
Jason Weston, Antoine Bordes, Sumit Chopra, Alexander M Rush, Bart van Merriënboer, Armand Joulin, and Tomas Mikolov · 2015
Later among the works it cites.
Show, attend and tell: Neural image caption generation with visual attention
Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron Courville, Ruslan Salakhutdinov, Richard S Zemel, and Yoshua Bengio · 2015
Later among the works it cites.
Visual madlibs: Fill in the blank image generation and question answering
Licheng Yu, Eunbyung Park, Alexander C Berg, and Tamara L Berg · 2015
Later among the works it cites.
Visual7w: Grounded question answering in images
Yuke Zhu, Oliver Groth, Michael Bernstein, and Li Fei-Fei · 2015
Later among the works it cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A Shamma, Michael Bernstein, and Li Fei-Fei · 2016
Closest in time.
Transforming dependency structures to logical forms for semantic parsing
Siva Reddy, Oscar Táckstr0́m, Michael Collins, Tom Kwiatkowski, Dipanjan Das, Mark Steedman, and Mirella Lapata · 2016
Closest in time.