Fetching the paper…
Reading the bibliography…
Existing image captioning models do not generalize well to out-of-domain images containing novel scenes or objects.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
WordNet: An Electronic Lexical Database
Christiane Fellbaum. 1998 · 1998
Earlier work this paper cites.
Statistical Machine Translation
Philipp Koehn. 2010 · 2010
Earlier work this paper cites.
Introduction to the Theory of Computation
Michael Sipser. 2012 · 2012
Earlier work this paper cites.
Fast image tagging
Minmin Chen, Alice X Zheng, and Kilian Q Weinberger. 2013 · 2013
Earlier work this paper cites.
Pushdown automata in statistical machine translation
Cyril Allauzen, Bill Byrne, Adrià de Gispert, Gonzalo Iglesias, and Michael Riley. 2014 · 2014
Earlier work this paper cites.
Meteor universal: Language specific translation evaluation for any target language
Michael Denkowski and Alon Lavie. 2014 · 2014
Earlier work this paper cites.
Caffe: Convolutional architecture for fast feature embedding
Yangqing Jia, Evan Shelhamer, Jeff Donahue, Sergey Karayev, Jonathan Long, Ross Girshick, Sergio Guadarrama, and Trevor Darrell. 2014 · 2014
Earlier work this paper cites.
Microsoft COCO: Common objects in context
T.Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollar, and C. L. Zitnick. 2014 · 2014
Earlier work this paper cites.
GloVe: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher D. Manning. 2014 · 2014
Earlier work this paper cites.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Peter Young, Alice Lai, Micah Hodosh, and Julia Hockenmaier. 2014 · 2014
Earlier work this paper cites.
Microsoft COCO captions: Data collection and evaluation server
Xinlei Chen, Tsung-Yi Lin Hao Fang, Ramakrishna Vedantam, Saurabh Gupta, Piotr Dollar, and C. Lawrence Zitnick. 2015 · 2015
Cited alongside, same era.
Language models for image captioning: The quirks and what works
Jacob Devlin, Hao Cheng, Hao Fang, Saurabh Gupta, Li Deng, Xiaodong He, Geoffrey Zweig, and Margaret Mitchell. 2015 · 2015
Cited alongside, same era.
Long-term recurrent convolutional networks for visual recognition and description
Jeffrey Donahue, Lisa A. Hendricks, Sergio Guadarrama, Marcus Rohrbach, Subhashini Venugopalan, Kate Saenko, and Trevor Darrell. 2015 · 2015
Cited alongside, same era.
Describing images using inferred visual dependency representations
Desmond Elliot and Arjen P. de Vries. 2015 · 2015
Cited alongside, same era.
From captions to visual concepts and back
Hao Fang, Saurabh Gupta, Forrest N. Iandola, Rupesh Srivastava, Li Deng, Piotr Dollar, Jianfeng Gao, Xiaodong He, Margaret Mitchell, John C. Platt, C. Lawrence Zitnick, and Geoffrey Zweig. 2015 · 2015
Show and tell: A neural image caption generator
Oriol Vinyals, Alexander Toshev, Samy Bengio, and Dumitru Erhan. 2015 · 2015
Later among the works it cites.
SPICE: Semantic propositional image caption evaluation
Peter Anderson, Basura Fernando, Mark Johnson, and Stephen Gould. 2016 · 2016
Closest in time.
Generating topical poetry
Marjan Ghazvininejad, Xing Shi, Yejin Choi, and Kevin Knight. 2016 · 2016
Closest in time.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Closest in time.
Deep compositional captioning: Describing novel object categories without paired training data
Lisa Anne Hendricks, Subhashini Venugopalan, Marcus Rohrbach, Raymond Mooney, Kate Saenko, and Trevor Darrell. 2016 · 2016
Closest in time.
The unreasonable effectiveness of noisy data for fine-grained recognition
Jonathan Krause, Benjamin Sapp, Andrew Howard, Howard Zhou, Alexander Toshev, Tom Duerig, James Philbin, and Li Fei-Fei. 2016 · 2016
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Deep visual-semantic alignments for generating image descriptions
Andrej Karpathy and Li Fei-Fei. 2015 · 2015
Cited alongside, same era.
Deep captioning with multimodal recurrent neural networks (m-RNN)
Junhua Mao, Wei Xu, Yi Yang, Jiang Wang, and Alan L. Yuille. 2015 · 2015
Cited alongside, same era.
Faster R-CNN: Towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. 2015 · 2015
Cited alongside, same era.
Imagenet large scale visual recognition challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li Fei-Fei. 2015 · 2015
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman. 2015 · 2015
Cited alongside, same era.
CIDEr: Consensus-based image description evaluation
Ramakrishna Vedantam, C. Lawrence Zitnick, and Devi Parikh. 2015 · 2015
Cited alongside, same era.
Closest in time.
Rich image captioning in the wild
Kenneth Tran, Xiaodong He, Lei Zhang, Jian Sun, Cornelia Carapcea, Chris Thrasher, Chris Buehler, and Chris Sienkiewicz. 2016 · 2016
Closest in time.
Captioning images with diverse objects
Subhashini Venugopalan, Lisa Anne Hendricks, Marcus Rohrbach, Raymond J. Mooney, Trevor Darrell, and Kate Saenko. 2016 · 2016
Closest in time.
What Value Do Explicit High Level Concepts Have in Vision to Language Problems?
Q. Wu, C. Shen, L. Liu, A. Dick, and A. van den Hengel. 2016 · 2016
Closest in time.
Yang Zhang, Boqing Gong, and Mubarak Shah. 2016 · 2016
Closest in time.