Fetching the paper…
Reading the bibliography…
Recent progress on automatic generation of image captions has shown that it is possible to describe the most salient information conveyed by images with accurate and meaningful sentences.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Latent dirichlet allocation
David M Blei, Andrew Y Ng, and Michael I Jordan · 2003
Earlier work this paper cites.
Factored conditional restricted boltzmann machines for modeling motion style
Graham W Taylor and Geoffrey E Hinton · 2009
Earlier work this paper cites.
Collecting image annotations using amazon’s mechanical turk
Cyrus Rashtchian, Peter Young, Micah Hodosh, and Julia Hockenmaier · 2010
Earlier work this paper cites.
Composing simple image descriptions using web-scale n-grams
Siming Li, Girish Kulkarni, Tamara L Berg, Alexander C Berg, and Yejin Choi · 2011
Earlier work this paper cites.
Generating text with recurrent neural networks
Ilya Sutskever, James Martens, and Geoffrey E Hinton · 2011
Earlier work this paper cites.
Corpus-guided sentence generation of natural images
Yezhou Yang, Ching Lik Teo, Hal Daumé III, and Yiannis Aloimonos · 2011
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Collective generation of natural image descriptions
Polina Kuznetsova, Vicente Ordonez, Alexander C Berg, Tamara L Berg, and Yejin Choi · 2012
Earlier work this paper cites.
Midge: Generating image descriptions from computer vision detections
Margaret Mitchell, Xufeng Han, Jesse Dodge, Alyssa Mensch, Amit Goyal, Alex Berg, Kota Yamaguchi, Tamara Berg, Karl Stratos, and Hal Daumé III · 2012
Earlier work this paper cites.
Image description using visual dependency representations
Desmond Elliott and Frank Keller · 2013
Cited alongside, same era.
Babytalk: Understanding and generating simple image descriptions
Girish Kulkarni, Visruth Premraj, Vicente Ordonez, Sagnik Dhar, Siming Li, Yejin Choi, Alexander C Berg, and Tamara L Berg · 2013
Cited alongside, same era.
Selective search for object recognition
Jasper RR Uijlings, Koen EA van de Sande, Theo Gevers, and Arnold WM Smeulders · 2013
Cited alongside, same era.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio · 2014
Cited alongside, same era.
Long-term recurrent convolutional networks for visual recognition and description
Jeff Donahue, Lisa Anne Hendricks, Sergio Guadarrama, Marcus Rohrbach, Subhashini Venugopalan, Kate Saenko, and Trevor Darrell · 2014
Cited alongside, same era.
From captions to visual concepts and back
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Later among the works it cites.
The stanford corenlp natural language processing toolkit
Christopher D Manning, Mihai Surdeanu, John Bauer, Jenny Finkel, Steven J Bethard, and David McClosky · 2014
Later among the works it cites.
Explain images with multimodal recurrent neural networks
Junhua Mao, Wei Xu, Yi Yang, Jiang Wang, and Alan L Yuille · 2014
Later among the works it cites.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2014
Later among the works it cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2014
Later among the works it cites.
Show and tell: A neural image caption generator
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hao Fang, Saurabh Gupta, Forrest Iandola, Rupesh Srivastava, Li Deng, Piotr Dollár, Jianfeng Gao, Xiaodong He, Margaret Mitchell, John Platt, et al · 2014
Cited alongside, same era.
Deep visual-semantic alignments for generating image descriptions
Andrej Karpathy and Li Fei-Fei · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba · 2014
Cited alongside, same era.
Unifying visual-semantic embeddings with multimodal neural language models
Ryan Kiros, Ruslan Salakhutdinov, and Richard S Zemel · 2014
Cited alongside, same era.
Treetalk: Composition and compression of trees for image descriptions
Polina Kuznetsova, Vicente Ordonez, Tamara L Berg, and Yejin Choi · 2014
Cited alongside, same era.
Oriol Vinyals, Alexander Toshev, Samy Bengio, and Dumitru Erhan · 2014
Later among the works it cites.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Peter Young, Alice Lai, Micah Hodosh, and Julia Hockenmaier · 2014
Later among the works it cites.
Learning Deep Features for Scene Recognition using Places Database
B. Zhou, A. Lapedriza, J. Xiao, A. Torralba, and A. Oliva · 2014
Later among the works it cites.
Microsoft coco captions: Data collection and evaluation server
Xinlei Chen, Hao Fang, Tsung-Yi Lin, Ramakrishna Vedantam, Saurabh Gupta, Piotr Dollar, and C Lawrence Zitnick · 2015
Closest in time.
Show, attend and tell: Neural image caption generation with visual attention
Kelvin Xu, Jimmy Ba, Ryan Kiros, Aaron Courville, Ruslan Salakhutdinov, Richard Zemel, and Yoshua Bengio · 2015
Closest in time.