Fetching the paper…
Reading the bibliography…
We explore a variety of nearest neighbor baseline approaches for image captioning.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu · 2002
Earlier work this paper cites.
Meta-analysis of face recognition algorithms
P. Phillips and E. Newton · 2002
Earlier work this paper cites.
The iapr tc-12 benchmark: A new evaluation resource for visual information systems
M. Grubinger, P. Clough, H. Müller, and T. Deselaers · 2006
Earlier work this paper cites.
Building the gist of a scene: The role of global image features in recognition
A. Oliva and A. Torralba · 2006
Earlier work this paper cites.
ImageNet: A Large-Scale Hierarchical Image Database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Every picture tells a story: Generating sentences from images
A. Farhadi, M. Hejrati, M. A. Sadeghi, P. Young, C. Rashtchian, J. Hockenmaier, and D. Forsyth · 2010
Earlier work this paper cites.
Im2text: Describing images using 1 million captioned photographs
V. Ordonez, G. Kulkarni, and T. Berg · 2011
Earlier work this paper cites.
ImageNet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. Hinton · 2012
Earlier work this paper cites.
Devise: A deep visual-semantic embedding model
A. Frome, G. S. Corrado, J. Shlens, S. Bengio, J. Dean, T. Mikolov, et al · 2013
Earlier work this paper cites.
Framing image description as a ranking task: Data, models and evaluation metrics
M. Hodosh, P. Young, and J. Hockenmaier · 2013
Earlier work this paper cites.
Grounded compositional semantics for finding and describing images with sentences
R. Socher, Q. Le, C. Manning, and A. Ng · 2013
Earlier work this paper cites.
Meteor universal: Language specific translation evaluation for any target language
M. Denkowski and A. Lavie · 2014
Earlier work this paper cites.
Comparing automatic evaluation measures for image description
D. Elliott and F. Keller · 2014
Cited alongside, same era.
Caffe: Convolutional architecture for fast feature embedding
Y. Jia, E. Shelhamer, J. Donahue, S. Karayev, J. Long, R. Girshick, S. Guadarrama, and T. Darrell · 2014
Cited alongside, same era.
Mining text snippets for images on the web
A. Kannan, S. Baker, K. Ramnath, J. Fiss, D. Lin, L. Vanderwende, R. Ansary, A. Kapoor, Q. Ke, M. Uyttendaele, X.-J. Wang, and L. Zhang · 2014
Cited alongside, same era.
Deep fragment embeddings for bidirectional image sentence mapping
A. Karpathy, A. Joulin, and F.-F. Li · 2014
Cited alongside, same era.
Multimodal neural language models
R. Kiros, R. Salakhutdinov, and R. Zemel · 2014
Cited alongside, same era.
Unifying visual-semantic embeddings with multimodal neural language models
Cider: Consensus-based image description evaluation
R. Vedantam, C. L. Zitnick, and D. Parikh · 2014
Later among the works it cites.
P. Young, A. Lai, M. Hodosh, and J. Hockenmaier · 2014
Later among the works it cites.
Déjá image-captions: A corpus of expressive image descriptions in repetition
J. Chen, P. Kuznetsova, D. Warren, and Y. Choi · 2015
Closest in time.
Mind’s eye: A recurrent visual representation for image caption generation
X. Chen and C. L. Zitnick · 2015
Closest in time.
Language models for image captioning: The quirks and what works
J. Devlin, H. Cheng, H. Fang, S. Gupta, L. Deng, X. He, G. Zweig, and M. Mitchell · 2015
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
R. Kiros, R. Salakhutdinov, and R. S. Zemel · 2014
Cited alongside, same era.
Simple image description generator via a linear phrase-based approach
R. Lebret, P. O. Pinheiro, and R. Collobert · 2014
Cited alongside, same era.
Microsoft COCO: Common objects in context
T. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Cited alongside, same era.
Deep captioning with multimodal recurrent neural networks (m-rnn)
J. Mao, W. Xu, Y. Yang, J. Wang, and A. L. Yuille · 2014
Cited alongside, same era.
Explain images with multimodal recurrent neural networks
J. Mao, W. Xu, Y. Yang, J. Wang, and A. L. Yuille · 2014
Cited alongside, same era.
Domain-specific image captioning
R. Mason and E. Charniak · 2014
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2014
Cited alongside, same era.
Closest in time.
Long-term recurrent convolutional networks for visual recognition and description
J. Donahue, L. A. Hendricks, S. Guadarrama, M. Rohrbach, S. Venugopalan, K. Saenko, and T. Darrell · 2015
Closest in time.
From captions to visual concepts and back
H. Fang, S. Gupta, F. Iandola, R. Srivastava, L. Deng, P. Dollár, J. Gao, X. He, M. Mitchell, J. Platt, et al · 2015
Closest in time.
Deep visual-semantic alignments for generating image descriptions
A. Karpathy and L. Fei-Fei · 2015
Closest in time.
Combining language and vision with a multimodal skip-gram model
A. Lazaridou, N. T. Pham, and M. Baroni · 2015
Closest in time.
R. Lebret, P. O. Pinheiro, and R. Collobert · 2015
Closest in time.
Show and tell: A neural image caption generator
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan · 2015
Closest in time.
Show, attend and tell: Neural image caption generation with visual attention
K. Xu, J. Ba, R. Kiros, K. Cho, A. C. Courville, R. Salakhutdinov, R. S. Zemel, and Y. Bengio · 2015
Closest in time.