Fetching the paper…
Reading the bibliography…
In this paper, we present a multimodal Recurrent Neural Network (m-RNN) model for generating novel image captions.
Learning representations by back-propagating errors
Rumelhart, David E, Hinton, Geoffrey E, and Williams, Ronald J · 1988
Earlier work this paper cites.
Finding structure in time
Elman, Jeffrey L · 1990
Earlier work this paper cites.
Long short-term memory
Hochreiter, Sepp and Schmidhuber, Jürgen · 1997
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Papineni, Kishore, Roukos, Salim, Ward, Todd, and Zhu, Wei-Jing · 2002
Earlier work this paper cites.
Matching words and pictures
Barnard, Kobus, Duygulu, Pinar, Forsyth, David, De Freitas, Nando, Blei, David M, and Jordan, Michael I · 2003
Earlier work this paper cites.
The iapr tc-12 benchmark: A new evaluation resource for visual information systems
Grubinger, Michael, Clough, Paul, Müller, Henning, and Deselaers, Thomas · 2006
Earlier work this paper cites.
Three new graphical models for statistical language modelling
Mnih, Andriy and Hinton, Geoffrey · 2007
Earlier work this paper cites.
Every picture tells a story: Generating sentences from images
Farhadi, Ali, Hejrati, Mohsen, Sadeghi, Mohammad Amin, Young, Peter, Rashtchian, Cyrus, Hockenmaier, Julia, and Forsyth, David · 2010
Earlier work this paper cites.
Multiple instance metric learning from automatically labeled bags of faces
Guillaumin, Matthieu, Verbeek, Jakob, and Schmid, Cordelia · 2010
Earlier work this paper cites.
Recurrent neural network based language model
Mikolov, Tomas, Karafiát, Martin, Burget, Lukas, Cernockỳ, Jan, and Khudanpur, Sanjeev · 2010
Earlier work this paper cites.
Rectified linear units improve restricted boltzmann machines
Nair, Vinod and Hinton, Geoffrey E · 2010
Earlier work this paper cites.
Collecting image annotations using amazon’s mechanical turk
Rashtchian, Cyrus, Young, Peter, Hodosh, Micah, and Hockenmaier, Julia · 2010
Earlier work this paper cites.
Learning cross-modality similarity for multinomial data
Jia, Yangqing, Salzmann, Mathieu, and Darrell, Trevor · 2011
Earlier work this paper cites.
Baby talk: Understanding and generating image descriptions
Kulkarni, Girish, Premraj, Visruth, Dhar, Sagnik, Li, Siming, Choi, Yejin, Berg, Alexander C, and Berg, Tamara L · 2011
Earlier work this paper cites.
Extensions of recurrent neural network language model
Mikolov, Tomas, Kombrink, Stefan, Burget, Lukas, Cernocky, JH, and Khudanpur, Sanjeev · 2011
Earlier work this paper cites.
From image annotation to image description
Gupta, Ankush and Mannem, Prashanth · 2012
Cited alongside, same era.
Choosing linguistics over vision to describe images
Gupta, Ankush, Verma, Yashaswi, and Jawahar, CV · 2012
Cited alongside, same era.
Imagenet classification with deep convolutional neural networks
Krizhevsky, Alex, Sutskever, Ilya, and Hinton, Geoffrey E · 2012
Cited alongside, same era.
Efficient backprop
LeCun, Yann A, Bottou, Léon, Orr, Genevieve B, and Müller, Klaus-Robert · 2012
Cited alongside, same era.
Midge: Generating image descriptions from computer vision detections
Mitchell, Margaret, Han, Xufeng, Dodge, Jesse, Mensch, Alyssa, Goyal, Amit, Berg, Alex, Yamaguchi, Kota, Berg, Tamara, Stratos, Karl, and Daumé III, Hal · 2012
Cited alongside, same era.
Multimodal learning with deep boltzmann machines
Srivastava, Nitish and Salakhutdinov, Ruslan · 2012
Cited alongside, same era.
Deep visual-semantic alignments for generating image descriptions
Karpathy, Andrej and Fei-Fei, Li · 2014
Closest in time.
Deep fragment embeddings for bidirectional image sentence mapping
Karpathy, Andrej, Joulin, Armand, and Fei-Fei, Li · 2014
Closest in time.
Treetalk: Composition and compression of trees for image descriptions
Kuznetsova, Polina, Ordonez, Vicente, Berg, Tamara L, and Choi, Yejin · 2014
Closest in time.
Microsoft coco: Common objects in context
Lin, Tsung-Yi, Maire, Michael, Belongie, Serge, Hays, James, Perona, Pietro, Ramanan, Deva, Dollár, Piotr, and Zitnick, C Lawrence · 2014
Closest in time.
Explain images with multimodal recurrent neural networks
Mao, Junhua, Xu, Wei, Yang, Yi, Wang, Jiang, and Yuille, Alan L · 2014
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Devise: A deep visual-semantic embedding model
Frome, Andrea, Corrado, Greg S, Shlens, Jon, Bengio, Samy, Dean, Jeff, Mikolov, Tomas, et al · 2013
Cited alongside, same era.
Framing image description as a ranking task: Data, models and evaluation metrics
Hodosh, Micah, Young, Peter, and Hockenmaier, Julia · 2013
Cited alongside, same era.
Recurrent continuous translation models
Kalchbrenner, Nal and Blunsom, Phil · 2013
Cited alongside, same era.
Distributed representations of words and phrases and their compositionality
Mikolov, Tomas, Sutskever, Ilya, Chen, Kai, Corrado, Greg S, and Dean, Jeff · 2013
Cited alongside, same era.
Learning a recurrent visual representation for image caption generation
Chen, Xinlei and Zitnick, C Lawrence · 2014
Cited alongside, same era.
Learning phrase representations using rnn encoder-decoder for statistical machine translation
Cho, Kyunghyun, van Merrienboer, Bart, Gulcehre, Caglar, Bougares, Fethi, Schwenk, Holger, and Bengio, Yoshua · 2014
Cited alongside, same era.
ImageNet Large Scale Visual Recognition Challenge, 2014
Russakovsky, Olga, Deng, Jia, Su, Hao, Krause, Jonathan, Satheesh, Sanjeev, Ma, Sean, Huang, Zhiheng, Karpathy, Andrej, Khosla, Aditya, Bernstein, Michael, Berg, Alexander C., and Fei-Fei, Li · 2014
Closest in time.
Very deep convolutional networks for large-scale image recognition
Simonyan, Karen and Zisserman, Andrew · 2014
Closest in time.
Grounded compositional semantics for finding and describing images with sentences
Socher, Richard, Le, Q, Manning, C, and Ng, A · 2014
Closest in time.
Sequence to sequence learning with neural networks
Sutskever, Ilya, Vinyals, Oriol, and Le, Quoc VV · 2014
Closest in time.
Cider: Consensus-based image description evaluation
Vedantam, Ramakrishna, Zitnick, C Lawrence, and Parikh, Devi · 2014
Closest in time.
Show and tell: A neural image caption generator
Vinyals, Oriol, Toshev, Alexander, Bengio, Samy, and Erhan, Dumitru · 2014
Closest in time.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Young, Peter, Lai, Alice, Hodosh, Micah, and Hockenmaier, Julia · 2014
Closest in time.
Microsoft coco captions: Data collection and evaluation server
Chen, X., Fang, H., Lin, TY, Vedantam, R., Gupta, S., Dollár, P., and Zitnick, C. L · 2015
Closest in time.
Learning like a child: Fast novel visual concept learning from sentence descriptions of images
Mao, Junhua, Xu, Wei, Yang, Yi, Wang, Jiang, Huang, Zhiheng, and Yuille, Alan · 2015
Closest in time.