Fetching the paper…
Reading the bibliography…
In this paper we present an approach to multi-language image description bringing together insights from neural machine translation and neural image description.
Long short-term memory
Hochreiter, Sepp and Schmidhuber, Jürgen · 1997
Earlier work this paper cites.
BLEU: A method for automatic evaluation of machine translation
Papineni, K., Roukos, S., Ward, T., and Zhu, W.-J · 2002
Earlier work this paper cites.
The IAPR TC-12 benchmark: A new evaluation resource for visual information systems
Grubinger, M., Clough, P. D., Muller, H., and Thomas, D · 2006
Earlier work this paper cites.
Every picture tells a story: Generating sentences from images
Farhadi, Ali, Hejrati, M, Sadeghi, Mohammad Amin, Young, P, Rashtchian, C, Hockenmaier, J, and Forsyth, David · 2010
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Glorot, Xavier and Bengio, Yoshua · 2010
Earlier work this paper cites.
Recurrent neural network based language model
Mikolov, Tomas, Karafiát, Martin, Burget, Lukas, Cernocký, Jan, and Khudanpur, Sanjeev · 2010
Earlier work this paper cites.
Collecting image annotations using Amazon’s Mechanical Turk
Rashtchian, C., Young, P., Hodosh, M., and Hockenmaier, J · 2010
Earlier work this paper cites.
Building a persistent workforce on Mechanical Turk for multilingual data collection
Chen, David L. and Dolan, William B · 2011
Earlier work this paper cites.
Better hypothesis testing for statistical machine translation: Controlling for optimizer instability
Clark, Jonathan H., Dyer, Chris, Lavie, Alon, and Smith, Noah A · 2011
Earlier work this paper cites.
Composing simple image descriptions using web-scale n-grams
Li, S, Kulkarni, G, Berg, T L, Berg, A C, and Choi, Y Young · 2011
Earlier work this paper cites.
Corpus-guided sentence generation of natural images
Yang, Y, Teo, C L, Daume, III, Hal, and Aloimonos, Y · 2011
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, Alex, Sutskever, Ilya, and Hinton, Geoffrey E · 2012
Earlier work this paper cites.
Continuous space translation models with neural networks
Le, Hai-Son, Allauzen, Alexandre, and Yvon, Francois · 2012
Earlier work this paper cites.
Midge: generating image descriptions from computer vision detections
Mitchell, Margaret, Han, Xufeng, Dodge, Jesse, Mensch, Alyssa, Goyal, Amit, Berg, A C, Yamaguchi, Kota, Berg, T L, Stratos, Karl, Daume, III, Hal, and III · 2012
Earlier work this paper cites.
Continuous space translation models for phrase-based statistical machine translation
Schwenk, Holger · 2012
Earlier work this paper cites.
Joint language and translation modeling with recurrent neural networks
Auli, Michael, Galley, Michel, Quirk, Chris, and Zweig, Geoffrey · 2013
Cited alongside, same era.
Image Description using Visual Dependency Representations
Elliott, Desmond and Keller, Frank · 2013
Cited alongside, same era.
Framing Image Description as a Ranking Task: Data, Models and Evaluation Metrics
Hodosh, Micah, Young, P, and Hockenmaier, J · 2013
Cited alongside, same era.
Learning phrase representations using RNN encoder–decoder for statistical machine translation
Cho, Kyunghyun, van Merrienboer, Bart, Gulcehre, Caglar, Bahdanau, Dzmitry, Bougares, Fethi, Schwenk, Holger, and Bengio, Yoshua · 2014
Cited alongside, same era.
Meteor Universal: Language Specific Translation Evaluation for Any Target Language
Denkowski, M. and Lavie, A · 2014
Cited alongside, same era.
Fast and robust neural network joint models for statistical machine translation
See no evil, say no evil: Description generation from densely labeled images
Yatskar, M, Vanderwende, L, and Zettlemoyer, L · 2014
Later among the works it cites.
Neural machine translation by jointly learning to align and translate
Bahdanau, Dzmitry, Cho, Kyunghyun, and Bengio, Yoshua · 2015
Closest in time.
Findings of the 2015 workshop on statistical machine translation
Bojar, Ondřej, Chatterjee, Rajen, Federmann, Christian, Haddow, Barry, Huck, Matthias, Hokamp, Chris, Koehn, Philipp, Logacheva, Varvara, Monz, Christof, Negri, Matteo, Post, Matt, Scarton, Carolina, Specia, Lucia, and Turchi, Marco · 2015
Closest in time.
Microsoft COCO captions: Data collection and evaluation server
Chen, Xinlei, Fang, Hao, Lin, Tsung-Yi, Vedantam, Ramakrishna, Gupta, Saurabh, Dollár, Piotr, and Zitnick, C. Lawrence · 2015
Closest in time.
Describing images using inferred visual dependency representations
Elliott, Desmond and de Vries, Arjen P · 2015
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Devlin, Jacob, Zbib, Rabih, Huang, Zhongqiang, Lamar, Thomas, Schwartz, Richard, and Makhoul, John · 2014
Cited alongside, same era.
Long-term recurrent convolutional networks for visual recognition and description
Donahue, J., Hendricks, L. A., Guadarrama, S., Rohrbach, M., Venugopalan, S., Saenko, K., and Darrell, T · 2014
Cited alongside, same era.
Comparing Automatic Evaluation Measures for Image Description
Elliott, Desmond and Keller, Frank · 2014
Cited alongside, same era.
Learning Image Embeddings using Convolutional Neural Networks for Improved Multi-Modal Semantics
Kiela, Douwe and Bottou, Léon · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, Diederik P. and Ba, Jimmy · 2014
Cited alongside, same era.
Multimodal neural language models
Kiros, Ryan, Salakhutdinov, Ruslan, and Zemel, Rich · 2014
Cited alongside, same era.
Learning grounded meaning representations with autoencoders
Silberer, Carina and Lapata, Mirella · 2014
Cited alongside, same era.
Closest in time.
From captions to visual concepts and back
Fang, Hao, Gupta, Saurabh, Iandola, Forrest, Srivastava, Rupesh K., Deng, Li, Dollar, Piotr, Gao, Jianfeng, He, Xiaodong, Mitchell, Margaret, Platt, John C., Lawrence Zitnick, C., and Zweig, Geoffrey · 2015
Closest in time.
Image-mediated learning for zero-shot cross-lingual document retrieval
Funaki, Ruka and Nakayama, Hideki · 2015
Closest in time.
Montreal neural machine translation systems for wmt’15
Jean, Sébastien, Firat, Orhan, Cho, Kyunghyun, Memisevic, Roland, and Bengio, Yoshua · 2015
Closest in time.
Deep visual-semantic alignments for generating image descriptions
Karpathy, Andrej and Fei-Fei, Li · 2015
Closest in time.
Visual bilingual lexicon induction with transferred convnet features
Kiela, Douwe, Vulić, Ivan, and Clark, Stephen · 2015
Closest in time.
Deep captioning with multimodal recurrent neural networks (m-RNN)
Mao, Junhua, Xu, Wei, Yang, Yi, Wang, Jiang, Huang, Zhiheng, and Yuille, Alan · 2015
Closest in time.
Very deep convolutional networks for large-scale image recognition
Simonyan, Karen and Zisserman, Andrew · 2015
Closest in time.
Show and tell: A neural image caption generator
Vinyals, Oriol, Toshev, Alexander, Bengio, Samy, and Erhan, Dumitru · 2015
Closest in time.
Show, attend and tell: Neural image caption generation with visual attention
Xu, Kelvin, Ba, Jimmy, Kiros, Ryan, Cho, Kyunghyun, Courville, Aaron C., Salakhutdinov, Ruslan, Zemel, Richard S., and Bengio, Yoshua · 2015
Closest in time.
Automatic description generation from images: A survey of models, datasets, and evaluation measures
Bernardi, Raffaella, Cakici, Ruken, Elliott, Desmond, Erdem, Aykut, Erdem, Erkut, Ikizler-Cinbis, Nazli, Keller, Frank, Muscat, Adrian, and Plank, Barbara · 2016
Closest in time.