Fetching the paper…
Reading the bibliography…
Inspired by recent advances in multimodal learning and machine translation, we introduce an encoder-decoder pipeline that learns (a): a multimodal joint embedding space with images and text and (b): a novel language model for decoding distributed representations from our space.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
A neural probabilistic language model
Yoshua Bengio, Réjean Ducharme, Pascal Vincent, and Christian Janvin · 2003
Earlier work this paper cites.
Three new graphical models for statistical language modelling
Andriy Mnih and Geoffrey Hinton · 2007
Earlier work this paper cites.
Unsupervised learning of image transformations
Roland Memisevic and Geoffrey Hinton · 2007
Earlier work this paper cites.
A novel connectionist system for unconstrained handwriting recognition
Alex Graves, Marcus Liwicki, Santiago Fernández, Roman Bertolami, Horst Bunke, and Jürgen Schmidhuber · 2009
Earlier work this paper cites.
Large scale image annotation: learning to rank with joint word-image embeddings
Jason Weston, Samy Bengio, and Nicolas Usunier · 2010
Earlier work this paper cites.
Every picture tells a story: Generating sentences from images
Ali Farhadi, Mohsen Hejrati, Mohammad Amin Sadeghi, Peter Young, Cyrus Rashtchian, Julia Hockenmaier, and David Forsyth · 2010
Earlier work this paper cites.
Factored 3-way restricted boltzmann machines for modeling natural images
Alex Krizhevsky, Geoffrey E Hinton, et al · 2010
Earlier work this paper cites.
Multimodal deep learning
Jiquan Ngiam, Aditya Khosla, Mingyu Kim, Juhan Nam, Honglak Lee, and Andrew Ng · 2011
Earlier work this paper cites.
Learning cross-modality similarity for multinomial data
Yangqing Jia, Mathieu Salzmann, and Trevor Darrell · 2011
Earlier work this paper cites.
Baby talk: Understanding and generating simple image descriptions
Girish Kulkarni, Visruth Premraj, Sagnik Dhar, Siming Li, Yejin Choi, Alexander C Berg, and Tamara L Berg · 2011
Earlier work this paper cites.
Composing simple image descriptions using web-scale n-grams
Siming Li, Girish Kulkarni, Tamara L Berg, Alexander C Berg, and Yejin Choi · 2011
Earlier work this paper cites.
Corpus-guided sentence generation of natural images
Yezhou Yang, Ching Lik Teo, Hal Daumé III, and Yiannis Aloimonos · 2011
Earlier work this paper cites.
Im2text: Describing images using 1 million captioned photographs
Vicente Ordonez, Girish Kulkarni, and Tamara L Berg · 2011
Earlier work this paper cites.
Multimodal learning with deep boltzmann machines
Nitish Srivastava and Ruslan Salakhutdinov · 2012
Earlier work this paper cites.
Midge: Generating image descriptions from computer vision detections
Margaret Mitchell, Xufeng Han, Jesse Dodge, Alyssa Mensch, Amit Goyal, Alex Berg, Kota Yamaguchi, Tamara Berg, Karl Stratos, and Hal Daumé III · 2012
Earlier work this paper cites.
Collective generation of natural image descriptions
Polina Kuznetsova, Vicente Ordonez, Alexander C Berg, Tamara L Berg, and Yejin Choi · 2012
Cited alongside, same era.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoff Hinton · 2012
Cited alongside, same era.
Framing image description as a ranking task: Data, models and evaluation metrics
Micah Hodosh, Peter Young, and Julia Hockenmaier · 2013
Cited alongside, same era.
Devise: A deep visual-semantic embedding model
Andrea Frome, Greg S Corrado, Jon Shlens, Samy Bengio, Jeffrey Dean, and Tomas Mikolov MarcAurelio Ranzato · 2013
Cited alongside, same era.
Recurrent continuous translation models
Nal Kalchbrenner and Phil Blunsom · 2013
Cited alongside, same era.
Linguistic regularities in continuous space word representations
Tomas Mikolov, Wen-tau Yih, and Geoffrey Zweig · 2013
Sequence to sequence learning with neural networks
Ilya Sutskever, Oriol Vinyals, and Quoc V Le · 2014
Closest in time.
Multilingual distributed representations without word alignment
Karl Moritz Hermann and Phil Blunsom · 2014
Closest in time.
Multilingual models for compositional distributional semantics
Karl Moritz Hermann and Phil Blunsom · 2014
Closest in time.
Deep fragment embeddings for bidirectional image sentence mapping
Andrej Karpathy, Armand Joulin, and Li Fei-Fei · 2014
Closest in time.
Improving image-sentence embeddings using large weakly annotated photo collections
Yunchao Gong, Liwei Wang, Micah Hodosh, Julia Hockenmaier, and Svetlana Lazebnik · 2014
Closest in time.
A deep architecture for semantic parsing
Phil Blunsom, Nando de Freitas, Edward Grefenstette, Karl Moritz Hermann, et al · 2014
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Translating video content to natural language descriptions
Marcus Rohrbach, Wei Qiu, Ivan Titov, Stefan Thater, Manfred Pinkal, and Bernt Schiele · 2013
Cited alongside, same era.
Generating sequences with recurrent neural networks
Alex Graves · 2013
Cited alongside, same era.
Hybrid speech recognition with deep bidirectional lstm
Alex Graves, Navdeep Jaitly, and Abdel-rahman Mohamed · 2013
Cited alongside, same era.
Efficient estimation of word representations in vector space
Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean · 2013
Cited alongside, same era.
Multimodal neural language models
Ryan Kiros, Richard S Zemel, and Ruslan Salakhutdinov · 2014
Cited alongside, same era.
Grounded compositional semantics for finding and describing images with sentences
Richard Socher, Q Le, C Manning, and A Ng · 2014
Cited alongside, same era.
Treetalk : Composition and compression of trees for image descriptions
Polina Kuznetsova, Vicente Ordonez, Tamara L. Berg, and Yejin Choi · 2014
Closest in time.
A multiplicative model for learning distributed text-based attribute representations
Ryan Kiros, Richard S Zemel, and Ruslan Salakhutdinov · 2014
Closest in time.
Dropout: A simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Closest in time.
Recurrent neural network regularization
Wojciech Zaremba, Ilya Sutskever, and Oriol Vinyals · 2014
Closest in time.
Fast and robust neural network joint models for statistical machine translation
Jacob Devlin, Rabih Zbib, Zhongqiang Huang, Thomas Lamar, Richard Schwartz, and John Makhoul · 2014
Closest in time.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Peter Young Alice Lai Micah Hodosh and Julia Hockenmaier · 2014
Closest in time.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2014
Closest in time.
Rich feature hierarchies for accurate object detection and semantic segmentation
Ross Girshick, Jeff Donahue, Trevor Darrell, and Jitendra Malik · 2014
Closest in time.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Closest in time.