Fetching the paper…
Reading the bibliography…
We propose a model to learn visually grounded word embeddings (vis-w2v) to capture visual notions of semantic relatedness.
Estimation of probabilities from sparse data for the language model component of a speech recognizer
S. M. Katz · 1987
Earlier work this paper cites.
An empirical study of smoothing techniques for language modeling
S. F. Chen, S. F. Chen, J. Goodman, and J. Goodman · 1998
Earlier work this paper cites.
Nltk: The natural language toolkit
E. Loper and S. Bird · 2002
Earlier work this paper cites.
A neural probabilistic language model
Y. Bengio, R. Ducharme, P. Vincent, and C. Jauvin · 2003
Earlier work this paper cites.
A unified architecture for natural language processing: Deep neural networks with multitask learning
R. Collobert and J. Weston · 2008
Earlier work this paper cites.
Neural network based language models for highly inflective languages
T. Mikolov, J. Kopecky, L. Burget, O. Glembek, and J. Cernocky · 2009
Earlier work this paper cites.
Baby talk: Understanding and generating image descriptions
G. Kulkarni, V. Premraj, S. Dhar, S. Li, Y. Choi, A. C. Berg, and T. L. Berg · 2011
Earlier work this paper cites.
Action recognition from a distributed representation of pose and appearance
S. Maji, L. Bourdev, and J. Malik · 2011
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Earlier work this paper cites.
Midge: Generating descriptions of images
M. Mitchell, X. Han, and J. Hayes · 2012
Earlier work this paper cites.
Framing image description as a ranking task: Data, models and evaluation metrics
M. Hodosh, P. Young, and J. Hockenmaier · 2013
Earlier work this paper cites.
Efficient Estimation of Word Representations in Vector Space
T. Mikolov, K. Chen, G. Corrado, and J. Dean · 2013
Earlier work this paper cites.
Distributed Representations of Words and Phrases and their Compositionality
T. Mikolov, I. Sutskever, K. Chen, G. S. Corrado, and J. Dean · 2013
Earlier work this paper cites.
Translating video content to natural language descriptions
M. Rohrbach, W. Qiu, I. Titov, S. Thater, M. Pinkal, and B. Schiele · 2013
Earlier work this paper cites.
Models of semantic representation with visual attributes
C. Silberer, V. Ferrari, and M. Lapata · 2013
Earlier work this paper cites.
Bringing semantics into focus using visual abstraction
C. L. Zitnick and D. Parikh · 2013
Cited alongside, same era.
Learning the visual interpretation of sentences
C. L. Zitnick, D. Parikh, and L. Vanderwende · 2013
Cited alongside, same era.
Zero-shot learning via visual abstraction
S. Antol, C. L. Zitnick, and D. Parikh · 2014
Cited alongside, same era.
Learning a recurrent visual representation for image caption generation
X. Chen and C. L. Zitnick · 2014
Cited alongside, same era.
Long-term recurrent convolutional networks for visual recognition and description
J. Donahue, L. A. Hendricks, S. Guadarrama, M. Rohrbach, S. Venugopalan, K. Saenko, and T. Darrell · 2014
Cited alongside, same era.
Adopting abstract images for semantic scene understanding
C. Zitnick, R. Vedantam, and D. Parikh · 2014
Later among the works it cites.
VQA: Visual question answering
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. L. Zitnick, and D. Parikh · 2015
Closest in time.
Unsupervised visual representation learning by context prediction
C. Doersch, A. Gupta, and A. A. Efros · 2015
Closest in time.
Are you talking to a machine? dataset and methods for multilingual image question answering
H. Gao, J. Mao, J. Zhou, Z. Huang, and A. Yuille · 2015
Closest in time.
Visual turing test for computer vision systems
D. Geman, S. Geman, N. Hallonquist, and L. Younes · 2015
Closest in time.
Image Specificity
M. Jas and D. Parikh · 2015
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Discriminative unsupervised feature learning with convolutional neural networks
A. Dosovitskiy, J. T. Springenberg, M. Riedmiller, and T. Brox · 2014
Cited alongside, same era.
Predicting object dynamics in scenes
D. F. Fouhey and C. L. Zitnick · 2014
Cited alongside, same era.
Rich feature hierarchies for accurate object detection and semantic segmentation
R. Girshick, J. Donahue, T. Darrell, and J. Malik · 2014
Cited alongside, same era.
Unifying visual-semantic embeddings with multimodal neural language models
R. Kiros, R. Salakhutdinov, and R. S. Zemel · 2014
Cited alongside, same era.
Microsoft COCO: Common objects in context
T. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Cited alongside, same era.
A multi-world approach to question answering about real-world scenes based on uncertain input
M. Malinowski and M. Fritz · 2014
Cited alongside, same era.
Explain images with multimodal recurrent neural networks
J. Mao, W. Xu, Y. Yang, J. Wang, and A. L. Yuille · 2014
Cited alongside, same era.
Deep visual-semantic alignments for generating image descriptions
A. Karpathy and L. Fei-Fei · 2015
Closest in time.
Combining language and vision with a multimodal skip-gram model
A. Lazaridou, N. T. Pham, and M. Baroni · 2015
Closest in time.
Don’t just listen, use your imagination: Leveraging visual common sense for non-visual tasks
X. Lin and D. Parikh · 2015
Closest in time.
Fully convolutional networks for semantic segmentation
J. Long, E. Shelhamer, and T. Darrell · 2015
Closest in time.
Ask your neurons: A neural-based approach to answering questions about images
M. Malinowski, M. Rohrbach, and M. Fritz · 2015
Closest in time.
Learning common sense through visual abstraction
T. B. C. L. Z. D. P. Ramakrishna Vedantam, Xiao Lin · 2015
Closest in time.
Image question answering: A visual semantic embedding model and a new dataset
M. Ren, R. Kiros, and R. S. Zemel · 2015
Closest in time.
Viske: Visual knowledge extraction and question answering by visual verification of relation phrases
F. Sadeghi, S. K. Divvala, and A. Farhadi · 2015
Closest in time.
Show and tell: A neural image caption generator
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan · 2015
Closest in time.