Fetching the paper…
Reading the bibliography…
In this paper, we present a model which takes as input a corpus of images with relevant spoken captions and finds a correspondence between the two modalities.
“The design for the wall street journal-based csr corpus,”
D. B. Paul and J. M. Baker, · 1992
Earlier work this paper cites.
“Matching words and pictures,”
K. Barnard, P. Duygulu, D. Forsyth, N. DeFreitas, D. M. Blei, and M. I. Jordan, · 2003
Earlier work this paper cites.
“Imagenet: A large scale hierarchical image database,”
J. Deng, W. Dong, R. Socher, L. J. Li, K. Li, and L. Fei-Fei, · 2009
Earlier work this paper cites.
“Connecting modalities: Semi-supervised segmentation and annotation of images using unaligned text corpora,”
R. Socher and L. Fei-Fei, · 2010
Earlier work this paper cites.
“Collecting image annotations using amazon’s mechanical turk,”
C. Rashtchian, P. Young, M. Hodosh, and J. Hockenmaier, · 2010
Earlier work this paper cites.
“The Kaldi speech recognition toolkit,”
D. Povey, A. Ghoshal, G. Boulianne, L. Burget, O. Glembek, N. Goel, M. Hannemann, P. Motlicek, Y. Qian, P. Schwarz, J. Silovsky, G. Stemmer, and K. Vesely, · 2011
Earlier work this paper cites.
“A joint model of language and perception for grounded attribute learning,”
C. Matuszek, N. Fitzgerald, L. Zettlemoyer, L. Bo, and D. Fox, · 2012
Earlier work this paper cites.
“Improving word representations via global context and multiple word prototypes,”
E. Huang, R. Socher, C. D. Manning, and A. Y. Ng, · 2012
Earlier work this paper cites.
“Devise: A deep visual-semantic embedding model,”
A. Frome, G. S. Corrado, J. Shlens, S. Bengio, J. Dean, M. Ranzato, and T. Mikolov, · 2013
Cited alongside, same era.
“Rich feature hierarchies for accurate object detection and semantic segmentation,”
R. Girshick, J. Donahue, T. Darrell, and J. Malik, · 2013
Cited alongside, same era.
“Selective search for object recognition,”
J. Uijlings, K. van de Sande, T. Gevers, and A. Smeulders, · 2013
Cited alongside, same era.
“What are you talking about? text-to-image coreference,”
C. Kong, K. Lin, M. Bansal, R. Urtasun, and S. Fidler, · 2014
Cited alongside, same era.
“Visual semantic search: Retrieving videos via complex textual queries,”
D. Lin, S. Fidler, C. Kong, and R. Urtasun, · 2014
Cited alongside, same era.
“Deep fragment embeddings for bidirectional image sentence mapping,”
A. Karpathy, A. Joulin, and L. Fei-Fei, · 2014
“Caffe: Convolutional architecture for fast feature embedding,”
Y. Jia, E. Shelhamer, J. Donahue, S. Karayev, J. Long, R. Girshick, S. Guadarrama, and T. Darrell, · 2014
Later among the works it cites.
“From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions,”
P. Young, A. Lai, M. Hodosh, and J. Hockenmaier, · 2014
Later among the works it cites.
“Microsoft coco: Common objects in context,”
T. Y. Lin, M. Marie, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollar, and C. L. Zitnick, · 2014
Later among the works it cites.
“Grounded compositional semantics for finding and describing images with sentences,”
R. Socher, A. Karpathy, Q. V. Le, C. D. Manning, and A. Y. Ng, · 2014
Later among the works it cites.
“Deep visual-semantic alignments for generating image descriptions,”
A. Karpathy and L. Fei-Fei, · 2015
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“Word embeddings for speech recognition,”
S. Bengio and G. Heigold, · 2014
Cited alongside, same era.
Closest in time.
“Show and tell: A neural image caption generator,”
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan, · 2015
Closest in time.
“Spoke: A framework for building speech-enabled websites,”
P. Saylor, · 2015
Closest in time.