Fetching the paper…
Reading the bibliography…
We capitalize on large amounts of readily-available, synchronous data to learn a deep discriminative representations shared across three major natural modalities: vision, sound and language.
Fundamentals of speech recognition
L. Rabiner and B.-H. Juang · 1993
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
Auditory-visual integration during multimodal object recognition in humans: a behavioral and electrophysiological study
M. H. Giard and F. Peronnet · 1999
Earlier work this paper cites.
Evidence from functional magnetic resonance imaging of crossmodal binding in the human heteromodal cortex
G. A. Calvert, R. Campbell, and M. J. Brammer · 2000
Earlier work this paper cites.
Semantic-audio retrieval
M. Slaney · 2002
Earlier work this paper cites.
Multimedia content processing through cross-modal association
D. Li, N. Dimitrova, M. Li, and I. K. Sethi · 2003
Earlier work this paper cites.
Cross-modal correlation learning for clustering on image-audio dataset
H. Zhang, Y. Zhuang, and F. Wu · 2007
Earlier work this paper cites.
Large-scale content-based audio retrieval from text queries
G. Chechik, E. Ie, M. Rehn, S. Bengio, and D. Lyon · 2008
Earlier work this paper cites.
Utility data annotation with amazon mechanical turk
A. Sorokin and D. Forsyth · 2008
Earlier work this paper cites.
Semantic annotation and retrieval of music and sound effects
D. Turnbull, L. Barrington, D. Torres, and G. Lanckriet · 2008
Earlier work this paper cites.
Every picture tells a story: Generating sentences from images
A. Farhadi, M. Hejrati, M. A. Sadeghi, P. Young, C. Rashtchian, J. Hockenmaier, and D. Forsyth · 2010
Earlier work this paper cites.
Weakly-paired maximum covariance analysis for multimodal dimensionality reduction and transfer learning
C. H. Lampert and O. Krömer · 2010
Earlier work this paper cites.
Collecting image annotations using amazon’s mechanical turk
C. Rashtchian, P. Young, M. Hodosh, and J. Hockenmaier · 2010
Earlier work this paper cites.
A new approach to cross-modal multimedia retrieval
N. Rasiwasia, J. Costa Pereira, E. Coviello, G. Doyle, G. R. Lanckriet, R. Levy, and N. Vasconcelos · 2010
Earlier work this paper cites.
Multimodal deep learning
J. Ngiam, A. Khosla, M. Kim, J. Nam, H. Lee, and A. Y. Ng · 2011
Earlier work this paper cites.
Im2text: Describing images using 1 million captioned photographs
V. Ordonez, G. Kulkarni, and T. L. Berg · 2011
Cited alongside, same era.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Cited alongside, same era.
Multimodal human behavior analysis: learning correlation and interaction across modalities
Y. Song, L.-P. Morency, and R. Davis · 2012
Cited alongside, same era.
Devise: A deep visual-semantic embedding model
A. Frome, G. S. Corrado, J. Shlens, S. Bengio, J. Dean, T. Mikolov, et al · 2013
Cited alongside, same era.
Babytalk: Understanding and generating simple image descriptions
G. Kulkarni, V. Premraj, V. Ordonez, S. Dhar, S. Li, Y. Choi, A. C. Berg, and T. Berg · 2013
Cited alongside, same era.
Distributed representations of words and phrases and their compositionality
T. Mikolov, I. Sutskever, K. Chen, G. S. Corrado, and J. Dean · 2013
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2014
Later among the works it cites.
Grounded compositional semantics for finding and describing images with sentences
R. Socher, A. Karpathy, Q. V. Le, C. D. Manning, and A. Y. Ng · 2014
Later among the works it cites.
Object detectors emerge in deep scene cnns
B. Zhou, A. Khosla, A. Lapedriza, A. Oliva, and A. Torralba · 2014
Later among the works it cites.
From captions to visual concepts and back
H. Fang, S. Gupta, F. Iandola, R. K. Srivastava, L. Deng, P. Dollár, J. Gao, X. He, M. Mitchell, J. C. Platt, et al · 2015
Later among the works it cites.
Cross modal distillation for supervision transfer
S. Gupta, J. Hoffman, and J. Malik · 2015
Later among the works it cites.
Deep visual-semantic alignments for generating image descriptions
A. Karpathy and L. Fei-Fei · 2015
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
The visual microphone: passive recovery of sound from video
A. Davis, M. Rubinstein, N. Wadhwa, G. J. Mysore, F. Durand, and W. T. Freeman · 2014
Cited alongside, same era.
A multi-view embedding space for modeling internet images, tags, and their semantics
Y. Gong, Q. Ke, M. Isard, and S. Lazebnik · 2014
Cited alongside, same era.
Improving image-sentence embeddings using large weakly annotated photo collections
Y. Gong, L. Wang, M. Hodosh, J. Hockenmaier, and S. Lazebnik · 2014
Cited alongside, same era.
Deep speech: Scaling up end-to-end speech recognition
A. Hannun, C. Case, J. Casper, B. Catanzaro, G. Diamos, E. Elsen, R. Prenger, S. Satheesh, S. Sengupta, A. Coates, et al · 2014
Cited alongside, same era.
Caffe: Convolutional architecture for fast feature embedding
Y. Jia, E. Shelhamer, J. Donahue, S. Karayev, J. Long, R. Girshick, S. Guadarrama, and T. Darrell · 2014
Cited alongside, same era.
Convolutional neural networks for sentence classification
Y. Kim · 2014
Cited alongside, same era.
Later among the works it cites.
Visually indicated sounds
A. Owens, P. Isola, J. McDermott, A. Torralba, E. H. Adelson, and W. T. Freeman · 2015
Later among the works it cites.
The new data and new challenges in multimedia research
B. Thomee, D. A. Shamma, G. Friedland, B. Elizalde, K. Ni, D. Poland, D. Borth, and L.-J. Li · 2015
Later among the works it cites.
Order-embeddings of images and language
I. Vendrov, R. Kiros, S. Fidler, and R. Urtasun · 2015
Later among the works it cites.
Show and tell: A neural image caption generator
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan · 2015
Later among the works it cites.
Y. Zhu, R. Kiros, R. Zemel, R. Salakhutdinov, R. Urtasun, A. Torralba, and S. Fidler · 2015
Later among the works it cites.
Soundnet: Learning sound representations from unlabeled video
Y. Aytar, C. Vondrick, and A. Torralba · 2016
Later among the works it cites.
Learning aligned cross-modal representations from weakly aligned data
L. Castrejon, Y. Aytar, C. Vondrick, H. Pirsiavash, and A. Torralba · 2016
Later among the works it cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
R. Krishna, Y. Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y. Kalanditis, L.-J. Li, D. A. Shamma, M. Bernstein, and L. Fei-Fei · 2016
Later among the works it cites.
Ambient sound provides supervision for visual learning
A. Owens, J. Wu, J. H. McDermott, W. T. Freeman, and A. Torralba · 2016
Later among the works it cites.