J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” in
2009
Earlier work this paper cites.
T. Mikolov, K. Chen, G. Corrado, and J. Dean, “Efficient estimation of word representations in vector space,”
Original
2013
Earlier work this paper cites.
G. Synnaeve, M. Versteegh, and E. Dupoux, “Learning words from images and speech,” in
2014
Earlier work this paper cites.
B. Zhou, A. Lapedriza, J. Xiao, A. Torralba, and A. Oliva, “Learning deep features for scene recognition using places database,” in
2014
Earlier work this paper cites.
D. Harwath and J. Glass, “Deep multimodal semantic embeddings for speech and images,” in
2015
Earlier work this paper cites.
D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” in
2015
Earlier work this paper cites.
B. Zhou, A. Khosla, A. Lapedriza, A. Oliva, and A. Torralba, “Object detectors emerge in deep scene cnns,” in
2015
Earlier work this paper cites.
D. Harwath, A. Torralba, and J. Glass, “Unsupervised learning of spoken language with visual context,” in
2016
Earlier work this paper cites.
Y. Aytar, C. Vondrick, and A. Torralba, “Soundnet: Learning sound representations from unlabeled video,” in
2016
Earlier work this paper cites.
J. Xu, T. Mei, T. Yao, and Y. Rui, “Msr-vtt: A large video description dataset for bridging video and language,” in
2016
Earlier work this paper cites.
A. Owens, J. Wu, J. H. McDermott, W. T. Freeman, and A. Torralba, “Ambient sound provides supervision for visual learning,” in
2016
Earlier work this paper cites.
J.-B. Alayrac, P. Bojanowski, N. Agrawal, J. Sivic, I. Laptev, and S. Lacoste-Julien, “Unsupervised learning from narrated instruction videos,” in
2016
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in
2016
Earlier work this paper cites.
G. Chrupała, L. Gelderloos, and A. Alishahi, “Representations of language in a model of visually grounded speech signal,” in
2017
Earlier work this paper cites.
R. Arandjelovic and A. Zisserman, “Look, listen and learn,” in
2017
Earlier work this paper cites.
A. Miech, I. Laptev, and J. Sivic, “Learnable pooling with context gating for video classification,” in
2017
Earlier work this paper cites.
A. Rohrbach, A. Torabi, M. Rohrbach, N. Tandon, C. Pal, H. Larochelle, A. Courville, and B. Schiele, “Movie description,” in
2017
Earlier work this paper cites.
W. Havard, L. Besacier, and O. Rosec, “Speech-coco: 600k visually grounded spoken captions aligned to mscoco data set,”
Original
2017
Earlier work this paper cites.
D. Harwath and J. Glass, “Learning word-like units from joint audio-visual analysis,” in
2017
Earlier work this paper cites.
K. Leidal, D. Harwath, and J. Glass, “Learning modality-invariant representations for speech and images,” in
2017
Earlier work this paper cites.
J. Carreira and A. Zisserman, “Quo vadis, action recognition? a new model and the kinetics dataset,” in
2017
Earlier work this paper cites.
D. Harwath, A. Recasens, D. Surís, G. Chuang, A. Torralba, and J. Glass, “Jointly discovering visual objects and spoken words from raw sensory input,” in
2018
Earlier work this paper cites.
——, “Objects that sound,” in
2018
Earlier work this paper cites.
A. Owens and A. A. Efros, “Audio-visual scene analysis with self-supervised multisensory features,” in
2018
Earlier work this paper cites.
B. Korbar, D. Tran, and L. Torresani, “Cooperative learning of audio and video models from self-supervised synchronization,” in
2018
Earlier work this paper cites.