Fetching the paper…
Reading the bibliography…
Recent work considered how images paired with speech can be used as supervision for building speech systems when transcriptions are not available.
J. G. Wilpon, L. R. Rabiner, C.-H. Lee, and E. Goldman, “Automatic recognition of keywords in unconstrained speech using hidden markov models,”
1990
Earlier work this paper cites.
P. Sheridan, M. Wechsler, and P. Schäuble, “Cross-language speech retrieval: Establishing a baseline performance,” in
1997
Earlier work this paper cites.
R. Caruana, “Multitask learning,”
1997
Earlier work this paper cites.
D. W. Oard and A. R. Diekema, “Cross-language information retrieval,”
1998
Earlier work this paper cites.
D. K. Roy and A. P. Pentland, “Learning words from sights and sounds: A computational model,”
2002
Earlier work this paper cites.
K. Barnard, P. Duygulu, D. Forsyth, N. d. Freitas, D. M. Blei, and M. I. Jordan, “Matching words and pictures,”
2003
Earlier work this paper cites.
I. Szöke, P. Schwarz, P. Matejka, L. Burget, M. Karafiát, M. Fapso, and J. Cernockỳ, “Comparison of keyword spotting approaches for informal continuous speech,” in
2005
Earlier work this paper cites.
A. Garcia and H. Gish, “Keyword spotting of arbitrary words using minimal speech resources,” in
2006
Earlier work this paper cites.
C. Chelba, T. J. Hazen, and M. Saraclar, “Retrieval and browsing of spoken content,”
2008
Earlier work this paper cites.
M. Guillaumin, T. Mensink, J. Verbeek, and C. Schmid, “Tagprop: Discriminative metric learning in nearest neighbor models for image auto-annotation,” in
2009
Earlier work this paper cites.
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” in
2009
Earlier work this paper cites.
T. J. Hazen, W. Shen, and C. White, “Query-by-example spoken term detection using phonetic posteriorgram templates,” in
2009
Earlier work this paper cites.
Y. Zhang and J. R. Glass, “Unsupervised spoken keyword spotting via segmental DTW on Gaussian posteriorgrams,” in
2009
Earlier work this paper cites.
H.-Y. Lee, T.-H. Wen, and L.-S. Lee, “Improved semantic retrieval of spoken content by language models enhanced with acoustic similarity graph,” in
2012
Earlier work this paper cites.
Y.-C. Li, H.-y. Lee, C.-T. Chung, C.-a. Chan, and L.-s. Lee, “Towards unsupervised semantic retrieval of spoken content with query expansion based on automatically discovered acoustic patterns,” in
2013
Earlier work this paper cites.
M. Chen, A. X. Zheng, and K. Q. Weinberger, “Fast image tagging,” in
2013
Cited alongside, same era.
L. Besacier, E. Barnard, A. Karpov, and T. Schultz, “Automatic speech recognition for under-resourced languages: A survey,”
2014
Cited alongside, same era.
G. Synnaeve, M. Versteegh, and E. Dupoux, “Learning words from images and speech,” in
2014
Cited alongside, same era.
K. Simonyan and A. Zisserman, “Very deep convolutional networks for large-scale image recognition,”
2014
Cited alongside, same era.
O. Räsänen and H. Rasilo, “A joint model of word segmentation and meaning acquisition through cross-situational learning,”
2015
Cited alongside, same era.
Y. Aytar, C. Vondrick, and A. Torralba, “Soundnet: Learning sound representations from unlabeled video,” in
2016
Later among the works it cites.
S. Gupta, J. Hoffman, and J. Malik, “Cross modal distillation for supervision transfer,” in
2016
Later among the works it cites.
G. Chrupała, L. Gelderloos, and A. Alishahi, “Representations of language in a model of visually grounded speech signal,”
2017
Later among the works it cites.
H. Kamper, S. Settle, G. Shakhnarovich, and K. Livescu, “Visually grounded learning of keyword prediction from untranscribed speech,”
2017
Later among the works it cites.
S. Bansal, H. Kamper, A. Lopez, and S. J. Goldwater, “Towards speech-to-text translation without speech recognition,” in
2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
L.-s. Lee, J. Glass, H.-y. Lee, and C.-a. Chan, “Spoken content retrieval—beyond cascading speech recognition with text retrieval,”
2015
Cited alongside, same era.
D. Harwath and J. Glass, “Deep multimodal semantic embeddings for speech and images,” in
2015
Cited alongside, same era.
D. Kingma and J. Ba, “Adam: A method for stochastic optimization,” in
2015
Cited alongside, same era.
G. Hinton, O. Vinyals, and J. Dean, “Distilling the knowledge in a neural network,”
2015
Cited alongside, same era.
D. Harwath, A. Torralba, and J. R. Glass, “Unsupervised learning of spoken language with visual context,” in
2016
Cited alongside, same era.
T. Taniguchi, T. Nagai, T. Nakamura, N. Iwahashi, T. Ogata, and H. Asoh, “Symbol emergence in robotics: A survey,”
2016
Cited alongside, same era.
L. Duong, A. Anastasopoulos, D. Chiang, S. Bird, and T. Cohn, “An attentional model for speech translation without transcription,” in
2016
Cited alongside, same era.
R. J. Weiss, J. Chorowski, N. Jaitly, Y. Wu, and Z. Chen, “Sequence-to-sequence models can directly translate foreign speech,” in
2017
Later among the works it cites.
J. Drexler and J. Glass, “Analysis of audio-visual features for unsupervised speech recognition,” 2017
2017
Later among the works it cites.
K. Leidal, D. Harwath, and J. Glass, “Learning modality-invariant representations for speech and images,”
2017
Later among the works it cites.
2017
Later among the works it cites.
D. Elliott and A. Kádár, “Imagination improves multimodal translation,”
2017
Later among the works it cites.
D. Elliott, S. Frank, L. Barrault, F. Bougares, and L. Specia, “Findings of the second shared task on multimodal machine translation and multilingual image description,” in
2017
Later among the works it cites.
A. Bérard, L. Besacier, A. C. Kocabiyikoglu, and O. Pietquin, “End-to-end automatic speech translation of audiobooks,”
2018
Closest in time.
2018
Closest in time.
D. Harwath, G. Chuang, and J. Glass, “Vision as an interlingua: Learning multilingual semantic embeddings of untranscribed speech,”
2018
Closest in time.