Fetching the paper…
Reading the bibliography…
There is growing interest in models that can learn from unlabelled speech paired with visual context.
S. Davis and P. Mermelstein, “Comparison of parametric representations for monosyllabic word recognition in continuously spoken sentences,” IEEE Trans. Acoust., Speech, Signal Process. , vol. 28, no. 4, pp. 357–366, 1980
1980
Earlier work this paper cites.
J. G. Wilpon, L. R. Rabiner, C.-H. Lee, and E. Goldman, “Automatic recognition of keywords in unconstrained speech using hidden Markov models,” IEEE Trans. Acoust. Speech Signal Process. , vol. 38, no. 11, pp. 1870–1878, 1990
1990
Earlier work this paper cites.
Z. Wu and M. Palmer, “Verbs semantics and lexical selection,” in Proc. ACL , 1994
1994
Earlier work this paper cites.
G. A. Miller, “WordNet: a lexical database for english,” Commun. ACM , vol. 38, no. 11, pp. 39–41, 1995
1995
Earlier work this paper cites.
J. M. Siskind, “A computational study of cross-situational techniques for learning word-to-meaning mappings,” Cognition , vol. 61, no. 1, pp. 39–91, 1996
1996
Earlier work this paper cites.
J. Xu and W. B. Croft, “Query expansion using local and global document analysis,” in Proc. SIGIR , 1996
1996
Earlier work this paper cites.
J. S. Garofolo, C. G. Auzanne, and E. M. Voorhees, “The TREC spoken document retrieval track: A success story,” in Content-Based Multimedia Information Access-Volume 1 , 2000, pp. 1–20
2000
Earlier work this paper cites.
D. K. Roy and A. P. Pentland, “Learning words from sights and sounds: A computational model,” Cognitive Sci. , vol. 26, no. 1, pp. 113–146, 2002
2002
Earlier work this paper cites.
K. Barnard, P. Duygulu, D. Forsyth, N. d. Freitas, D. M. Blei, and M. I. Jordan, “Matching words and pictures,” J. Mach. Learn. Res. , vol. 3, pp. 1107–1135, 2003
2003
Earlier work this paper cites.
C. Yu and D. H. Ballard, “A multimodal learning interface for grounding spoken language in sensory perceptions,” ACM T. Appl. Perception , vol. 1, no. 1, pp. 57–80, 2004
2004
Earlier work this paper cites.
E. V. Clark, “How language acquisition builds on cognitive development,” Trends Cogn. Sci. , vol. 8, no. 10, pp. 472–478, 2004
2004
Earlier work this paper cites.
J. Graupmann, R. Schenkel, and G. Weikum, “The SphereSearch engine for unified ranked retrieval of heterogeneous XML and web documents,” in Proc. VLDB , 2005
2005
Earlier work this paper cites.
I. Szöke, P. Schwarz, P. Matejka, L. Burget, M. Karafiát, M. Fapso, and J. Cernockỳ, “Comparison of keyword spotting approaches for informal continuous speech,” in Proc. Interspeech , 2005
2005
Earlier work this paper cites.
A. Garcia and H. Gish, “Keyword spotting of arbitrary words using minimal speech resources,” in Proc. ICASSP , 2006
2006
Earlier work this paper cites.
C. Yu and L. B. Smith, “Rapid word learning under uncertainty via cross-situational statistics,” Psychol. Sci. , vol. 18, no. 5, pp. 414–420, 2007
2007
Earlier work this paper cites.
L. Ten Bosch and B. Cranen, “A computational model for unsupervised word discovery,” in Proc. Interspeech , 2007
2007
Earlier work this paper cites.
T. J. Hazen, B. Sherry, and M. Adler, “Speech-based annotation and retrieval of digital photographs,” in Proc. Interspeech , 2007
2007
Earlier work this paper cites.
J. Luo, B. Caputo, A. Zweig, J.-H. Bach, and J. Anemüller, “Object category detection using audio-visual cues,” in Proc. ICVS , 2008
2008
Earlier work this paper cites.
B. Varadarajan, S. Khudanpur, and E. Dupoux, “Unsupervised learning of acoustic sub-word units,” in Proc. ACL , 2008
2008
Earlier work this paper cites.
A. S. Park and J. R. Glass, “Unsupervised pattern discovery in speech,” IEEE Trans. Audio, Speech, Language Process. , vol. 16, no. 1, pp. 186–197, 2008
2008
Earlier work this paper cites.
X. Anguera, J. Xu, and N. Oliver, “Multimodal photo annotation and retrieval on a mobile phone,” in Proc. ICMIR , 2008
2008
Earlier work this paper cites.
C. Chelba, T. J. Hazen, and M. Saraclar, “Retrieval and browsing of spoken content,” IEEE Signal Proc. Mag. , vol. 25, no. 3, 2008
2008
Earlier work this paper cites.
L. Van der Maaten and G. Hinton, “Visualizing data using t-SNE,” J. Mach. Learn. Res. , vol. 9, no. Nov, pp. 2579–2605, 2008
2008
Earlier work this paper cites.
V. Krunic, G. Salvi, A. Bernardino, L. Montesano, and J. Santos-Victor, “Affordance based word-to-meaning association,” in Proc. ICRA , 2009
2009
Earlier work this paper cites.
G. Aimetti, “Modelling early language acquisition skills: Towards a general statistical learning mechanism,” in Proc. EACL , 2009
2009
Earlier work this paper cites.
M. C. Frank, N. D. Goodman, and J. B. Tenenbaum, “Using speakers’ referential intentions to model early cross-situational word learning,” Psychol. Sci. , vol. 20, no. 5, pp. 578–585, 2009
2009
Earlier work this paper cites.
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “ImageNet: A large-scale hierarchical image database,” in Proc. CVPR , 2009
2009
Earlier work this paper cites.
M. Guillaumin, T. Mensink, J. Verbeek, and C. Schmid, “TagProp: Discriminative metric learning in nearest neighbor models for image auto-annotation,” in Proc. ICCV , 2009
2009
Earlier work this paper cites.
E. Agirre, E. Alfonseca, K. Hall, J. Kravalova, M. Paşca, and A. Soroa, “A study on similarity and relatedness using distributional and WordNet-based approaches,” in Proc. HLT-NAACL , 2009
2009
Earlier work this paper cites.
T. J. Hazen, W. Shen, and C. White, “Query-by-example spoken term detection using phonetic posteriorgram templates,” in Proc. ASRU , 2009
2009
Earlier work this paper cites.
Y. Zhang and J. R. Glass, “Unsupervised spoken keyword spotting via segmental DTW on Gaussian posteriorgrams,” in Proc. ASRU , 2009
2009
Earlier work this paper cites.
T. Cunillera, E. Camara, M. Laine, and A. Rodriguez-Fornells, “Speech segmentation is facilitated by visual cues,” Q. J. Exp. Psychol. , vol. 63, no. 2, pp. 260–274, 2010
2010
Earlier work this paper cites.
E. D. Thiessen, “Effects of visual information on adults’ and infants’ auditory statistical learning,” Cognitive Sci. , vol. 34, no. 6, pp. 1093–1106, 2010
2010
Earlier work this paper cites.
R. Socher and L. Fei-Fei, “Connecting modalities: Semi-supervised segmentation and annotation of images using unaligned text corpora,” in Proc. CVPR , 2010
2010
Earlier work this paper cites.
A. Farhadi, M. Hejrati, M. A. Sadeghi, P. Young, C. Rashtchian, J. Hockenmaier, and D. Forsyth, “Every picture tells a story: Generating sentences from images,” in Proc. ECCV , 2010
2010
Cited alongside, same era.
Y. Feng and M. Lapata, “Visual information in semantic representation,” in Proc. NAACL . Association for Computational Linguistics, 2010, pp. 91–99
2010
Cited alongside, same era.
J. Driesen and H. Van hamme, “Modelling vocabulary acquisition, adaptation and generalization in infants using adaptive Bayesian PLSA,” Neurocomputing , vol. 74, no. 11, pp. 1874–1882, 2011
2011
Cited alongside, same era.
J. Weston, S. Bengio, and N. Usunier, “WSABIE: Scaling up to large vocabulary image annotation,” in Proc. IJCAI , 2011
2011
Cited alongside, same era.
Y. Yang, C. L. Teo, H. Daumé III, and Y. Aloimonos, “Corpus-guided sentence generation of natural images,” in Proc. EMNLP , 2011
L.-s. Lee, J. R. Glass, H.-y. Lee, and C.-a. Chan, “Spoken content retrieval—beyond cascading speech recognition with text retrieval,” IEEE Trans. Audio, Speech, Language Process. , vol. 23, no. 9, pp. 1389–1420, 2015
2015
Later among the works it cites.
D. Harwath and J. R. Glass, “Deep multimodal semantic embeddings for speech and images,” in Proc. ASRU , 2015
2015
Later among the works it cites.
D. Kingma and J. Ba, “Adam: A method for stochastic optimization,” in Proc. ICLR , 2015
2015
Later among the works it cites.
F. Hill, R. Reichart, and A. Korhonen, “SimLex-999: Evaluating semantic models with (genuine) similarity estimation,” Comput. Linguist. , vol. 41, no. 4, 2015
2015
Later among the works it cites.
J. Wieting, M. Bansal, K. Gimpel, and K. Livescu, “From paraphrase database to compositional paraphrase model and back,” Trans. ACL , vol. 3, pp. 345–358, 2015
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2011
Cited alongside, same era.
C.-y. Lee and J. R. Glass, “A nonparametric Bayesian approach to acoustic model discovery,” in Proc. ACL , 2012
2012
Cited alongside, same era.
Y. Zhang, R. Salakhutdinov, H.-A. Chang, and J. R. Glass, “Resource configurable spoken query detection using deep Boltzmann machines,” in Proc. ICASSP , 2012
2012
Cited alongside, same era.
C. Silberer and M. Lapata, “Grounded models of semantic representation,” in Proc. EMNLP , 2012
2012
Cited alongside, same era.
E. Bruni, G. Boleda, M. Baroni, and N.-K. Tran, “Distributional semantics in technicolor,” in Proc. ACL , 2012
2012
Cited alongside, same era.
H.-Y. Lee, T.-H. Wen, and L.-S. Lee, “Improved semantic retrieval of spoken content by language models enhanced with acoustic similarity graph,” in Proc. SLT , 2012
2012
Cited alongside, same era.
A. Jansen, E. Dupoux, S. J. Goldwater, M. Johnson, S. Khudanpur, K. Church, N. Feldman, H. Hermansky, F. Metze, R. Rose, M. Seltzer, P. Clark, I. McGraw, B. Varadarajan, E. Bennett, B. Borschinger, J. Chiu, E. Dunbar, A. Fourtassi, D. Harwath, C.-y. Lee, K. Levin, A. Norouzian, V. Peddinti, R. Richardson, T. Schatz, and S. Thomas, “A summary of the 2012 JHU CLSP workshop on zero resource speech technologies and models of early language acquisition,” in Proc. ICASSP , 2013
2013
Cited alongside, same era.
M. Hodosh, P. Young, and J. Hockenmaier, “Framing image description as a ranking task: Data, models and evaluation metrics,” J. Artif. Intell. Res. , vol. 47, pp. 853–899, 2013
2013
Cited alongside, same era.
2015
Later among the works it cites.
G. Adda, S. Stüker, M. Adda-Decker, O. Ambouroue, L. Besacier, D. Blachon, H. Bonneau-Maynard, P. Godard, F. Hamlaoui, D. Idiatov et al. , “Breaking the unwritten language barrier: The BULB project,” Proc. SLTU , 2016
2016
Later among the works it cites.
L. Duong, A. Anastasopoulos, D. Chiang, S. Bird, and T. Cohn, “An attentional model for speech translation without transcription,” in Proc. NAACL , 2016, pp. 949–959
2016
Later among the works it cites.
D. Palaz, G. Synnaeve, and R. Collobert, “Jointly learning to locate and classify words using convolutional networks,” in Proc. Interspeech , 2016
2016
Later among the works it cites.
T. Taniguchi, T. Nagai, T. Nakamura, N. Iwahashi, T. Ogata, and H. Asoh, “Symbol emergence in robotics: A survey,” Adv. Robotics , vol. 30, no. 11-12, pp. 706–728, 2016
2016
Later among the works it cites.
D. Harwath, A. Torralba, and J. R. Glass, “Unsupervised learning of spoken language with visual context,” in Proc. NIPS , 2016
2016
Later among the works it cites.
L. Gelderloos and G. Chrupała, “From phonemes to images: Levels of representation in a recurrent neural model of visually-grounded language learning,” Proc. COLING , 2016
2016
Later among the works it cites.
M. Versteegh, X. Anguera, A. Jansen, and E. Dupoux, “The Zero Resource Speech Challenge 2015: Proposed approaches and results,” in Proc. SLTU , 2016
2016
Later among the works it cites.
H. Kamper, A. Jansen, and S. J. Goldwater, “Unsupervised word segmentation and lexicon discovery using acoustic word embeddings,” IEEE Trans. Audio, Speech, Language Process. , vol. 24, no. 4, pp. 669–679, 2016
2016
Later among the works it cites.
H. Kamper, “Unsupervised neural and bayesian models for zero-resource speech processing,” Ph.D. dissertation, University of Edinburgh, UK, 2016
2016
Later among the works it cites.
R. Bernardi, R. Cakici, D. Elliott, A. Erdem, E. Erdem, N. Ikizler-Cinbis, F. Keller, A. Muscat, and B. Plank, “Automatic description generation from images: A survey of models, datasets, and evaluation measures,” J. Artif. Intell. Res. , vol. 55, pp. 409–442, 2016
2016
Later among the works it cites.
F. Sun, D. Harwath, and J. R. Glass, “Look, listen, and decode: Multimodal speech recognition with images,” in Proc. SLT , 2016
2016
Later among the works it cites.
F. Diaz, B. Mitra, and N. Craswell, “Query expansion with locally-trained word embeddings,” in Proc. ACL , 2016
2016
Later among the works it cites.
D. Roy, D. Paul, M. Mitra, and U. Garain, “Using word embeddings for automatic query expansion,” in SIGIR Workshop on Neural Information Retrieval , 2016
2016
Later among the works it cites.
Y. Aytar, C. Vondrick, and A. Torralba, “SoundNet: Learning sound representations from unlabeled video,” in Proc. NIPS , 2016
2016
Later among the works it cites.
——, “Towards universal paraphrastic sentence embeddings,” Proc. ICLR , 2016
2016
Later among the works it cites.
S. Settle, K. Levin, H. Kamper, and K. Livescu, “Query-by-example search with discriminative neural acoustic word embeddings,” in Proc. Interspeech , 2017
2017
Closest in time.
S. Bansal, H. Kamper, A. Lopez, and S. J. Goldwater, “Towards speech-to-text translation without speech recognition,” in Proc. EACL , 2017
2017
Closest in time.
R. J. Weiss, J. Chorowski, N. Jaitly, Y. Wu, and Z. Chen, “Sequence-to-sequence models can directly translate foreign speech,” in Proc. Interspeech , 2017
2017
Closest in time.
2017
Closest in time.
G. Chrupała, L. Gelderloos, and A. Alishahi, “Representations of language in a model of visually grounded speech signal,” in Proc. ACL , 2017
2017
Closest in time.
J. Drexler and J. R. Glass, “Analysis of audio-visual features for unsupervised speech recognition,” in Proc. GLU , 2017
2017
Closest in time.
K. Leidal, D. Harwath, and J. Glass, “Learning modality-invariant representations for speech and images,” Proc. ASRU , 2017
2017
Closest in time.
H. Kamper, S. Settle, G. Shakhnarovich, and K. Livescu, “Visually grounded learning of keyword prediction from untranscribed speech,” Proc. Interspeech , 2017
2017
Closest in time.
A. Gupta, Y. Miao, L. Neves, and F. Metze, “Visual features for context-aware speech recognition,” in Proc. ICASSP , 2017
2017
Closest in time.
D. Harwath and J. R. Glass, “Learning word-like units from joint audio-visual analysis,” in Proc. ACL , 2017
2017
Closest in time.
2018
Closest in time.
2018
Closest in time.
D. Harwath, G. Chuang, and J. R. Glass, “Vision as an interlingua: Learning multilingual semantic embeddings of untranscribed speech,” in Proc. ICASSP , 2018
2018
Closest in time.