Fetching the paper…
Reading the bibliography…
In this paper, we investigate the manner in which interpretable sub-word speech units emerge within a convolutional neural network model trained to associate raw speech waveforms with semantically related natural image scenes.
“The TIMIT acoustic-phonetic continuous speech corpus,” 1993
John Garofolo, Lori Lamel, William Fisher, Jonathan Fiscus, David Pallet, Nancy Dahlgren, and Victor Zue, · 1993
Earlier work this paper cites.
“Learning classification with unlabeled data,”
Virginia de Sa, · 1994
Earlier work this paper cites.
“Learning words from sights and sounds: a computational model,”
Deb Roy and Alex Pentland, · 2002
Earlier work this paper cites.
“Unsupervised pattern discovery in speech,”
Alex Park and James Glass, · 2008
Earlier work this paper cites.
“Unsupervised learning of acoustic sub-word units,”
Balakrishnan Varadarajan, Sanjeev Khudanpur, and Emmanuel Dupoux, · 2008
Earlier work this paper cites.
“Unsupervised training of an HMM-based speech recognizer for topic classification,”
Herb Gish, Man-Hung Siu, Arthur Chan, and William Belfield, · 2009
Earlier work this paper cites.
“Imagenet: A large scale hierarchical image database,”
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei, · 2009
Earlier work this paper cites.
“Toward spoken term discovery at scale with zero resources,”
Aren Jansen, Kenneth Church, and Hynek Hermansky, · 2010
Earlier work this paper cites.
“Unsupervised speech segmentation: An analy- sis of the hypothesized phone boundaries,”
Odette Scharenborg, Vincent Wan, and Mirjam Ernestus, · 2010
Earlier work this paper cites.
Nguyen Xuan Vinh, Julien Epps, and James Bailey, · 2010
Earlier work this paper cites.
“A nonparametric Bayesian approach to acoustic model discovery,”
Chia-Ying Lee and James Glass, · 2012
Earlier work this paper cites.
“Weak top-down constraints for unsupervised acoustic model training,”
Aren Jansen, Samuel Thomas, and Hynek Hermansky, · 2013
Cited alongside, same era.
“Learning words from images and speech,”
Gabriel Synnaeve, Maarten Versteegh, and Emmanuel Dupoux, · 2014
Cited alongside, same era.
“Learning deep features for scene recognition using places database,”
Bolei Zhou, Agata Lapedriza, Jianxiong Xiao, Antonio Torralba, and Aude Oliva, · 2014
Cited alongside, same era.
“Basic cuts revisited: Temporal segmentation of speech into phone-like units with statistical learning at a pre-linguistic level.,”
Okko Räsänen, · 2014
Cited alongside, same era.
“Unsupervised lexicon discovery from acoustic input,”
Chia-ying Lee, Timothy O’Donnell, and James Glass, · 2015
Cited alongside, same era.
“Deep multimodal semantic embeddings for speech and images,”
David Harwath and James Glass, · 2015
“Visually grounded learning of keyword prediction from untranscribed speech,”
Herman Kamper, Shane Settle, Gregory Shakhnarovich, and Karen Livescu, · 2017
Later among the works it cites.
“Representations of language in a model of visually grounded speech signal,”
Grzegorz Chrupala, Lieke Gelderloos, and Afra Alishahi, · 2017
Later among the works it cites.
“Analysis of audio-visual features for unsupervised speech recognition,”
Jennifer Drexler and James Glass, · 2017
Later among the works it cites.
“Encoding of phonology in a recurrent neural model of grounded speech,”
Afra Alishahi, Marie Barking, and Grzegorz Chrupala, · 2017
Later among the works it cites.
“Blind phoneme segmentation with temporal prediction errors,”
Paul Michel, Okko Räsänen, Roland Thiolliére, and Emmanuel Dupoux, · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“Variational inference for acoustic unit discovery,”
Lucas Ondel, Lukaš Burget, and Jan Černocký, · 2016
Cited alongside, same era.
“Unsupervised learning of spoken language with visual context,”
David Harwath, Antonio Torralba, and James R. Glass, · 2016
Cited alongside, same era.
“A segmental framework for fully-unsupervised large-vocabulary speech recognition,”
Herman Kamper, Aren Jansen, and Sharon Goldwater, · 2017
Cited alongside, same era.
“Learning word-like units from joint audio-visual analysis,”
David Harwath and James Glass, · 2017
Cited alongside, same era.
Odette Scharenborg, Laurent Besacier, Alan W. Black, Mark Hasegawa-Johnson, Florian Metze, Graham Neubig, Sebastian Stüker, Pierre Godard, Markus Müller, Lucas Ondel, Shruti Palaskar, Philip Arthur, Francesco Ciannella, Mingxing Du, Elin Larsen, Danny Merkx, Rachid Riad, Liming Wang, and Emmanuel Dupoux, · 2018
Later among the works it cites.
“Vision as an interlingua: Learning multilingual semantic embeddings of untranscribed speech,”
David Harwath, Galen Chuang, and James Glass, · 2018
Later among the works it cites.
“Jointly discovering visual objects and spoken words from raw sensory input,”
David Harwath, Adrià Recasens, Dídac Surís, Galen Chuang, Antonio Torralba, and James Glass, · 2018
Later among the works it cites.
“Facenet: A unified embedding for face recognition and clustering,”
Florian Schroff, Dmitry Kalenichenko, and James Philbin, · 2018
Later among the works it cites.
“Unsupervised learning of semantic audio representations,”
Aren Jansen, Manoj Plakal, Ratheet Pandya, Daniel PW Ellis, Shawn Hershey, Jiayang Liu, R Channing Moore, and Rif A Saurous, · 2018
Later among the works it cites.