Fetching the paper…
Reading the bibliography…
Producing a large amount of annotated speech data for training ASR systems remains difficult for more than 95% of languages all over the world which are low-resourced.
“A spe based distinctive feature composition of the cmu label set in the timit database,”
Tom Brøndsted, · 1998
Earlier work this paper cites.
“Unsupervised speech segmentation: An analysis of the hypothesized phone boundaries,”
Odette Scharenborg, Vincent Wan, and Mirjam Ernestus, · 2010
Earlier work this paper cites.
“Semi-supervised training of deep neural networks,”
Karel Vesely, Mirko Hannemann, and Lukas Burget, · 2013
Earlier work this paper cites.
“Deep neural network features and semi-supervised training for low resource speech recognition,”
Samuel Thomas, Michael L Seltzer, Kenneth Church, and Hynek Hermansky, · 2013
Earlier work this paper cites.
“Distributed representations of words and phrases and their compositionality,”
Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg S Corrado, and Jeff Dean, · 2013
Earlier work this paper cites.
“Combination of multilingual and semi-supervised training for under-resourced languages,”
František Grézl and Martin Karafiát, · 2014
Earlier work this paper cites.
“Word embeddings for speech recognition,”
Samy Bengio and Georg Heigold, · 2014
Earlier work this paper cites.
“Segmental acoustic indexing for zero resource keyword search,”
Keith Levin, Aren Jansen, and Benjamin Van Durme, · 2015
Earlier work this paper cites.
“Query-by-example keyword spotting using long short-term memory networks,”
Guoguo Chen, Carolina Parada, and Tara N. Sainath, · 2015
Cited alongside, same era.
“Librispeech: an asr corpus based on public domain audio books,”
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur, · 2015
Cited alongside, same era.
“Semi-supervised and unsupervised discriminative language model training for automatic speech recognition,”
Erinç Dikici and Murat Saraçlar, · 2016
Cited alongside, same era.
Yu-An Chung, Chao-Chung Wu, Chia-Hao Shen, Hung-Yi Lee, and Lin-Shan Lee, · 2016
Cited alongside, same era.
“Semi-supervised dnn training with word selection for asr.,”
Karel Veselỳ, Lukás Burget, and Jan Cernockỳ, · 2017
Cited alongside, same era.
“Towards learning semantic audio representations from unlabeled data,”
Aren Jansen, Manoj Plakal, Ratheet Pandya, Dan Ellis, Shawn Hershey, Jiayang Liu, Channing Moore, and Rif A. Saurous, · 2017
Later among the works it cites.
“Semi-supervised end-to-end speech recognition,”
Shigeki Karita, Shinji Watanabe, Tomoharu Iwata, Atsunori Ogawa, and Marc Delcroix, · 2018
Closest in time.
Da-Rong Liu, Kuan-Yu Chen, Hung-Yi Lee, and Lin-shan Lee, · 2018
Closest in time.
“Towards unsupervised automatic speech recognition trained by unaligned speech and text only,”
Yi-Chen Chen, Chia-Hao Shen, Sung-Feng Huang, and Hung-yi Lee, · 2018
Closest in time.
“Phonetic-and-semantic embedding of spoken words with applications in spoken content retrieval,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Segmental audio Word2Vec: Representing utterances as sequences of vectors with applications in spoken term detection,”
Yu-Hsuan Wang, Hung-Yi Lee, and Lin-Shan Lee, · 2017
Cited alongside, same era.
“Query-by-example search with discriminative neural acoustic word embeddings,”
Shane Settle, Keith Levin, Herman Kamper, and Karen Livescu, · 2017
Cited alongside, same era.
Yi-Chen Chen, Sung-Feng Huang, Chia-Hao Shen, Hung-yi Lee, and Lin-shan Lee, · 2018
Closest in time.
“Unsupervised cross-modal alignment of speech and text embedding spaces,”
Yu-An Chung, Wei-Hung Weng, Schrasing Tong, and James Glass, · 2018
Closest in time.
“An iterative closest point method for unsupervised word translation,”
Yedid Hoshen and Lior Wolf, · 2018
Closest in time.