Fetching the paper…
Reading the bibliography…
We consider the problem of training speech recognition systems without using any labeled data, under the assumption that the learner can only access to the input utterances and a phoneme language model estimated from a non-overlapping corpus.
Speaker-independent phone recognition using hidden markov models
K-F Lee and H-W Hon · 1989
Earlier work this paper cites.
Practical implementations of speaker-adaptive training
Spyros Matsoukas, Rich Schwartz, Hubert Jin, and Long Nguyen · 1997
Earlier work this paper cites.
Using untranscribed training data to improve performance
George Zavaliagkos, Man-Hung Siu, Thomas Colthurst, and Jayadev Billa · 1998
Earlier work this paper cites.
Unsupervised training of a speech recognizer: Recent experiments
Thomas Kemp and Alex Waibel · 1999
Earlier work this paper cites.
A probabilistic framework for segment-based speech recognition
James R Glass · 2003
Earlier work this paper cites.
Text independent methods for speech segmentation
Anna Esposito and Guido Aversano · 2005
Earlier work this paper cites.
Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks
Alex Graves, Santiago Fernández, Faustino Gomez, and Jürgen Schmidhuber · 2006
Earlier work this paper cites.
On the relation between maximum spectral transition positions and phone boundaries
Sorin Dusan and Lawrence Rabiner · 2006
Earlier work this paper cites.
Unsupervised optimal phoneme segmentation: Objectives, algorithm and comparisons
Yu Qiao, Naoya Shimomura, and Nobuaki Minematsu · 2008
Earlier work this paper cites.
Unsupervised pattern discovery in speech
Alex S Park and James R Glass · 2008
Earlier work this paper cites.
Unsupervised speech segmentation: An analysis of the hypothesized phone boundaries
Odette Scharenborg, Vincent Wan, and Mirjam Ernestus · 2010
Earlier work this paper cites.
The kaldi speech recognition toolkit
Daniel Povey, Arnab Ghoshal, Gilles Boulianne, Lukas Burget, Ondrej Glembek, Nagendra Goel, Mirko Hannemann, Petr Motlicek, Yanmin Qian, Petr Schwarz, et al · 2011
Earlier work this paper cites.
Blind segmentation of speech using non-linear filtering methods
Okko Rasanen, Unto Laine, and Toomas Altosaar · 2011
Earlier work this paper cites.
Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups
Geoffrey Hinton, Li Deng, Dong Yu, George E Dahl, Abdel-rahman Mohamed, Navdeep Jaitly, Andrew Senior, Vincent Vanhoucke, Patrick Nguyen, Tara N Sainath, et al · 2012
Earlier work this paper cites.
Context-dependent pre-trained deep neural networks for large-vocabulary speech recognition
George E Dahl, Dong Yu, Li Deng, and Alex Acero · 2012
Cited alongside, same era.
Sequence transduction with recurrent neural networks
Alex Graves · 2012
Cited alongside, same era.
Active learning for accent adaptation in automatic speech recognition
Udhyakumar Nallasamy, Florian Metze, and Tanja Schultz · 2012
Cited alongside, same era.
A nonparametric bayesian approach to acoustic model discovery
Chia-ying Lee and James Glass · 2012
Cited alongside, same era.
Towards unsupervised speech processing
James Glass · 2012
Cited alongside, same era.
Fast word acquisition in an nmf-based learning framework
Joris Driesen et al · 2012
Blind phoneme segmentation with temporal prediction errors
Paul Michel, Okko Räsänen, Roland Thiolliere, and Emmanuel Dupoux · 2016
Later among the works it cites.
The zero resource speech challenge 2015: Proposed approaches and results
Maarten Versteegh, Xavier Anguera, Aren Jansen, and Emmanuel Dupoux · 2016
Later among the works it cites.
Building speech recognition system from untranscribed data report from jhu workshop 2016
Lukáš Burget, Sanjeev Khudanpur, Najim Dehak, Jan Trmal, Reinhold Haeb-Umbach, Graham Neubig, Shinji Watanabe, Daichi Mochihashi, Takahiro Shinozaki, Ming Sun, et al · 2016
Later among the works it cites.
Variational inference for acoustic unit discovery
Lucas Ondel, Lukáš Burget, and Jan Černockỳ · 2016
Later among the works it cites.
Unsupervised sequence classification using sequential output statistics
Yu Liu, Jianshu Chen, and Li Deng · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Speech recognition with deep recurrent neural networks
Alex Graves, Abdel-rahman Mohamed, and Geoffrey Hinton · 2013
Cited alongside, same era.
A hierarchical system for word discovery exploiting dtw-based initialization
Oliver Walter, Timo Korthals, Reinhold Haeb-Umbach, and Bhiksha Raj · 2013
Cited alongside, same era.
Basic cuts revisited: Temporal segmentation of speech into phone-like units with statistical learning at a pre-linguistic level
Okko Rasanen · 2014
Cited alongside, same era.
Phonetic segmentation of speech signal using local singularity analysis
Vahid Khanagha, Khalid Daoudi, Oriol Pont, and Hussein Yahia · 2014
Cited alongside, same era.
Unsupervised training of an hmm-based self-organizing unit recognizer with applications to topic classification and keyword discovery
Man-hung Siu, Herbert Gish, Arthur Chan, William Belfield, and Steve Lowe · 2014
Cited alongside, same era.
Blind phone segmentation based on spectral change detection using legendre polynomial approximation
Dac-Thang Hoang and Hsiao-Chuan Wang · 2015
Cited alongside, same era.
Yu-Hsuan Wang, Cheng-Tao Chung, and Hung-yi Lee · 2017
Later among the works it cites.
The zero resource speech challenge 2017
Ewan Dunbar, Xuan Nga Cao, Juan Benjumea, Julien Karadayi, Mathieu Bernard, Laurent Besacier, Xavier Anguera, and Emmanuel Dupoux · 2017
Later among the works it cites.
An embedded segmental k-means model for unsupervised segmentation and clustering of speech
Herman Kamper, Karen Livescu, and Sharon Goldwater · 2017
Later among the works it cites.
Bayesian phonotactic language model for acoustic unit discovery
Lucas Ondel, Lukaš Burget, Jan Černockỳ, and Santosh Kesiraju · 2017
Later among the works it cites.
Unsupervised neural machine translation
Mikel Artetxe, Gorka Labaka, Eneko Agirre, and Kyunghyun Cho · 2018
Closest in time.
Unsupervised machine translation using monolingual corpora only
Guillaume Lample, Ludovic Denoyer, and Marc’Aurelio Ranzato · 2018
Closest in time.
Completely unsupervised phoneme recognition by adversarially learning mapping relationships from audio embeddings
Da-Rong Liu, Kuan-Yu Chen, Hung-yi Lee, and Lin-Shan Lee · 2018
Closest in time.
Unsupervised cross-modal alignment of speech and text embedding spaces
Yu-An Chung, Wei-Hung Weng, Schrasing Tong, and James Glass · 2018
Closest in time.
Towards unsupervised automatic speech recognition trained by unaligned speech and text only
Yi-Chen Chen, Chia-Hao Shen, Sung-Feng Huang, and Hung-yi Lee · 2018
Closest in time.