Fetching the paper…
Reading the bibliography…
We present a method for cross-lingual training an ASR system using absolutely no transcribed training data from the target language, and with no phonetic knowledge of the language in question.
L. E. Baum, T. Petrie, G. Soules, and N. Weiss, “A maximization technique occurring in the statistical analysis of probabilistic functions of Markov chains,” The annals of mathematical statistics , vol. 41, no. 1, pp. 164–171, 1970
1970
Earlier work this paper cites.
K. Knight, “Decoding complexity in word-replacement translation models,” Computational linguistics , vol. 25, no. 4, pp. 607–615, 1999
1999
Earlier work this paper cites.
L. Lamel, J.-L. Gauvain, and G. Adda, “Lightly supervised and unsupervised acoustic model training,” Computer Speech and Language , vol. 16, 2002
2002
Earlier work this paper cites.
——, “Unsupervised acoustic model training,” in ICASSP , 2002
2002
Earlier work this paper cites.
A. Stolcke, “SRILM-an extensible language modeling toolkit,” in ICSLP , 2002
2002
Earlier work this paper cites.
C. Allauzen, M. Riley, J. Schalkwyk, W. Skut, and M. Mohri, “OpenFst: A general and efficient weighted finite-state transducer library,” in CIAA , 2007
2007
Earlier work this paper cites.
C. Allauzen and M. Mohri, “3-way composition of weighted finite-state transducers,” in CIAA , 2008
2008
Earlier work this paper cites.
S. Ravi and K. Knight, “Learning phoneme mappings for transliteration without parallel data,” in NAACL-HLT , 2009
2009
Earlier work this paper cites.
N. T. Vu, F. Kraus, and T. Schultz, “Cross-language bootstrapping based on completely unsupervised training using multilingual a-stabil,” in ICASSP , 2011
2011
Earlier work this paper cites.
S. Ravi and K. Knight, “Deciphering foreign language,” in ACL-HLT , 2011
2011
Earlier work this paper cites.
D. Povey, A. Ghoshal, G. Boulianne, L. Burget, O. Glembek, N. Goel, M. Hannemann, P. Motlicek, Y. Qian, P. Schwarz et al. , “The Kaldi speech recognition toolkit,” in ASRU , 2011
2011
Earlier work this paper cites.
J. Glass, “Towards unsupervised speech processing,” in ISSPA , 2012
2012
Earlier work this paper cites.
B. Roark, R. Sproat, C. Allauzen, M. Riley, J. Sorensen, and T. Tai, “The OpenGrm open-source finite-state grammar software libraries,” in ACL , 2012
2012
Earlier work this paper cites.
H. Gelas, L. Besacier, and F. Pellegrino, “Developments of Swahili resources for an automatic speech recognition system,” in SLTU , 2012
2012
Earlier work this paper cites.
T. Schultz, N. T. Vu, and T. Schlippe, “Globalphone: A multilingual text & speech database in 20 languages,” in ICASSP , 2013
2013
Earlier work this paper cites.
T. Berg-Kirkpatrick and D. Klein, “Decipherment with a million random restarts,” in EMNLP , 2013
2013
Cited alongside, same era.
V. Hai, X. Xiao, E. S. Chng, and H. Li, “Cross-lingual phone mapping for large vocabulary speech recognition of under-resourced languages,” IEICE TRANSACTIONS on Information and Systems , vol. 97, no. 2, pp. 285–295, 2014
2014
Cited alongside, same era.
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, “Generative adversarial nets,” in NeurIPS , 2014
2014
Cited alongside, same era.
M. Nuhn and H. Ney, “EM decipherment for large vocabularies,” in ACL , 2014
2014
Cited alongside, same era.
C. Buck, K. Heafield, and B. van Ooyen, “N-gram Counts and Language Models from the Common Crawl,” in LREC , 2014
2014
Cited alongside, same era.
J. Fainberg, O. Klejch, S. Renals, and P. Bell, “Lattice-Based Lightly-Supervised Acoustic Model Training,” in Interspeech , 2019
2019
Later among the works it cites.
M. Prasad, D. van Esch, S. Ritchie, and J. F. Mortensen, “Building large-vocabulary asr systems for languages without any audio training data.” in Interspeech , 2019
2019
Later among the works it cites.
A. Carmantini, P. Bell, and S. Renals, “Untranscribed web audio for low resource speech recognition.” in Interspeech , 2019
2019
Later among the works it cites.
C. Liu, Q. Zhang, X. Zhang, K. Singh, Y. Saraf, and G. Zweig, “Multilingual graphemic hybrid asr with massive data augmentation,” in SLTU-CCURL , 2020
2020
Later among the works it cites.
J. L. Lee, L. F. Ashby, M. E. Garza, Y. Lee-Sikka, S. Miller, A. Wong, A. D. McCarthy, and K. Gorman, “Massively multilingual pronunciation modeling with WikiPron,” in LREC , 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: an ASR corpus based on public domain audio books,” in ICASSP , 2015
2015
Cited alongside, same era.
W. Chan, N. Jaitly, Q. Le, and O. Vinyals, “Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,” in ICASSP , 2016
2016
Cited alongside, same era.
D. Povey, V. Peddinti, D. Galvez, P. Ghahremani, V. Manohar, X. Na, Y. Wang, and S. Khudanpur, “Purely sequence-trained neural networks for ASR based on lattice-free MMI,” in Interspeech , 2016
2016
Cited alongside, same era.
H. Kamper, A. Jansen, and S. Goldwater, “A segmental framework for fully-unsupervised large-vocabulary speech recognition,” Computer Speech and Language , vol. 46, 2017
2017
Cited alongside, same era.
T. Shinozaki, S. Watanabe, D. Mochihashi, and G. Neubig, “Semi-supervised learning of a pronunciation dictionary from disjoint phonemic transcripts and text,” in Interspeech , 2017
2017
Cited alongside, same era.
V. Manohar, H. Hadian, D. Povey, and S. Khudanpur, “Semi-supervised training of acoustic models using lattice-free MMI,” in ICASSP , 2018
2018
Cited alongside, same era.
H. Hadian, H. Sameti, D. Povey, and S. Khudanpur, “Flat-start single-stage discriminatively trained HMM-based models for ASR,” IEEE/ACM TASLP , vol. 26, no. 11, pp. 1949–1961, 2018
2018
Cited alongside, same era.
2020
Later among the works it cites.
X. Li, S. Dalmia, J. Li, M. Lee, P. Littell, J. Yao, A. Anastasopoulos, D. R. Mortensen, G. Neubig, A. W. Black et al. , “Universal phone recognition with a multilingual allophone system,” in ICASSP , 2020
2020
Later among the works it cites.
A. Baevski, Y. Zhou, A. Mohamed, and M. Auli, “wav2vec 2.0: A framework for self-supervised learning of speech representations,” NeurIPS , 2020
2020
Later among the works it cites.
C. Chu, S. Fang, and K. Knight, “Learning to pronounce chinese without a pronunciation dictionary,” in EMNLP , 2020
2020
Later among the works it cites.
E. Hermann, H. Kamper, and S. Goldwater, “Multilingual and unsupervised subword modeling for zero-resource languages,” Computer Speech and Language , vol. 65, 2021
2021
Closest in time.
X. Li, J. Li, F. Metze, and A. W. Black, “Hierarchical phone recognition with compositional phonetics,” Interspeech , 2021
2021
Closest in time.
S. Feng, P. Żelasko, L. Moro-Velázquez, A. Abavisani, M. Hasegawa-Johnson, O. Scharenborg, and N. Dehak, “How phonotactics affect multilingual and zero-shot asr performance,” in ICASSP , 2021
2021
Closest in time.
H. Gao, J. Ni, Y. Zhang, K. Qian, S. Chang, and M. Hasegawa-Johnson, “Zero-shot cross-lingual phonetic recognition with external language embedding,” Interspeech , 2021
2021
Closest in time.
S. Khare, A. Mittal, A. Diwan, S. Sarawagi, P. Jyothi, and B. Samarth, “Low resource ASR: The surprising effectiveness of high resource transliteration,” in Interspeech , 2021
2021
Closest in time.
2021
Closest in time.
E. Wallington, B. Kershenbaum, O. Klejch, and P. Bell, “On the learning dynamics of semi-supervised training for ASR,” Interspeech , 2021
2021
Closest in time.