Fetching the paper…
Reading the bibliography…
Training automatic speech recognition (ASR) systems requires large amounts of data in the target language in order to achieve good performance.
B. Wheatley, K. Kondo, W. Anderson, and Y. Muthusamy, “An evaluation of cross-language adaptation for rapid hmm development in a new language,” in Acoustics, Speech, and Signal Processing, 1994. ICASSP-94., 1994 IEEE International Conference on , vol. 1. IEEE, 1994, pp. I–237
1994
Earlier work this paper cites.
M. W. et al., “JANUS 93: Towards Spontaneous Speech Translation,” in International Conference on Acoustics, Speech, and Signal Processing 1994 , Adelaide, Australia, 1994
1994
Earlier work this paper cites.
T. Schultz and A. Waibel, “Fast bootstrapping of lvcsr systems with multilingual phoneme sets.” in Eurospeech , 1997
1997
Earlier work this paper cites.
R. Caruana, “Multitask learning,” Machine learning , vol. 28, no. 1, pp. 41–75, 1997
1997
Earlier work this paper cites.
K. Schubert, “Grundfrequenzverfolgung und deren Anwendung in der Spracherkennung,” Master’s thesis, Universität Karlsruhe (TH), Germany, 1999, in German
1999
Earlier work this paper cites.
C. Schillo, G. A. Fink, and F. Kummert, “Grapheme based speech recognition for large vocabularies,” in Proceedings of the Sixth International Conference on Spoken Language Processing (ICSLP 2000) . Beijing, China: ISCA, October 2000, pp. 584–587
2000
Earlier work this paper cites.
——, “Polyphone decision tree specialization for language adaptation,” in Acoustics, Speech, and Signal Processing, 2000. ICASSP’00. Proceedings. 2000 IEEE International Conference on , vol. 3. IEEE, 2000, pp. 1707–1710
2000
Earlier work this paper cites.
T. Schultz and A. Waibel, “Language-independent and language-adaptive acoustic modeling for speech recognition,” Speech Communication , vol. 35, no. 1, pp. 31–51, 2001
2001
Earlier work this paper cites.
H. Soltau, F. Metze, C. Fugen, and A. Waibel, “A One-Pass Decoder Based on Polymorphic Linguistic Context Assignment,” in Automatic Speech Recognition and Understanding, 2001. ASRU’01. IEEE Workshop on . IEEE, 2001, pp. 214–217
2001
Earlier work this paper cites.
S. Kanthak and H. Ney, “Context-dependent acoustic modeling using graphemes for large vocabulary speech recognition,” in Proceedings the 2002 IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP’02) , vol. 1. Orlando, Florida, USA: IEEE, 2002, pp. 845–848
2002
Earlier work this paper cites.
M. Killer, S. Stüker, and T. Schultz, “Grapheme based speech recognition,” in Proceedings of the 8th European Conference on Speech Communication and Technology EUROSPEECH’03 . Geneva, Switzerland: ISCA, September 2003, pp. 3141–3144
2003
Earlier work this paper cites.
S. Kanthak and H. Ney, “Multilingual acoustic modeling using graphems,” in Proceedings of the 8th European Conference on Speech Communication and Technology EUROSPEECH’03 . Geneva, Switzerland: ISCA, September 2003, pp. 1145–1148
2003
Earlier work this paper cites.
M. Schröder and J. Trouvain, “The German text-to-speech synthesis system MARY: A tool for research, development and teaching,” International Journal of Speech Technology , vol. 6, no. 4, pp. 365–377, 2003
2003
Earlier work this paper cites.
A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber, “Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,” in Proceedings of the 23rd international conference on Machine learning . ACM, 2006, pp. 369–376
2006
Earlier work this paper cites.
M. Bisani and H. Ney, “Joint-sequence models for grapheme-to-phoneme conversion,” Speech communication , vol. 50, no. 5, pp. 434–451, 2008
2008
Earlier work this paper cites.
S. Stüker, “Modified polyphone decision tree specialization for porting multilingual grapheme based asr systems to new languages,” in Proceedings of the 2008 IEEE International Conference on Acoustics, Speech, and Signal Processing . Las Vegas, NV, USA: IEEE, April 2008, pp. 4249–4252
2008
Earlier work this paper cites.
——, “Integrating thai grapheme based acoustic models into the ml-mix framework - for language independent and cross-language asr,” in Proceedings of the First International Workshop on Spoken Languages Technologies for Under-resourced languages (SLTU) , Hanoi, Vietnam, May 2008
2008
Earlier work this paper cites.
S. Scanzio, P. Laface, L. Fissore, R. Gemello, and F. Mana, “On the use of a multilingual neural network front-end,” in Proceedings of the Interspeech , 2008, pp. 2711–2714
2008
Cited alongside, same era.
K. Laskowski, M. Heldner, and J. Edlund, “The Fundamental Frequency Variation Spectrum,” in Proceedings of the 21st Swedish Phonetics Conference (Fonetik 2008) , Gothenburg, Sweden, June 2008, pp. 29–32
2008
Cited alongside, same era.
S. Stüker, “Acoustic modelling for under-resourced languages,” Ph.D. dissertation, Karlsruhe, Univ., Diss., 2009, 2009
2009
Cited alongside, same era.
J. R. Novak, D. Yang, N. Minematsu, and K. Hirose, “Phonetisaurus: A wfst-driven phoneticizer,” The University of Tokyo, Tokyo Institute of Technology , pp. 221–222, 2011
2011
Cited alongside, same era.
P. Swietojanski, A. Ghoshal, and S. Renals, “Unsupervised cross-lingual knowledge transfer in DNN-based LVCSR,” in SLT , IEEE. IEEE, 2012, pp. 246–251
M. Müller and A. Waibel, “Using Language Adaptive Deep Neural Networks for Improved Multilingual Speech Recognition,” IWSLT , 2015
2015
Later among the works it cites.
2015
Later among the works it cites.
2016
Later among the works it cites.
M. Müller, S. Stüker, and A. Waibel, “Language Adaptive DNNs for Improved Low Resource Speech Recognition,” in Interspeech , 2016
2016
Later among the works it cites.
——, “Language Feature Vectors for Resource Constraint Speech Recognition,” in Speech Communication; 12. ITG Symposium; Proceedings of . VDE, 2016
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2012
Cited alongside, same era.
K. Vesely, M. Karafiat, F. Grezl, M. Janda, and E. Egorova, “The language-independent bottleneck features,” in Proceedings of the Spoken Language Technology Workshop (SLT), 2012 IEEE . IEEE, 2012, pp. 336–341
2012
Cited alongside, same era.
A. Ghoshal, P. Swietojanski, and S. Renals, “Multilingual training of Deep-Neural networks,” in Proceedings of the ICASSP , Vancouver, Canada, 2013
2013
Cited alongside, same era.
G. Heigold, V. Vanhoucke, A. Senior, P. Nguyen, M. Ranzato, M. Devin, and J. Dean, “Multilingual Acoustic Models Using Distributed Deep Neural Networks,” in Proceedings of the ICASSP , Vancouver, Canada, May 2013
2013
Cited alongside, same era.
G. Saon, H. Soltau, D. Nahamoo, and M. Picheny, “Speaker Adaptation of Neural Network Acoustic Models Using i-Vectors,” in ASRU . IEEE, 2013, pp. 55–59
2013
Cited alongside, same era.
F. Metze, Z. Sheikh, A. Waibel, J. Gehring, K. Kilgour, Q. B. Nguyen, V. H. Nguyen, et al. , “Models of Tone for Tonal and Non-tonal Languages,” in Automatic Speech Recognition and Understanding (ASRU), 2013 IEEE Workshop on . IEEE, 2013, pp. 261–266
2013
Cited alongside, same era.
I. Sutskever, J. Martens, G. Dahl, and G. Hinton, “On the importance of initialization and momentum in deep learning,” in Proceedings of the 30th International Conference on Machine Learning (ICML-13) , 2013, pp. 1139–1147
2013
Cited alongside, same era.
F. Grézl, M. Karafiát, and K. Vesely, “Adaptation of multilingual stacked bottle-neck neural network structure for new language,” in Acoustics, Speech and Signal Processing (ICASSP), 2014 IEEE International Conference on . IEEE, 2014, pp. 7654–7658
2014
Cited alongside, same era.
2016
Later among the works it cites.
2016
Later among the works it cites.
2016
Later among the works it cites.
Y. Miao, M. Gowayyed, X. Na, T. Ko, F. Metze, and A. Waibel, “An empirical exploration of ctc acoustic models,” in Acoustics, Speech and Signal Processing (ICASSP), 2016 IEEE International Conference on . IEEE, 2016, pp. 2623–2627
2016
Later among the works it cites.
D. Amodei, S. Ananthanarayanan, R. Anubhai, J. Bai, E. Battenberg, C. Case, J. Casper, B. Catanzaro, Q. Cheng, G. Chen, et al. , “Deep speech 2: End-to-end speech recognition in english and mandarin,” in International Conference on Machine Learning , 2016, pp. 173–182
2016
Later among the works it cites.
——, “The microsoft 2016 conversational speech recognition system,” in Acoustics, Speech and Signal Processing (ICASSP), 2017 IEEE International Conference on . IEEE, 2017, pp. 5255–5259
2017
Closest in time.
M. Müller, S. Stüker, and A. Waibel, “Multilingual ctc speech recognition,” in SPECOM , 2017
2017
Closest in time.
2017
Closest in time.
H. Sak and K. Rao, “Multi-accent speech recognition with hierarchical grapheme based models,” 2017
2017
Closest in time.
“PyTorch,” http://pytorch.org, accessed: 2017-04-13
2017
Closest in time.
“warp-ctc,” https://github.com/baidu-research/warp-ctc, accessed: 2017-04-13
2017
Closest in time.