Fetching the paper…
Reading the bibliography…
In this work, we focus on multilingual systems based on recurrent neural networks (RNNs), trained using the Connectionist Temporal Classification (CTC) loss function.
“The Meta-Pi network: Building distributed knowledge representations for robust multisource pattern recognitio,”
John B Hampshire and Alex Waibel, · 1992
Earlier work this paper cites.
“Fast bootstrapping of LVCSR systems with multilingual phoneme sets,”
Tanja Schultz and Alex Waibel, · 1997
Earlier work this paper cites.
“Grundfrequenzverfolgung und deren anwendung in der spracherkennung,”
Kjell Schubert, · 1999
Earlier work this paper cites.
“Polyphone decision tree specialization for language adaptation,”
Tanja Schultz and Alex Waibel, · 2000
Earlier work this paper cites.
“The German text-to-speech synthesis system MARY: A tool for research, development and teaching,”
Marc Schröder and Jürgen Trouvain, · 2003
Earlier work this paper cites.
“Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,”
Alex Graves, Santiago Fernández, Faustino Gomez, and Jürgen Schmidhuber, · 2006
Earlier work this paper cites.
“The fundamental frequency variation spectrum,”
Kornel Laskowski, Mattias Heldner, and Jens Edlund, · 2008
Earlier work this paper cites.
Acoustic modelling for under-resourced languages
Sebastian Stüker, · 2009
Earlier work this paper cites.
“Unsupervised cross-lingual knowledge transfer in DNN-based LVCSR,”
Pawel Swietojanski, Arnab Ghoshal, and Steve Renals, · 2012
Cited alongside, same era.
“The language-independent bottleneck features,”
Karel Vesely, Martin Karafiat, Frantisek Grezl, Milos Janda, and Ekaterina Egorova, · 2012
Cited alongside, same era.
“Improving neural networks by preventing co-adaptation of feature detectors,”
Geoffrey E Hinton, Nitish Srivastava, Alex Krizhevsky, Ilya Sutskever, and Ruslan R Salakhutdinov, · 2012
Cited alongside, same era.
“Speaker adaptation of neural network acoustic models using i-Vectors,”
George Saon, Hagen Soltau, David Nahamoo, and Michael Picheny, · 2013
Cited alongside, same era.
“On the importance of initialization and momentum in deep learning,”
Ilya Sutskever, James Martens, George Dahl, and Geoffrey Hinton, · 2013
Cited alongside, same era.
“Joint acoustic modeling of triphones and trigraphemes by multi-task learning deep neural networks for low-resource speech recognition,”
Dongpeng Chen, Brian Mak, Cheung-Chi Leung, and Sunil Sivadas, · 2014
Later among the works it cites.
“Euronews: A multilingual benchmark for ASR and LID,”
Roberto Gretter, · 2014
Later among the works it cites.
“An investigation of augmenting speaker representations to improve speaker normalisation for dnn-based speech recognition,”
Hengguan Huang and Khe Chai Sim, · 2015
Later among the works it cites.
“Using language adaptive deep neural networks for improved multilingual speech recognition,”
Markus Müller and Alex Waibel, · 2015
Later among the works it cites.
“Language adaptive DNNs for improved low resource speech recognition,”
Markus Müller, Sebastian Stüker, and Alex Waibel, · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Adaptation of multilingual stacked bottle-neck neural network structure for new language,”
Frantisek Grézl, Martin Karafiát, and Karel Vesely, · 2014
Cited alongside, same era.
“Towards speaker adaptive training of deep neural network acoustic models,”
Yajie Miao, Hao Zhang, and Florian Metze, · 2014
Cited alongside, same era.
Hagen Soltau, Hank Liao, and Hasim Sak, · 2016
Later among the works it cites.
“Comparison of decoding strategies for CTC acoustic models,”
Thomas Zenkel, Ramon Sanabria, Florian Metze, Jan Niehues, Matthias Sperber, Sebastian Stüker, and Alex Waibel, · 2017
Closest in time.