Fetching the paper…
Reading the bibliography…
End-to-end multilingual speech recognition involves using a single model training on a compositional speech corpus including many languages, resulting in a single neural network to handle transcribing different languages.
J. B. Hampshire II and A. Waibel, “The meta-pi network: Building distributed knowledge representations for robust multisource pattern recognition,”
1992
Earlier work this paper cites.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,”
1997
Earlier work this paper cites.
A. Waibel, H. Soltau, T. Schultz, T. Schaaf, and F. Metze,
2000
Earlier work this paper cites.
H. Lin, L. Deng, D. Yu, Y.-f. Gong, A. Acero, and C.-H. Lee, “A study on multilingual acoustic modeling for large vocabulary asr,” in
2009
Earlier work this paper cites.
L. Burget, P. Schwarz, M. Agarwal, P. Akyazi, K. Feng, A. Ghoshal, O. Glembek, N. Goel, M. Karafiát, D. Povey
2010
Earlier work this paper cites.
G. Heigold, V. Vanhoucke, A. Senior, P. Nguyen, M. Ranzato, M. Devin, and J. Dean, “Multilingual acoustic models using distributed deep neural networks,” in
2013
Earlier work this paper cites.
D. Bahdanau, K. Cho, and Y. Bengio, “Neural machine translation by jointly learning to align and translate,”
2014
Earlier work this paper cites.
R. Gretter, “Euronews: a multilingual speech corpus for asr.” in
2014
Earlier work this paper cites.
2015
Earlier work this paper cites.
2016
Earlier work this paper cites.
D. Bahdanau, J. Chorowski, D. Serdyuk, P. Brakel, and Y. Bengio, “End-to-end attention-based large vocabulary speech recognition,” in
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
J. L. Ba, J. R. Kiros, and G. E. Hinton, “Layer normalization,”
2016
Earlier work this paper cites.
M. Johnson, M. Schuster, Q. V. Le, M. Krikun, Y. Wu, Z. Chen, N. Thorat, F. Viégas, M. Wattenberg, G. Corrado
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in
2017
Cited alongside, same era.
J. Gehring, M. Auli, D. Grangier, D. Yarats, and Y. N. Dauphin, “Convolutional sequence to sequence learning,” in
2017
Cited alongside, same era.
Y. Zhang, W. Chan, and N. Jaitly, “Very deep convolutional networks for end-to-end speech recognition,” in
2017
Cited alongside, same era.
E. A. Platanios, M. Sachan, G. Neubig, and T. Mitchell, “Contextual parameter generation for universal neural machine translation,” in
2018
Cited alongside, same era.
M. Müller, S. Stüker, and A. Waibel, “Neural language codes for multilingual acoustic models,”
2018
Cited alongside, same era.
M. Müller, S. Stüker, and A. Waibel, “Neural codes to factor language in multilingual speech recognition,” in
2019
Later among the works it cites.
2019
Later among the works it cites.
A. Zeyer, P. Bahar, K. Irie, R. Schlüter, and H. Ney, “A comparison of transformer and lstm encoder decoder models for asr,” in
2019
Later among the works it cites.
D. S. Park, W. Chan, Y. Zhang, C.-C. Chiu, B. Zoph, E. D. Cubuk, and Q. V. Le, “SpecAugment: A simple data augmentation method for automatic speech recognition,”
2019
Later among the works it cites.
Y. Zhu, P. Haghani, A. Tripathi, B. Ramabhadran, B. Farris, H. Xu, H. Lu, H. Sak, I. Leal, N. Gaur, P. J. Moreno, and Q. Zhang, “Multilingual Speech Recognition with Self-Attention Structured Parameterization,” in
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
S. Kim and M. L. Seltzer, “Towards language-universal end-to-end speech recognition,” in
2018
Cited alongside, same era.
S. Toshniwal, T. N. Sainath, R. J. Weiss, B. Li, P. Moreno, E. Weinstein, and K. Rao, “Multilingual speech recognition with a single end-to-end model,” in
2018
Cited alongside, same era.
2018
Cited alongside, same era.
2019
Cited alongside, same era.
N.-Q. Pham, T.-S. Nguyen, J. Niehues, M. Müller, and A. Waibel, “Very Deep Self-Attention Networks for End-to-End Speech Recognition,” in
2019
Cited alongside, same era.
2019
Cited alongside, same era.
O. Adams, M. Wiesner, S. Watanabe, and D. Yarowsky, “Massively multilingual adversarial speech recognition,” in
2019
Cited alongside, same era.
2020
Later among the works it cites.
Y. Wen, D. Tran, and J. Ba, “Batchensemble: an alternative approach to efficient ensemble and lifelong learning,”
2020
Later among the works it cites.
B. Zhang, P. Williams, I. Titov, and R. Sennrich, “Improving massively multilingual neural machine translation and zero-shot translation,” in
2020
Later among the works it cites.
V. Pratap, A. Sriram, P. Tomasello, A. Hannun, V. Liptchinsky, G. Synnaeve, and R. Collobert, “Massively multilingual asr: 50 languages, 1 model, 1 billion parameters,”
2020
Later among the works it cites.
J. Philip, A. Berard, M. Gallé, and L. Besacier, “Language adapters for zero shot neural machine translation,” in
2020
Later among the works it cites.
M. Dusenberry, G. Jerfel, Y. Wen, Y. Ma, J. Snoek, K. Heller, B. Lakshminarayanan, and D. Tran, “Efficient and scalable bayesian neural nets with rank-1 factors,” in
2020
Later among the works it cites.
J. Iranzo-Sánchez, J. A. Silvestre-Cerdà, J. Jorge, N. Roselló, A. Giménez, A. Sanchis, J. Civera, and A. Juan, “Europarl-st: A multilingual corpus for speech translation of parliamentary debates,” in
2020
Later among the works it cites.
N.-Q. Pham, T.-L. Ha, T.-N. Nguyen, T.-S. Nguyen, E. Salesky, S. Stüker, J. Niehues, and A. Waibel, “Relative Positional Encoding for Speech Recognition and Direct Translation,” in
2020
Later among the works it cites.
X. Wang, Y. Tsvetkov, and G. Neubig, “Balancing training for multilingual neural machine translation,” in
2020
Later among the works it cites.