Fetching the paper…
Reading the bibliography…
Previous work has shown that for low-resource source languages, automatic speech-to-text translation (AST) can be improved by pretraining an end-to-end model on automatic speech recognition (ASR) data from a high-resource language.
R. J. Williams and D. Zipser, “A learning algorithm for continually running fully recurrent neural networks,” Neural Computation , 1989
1989
Earlier work this paper cites.
J. Godfrey and E. Holliman, “Switchboard-1 Release 2 (LDC97S62),” 1993, https://catalog.ldc.upenn.edu/LDC97S62
1993
Earlier work this paper cites.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural Computation , 1997
1997
Earlier work this paper cites.
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu, “BLEU: a method for automatic evaluation of machine translation,” in Proc. ACL , 2002
2002
Earlier work this paper cites.
T. Schultz, “Globalphone: a multilingual speech and text database developed at Karlsruhe University,” in ICSLP , 2002
2002
Earlier work this paper cites.
V. Nair and G. E. Hinton, “Rectified linear units improve restricted Boltzmann machines,” in Proc. ICML , 2010
2010
Earlier work this paper cites.
D. Povey, A. Ghoshal, G. Boulianne, L. Burget, O. Glembek, N. Goel, M. Hannemann, P. Motlicek, Y. Qian, P. Schwarz, J. Silovsky, G. Stemmer, and K. Vesely, “The Kaldi Speech Recognition Toolkit,” in Proc. ASRU , 2011
2011
Earlier work this paper cites.
M. Post, G. Kumar, A. Lopez, D. Karakos, C. Callison-Burch, and S. Khudanpur, “Fisher and CALLHOME Spanish-English Speech Translation,” 2014, https://catalog.ldc.upenn.edu/LDC2014T23
2014
Earlier work this paper cites.
N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov, “Dropout: A simple way to prevent neural networks from overfitting,” J. Mach. Learn. Res. , 2014
2014
Earlier work this paper cites.
T. Ko, V. Peddinti, D. Povey, and S. Khudanpur, “Audio augmentation for speech recognition,” in Proc. Interspeech , 2015
2015
Earlier work this paper cites.
S. Ioffe and C. Szegedy, “Batch normalization: Accelerating deep network training by reducing internal covariate shift,” in Proc. ICML , 2015
2015
Earlier work this paper cites.
M.-T. Luong, H. Pham, and C. D. Manning, “Effective approaches to attention-based neural machine translation,” in Proc. EMNLP , 2015
2015
Earlier work this paper cites.
D. Kingma and J. Ba, “Adam: A method for stochastic optimization,” in Proc. ICLR , 2015
2015
Cited alongside, same era.
A. Bérard, O. Pietquin, C. Servan, and L. Besacier, “Listen and translate: A proof of concept for end-to-end speech-to-text translation,” in NIPS Workshop on end-to-end learning for speech and audio processing , 2016
2016
Cited alongside, same era.
R. Sennrich, B. Haddow, and A. Birch, “Neural machine translation of rare words with subword units,” in Proc. ACL , 2016
2016
Cited alongside, same era.
Y. Gal, “A theoretically grounded application of dropout in recurrent neural networks,” in Proc. NIPS , 2016
2016
Cited alongside, same era.
2016
A. Anastasopoulos and D. Chiang, “Tied multitask learning for neural speech translation,” in Proc. NAACL HLT , 2018
2018
Later among the works it cites.
S. Toshniwal, T. N. Sainath, R. J. Weiss, B. Li, P. Moreno, E. Weinstein, and K. Rao, “Multilingual Speech Recognition with A Single End-To-End Model,” in Proc. ICASSP , 2018
2018
Later among the works it cites.
2018
Later among the works it cites.
J. Cho, M. K. Baskar, R. Li, M. Wiesner, S. Mallidi, N. Yalta, M. Karafiát, S. Watanabe, and T. Hori, “Multilingual sequence-to-sequence speech recognition: Architecture, transfer learning, and language modeling,” in Proc. SLT , 2018
2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
R. J. Weiss, J. Chorowski, N. Jaitly, Y. Wu, and Z. Chen, “Sequence-to-sequence models can directly transcribe foreign speech,” in Proc. Interspeech , 2017
2017
Cited alongside, same era.
H. Bu, J. Du, X. Na, B. Wu, and H. Zheng, “Aishell-1: An open-source Mandarin speech corpus and a speech recognition baseline,” in O-COCOSDA , 2017
2017
Cited alongside, same era.
Y. Belinkov and J. R. Glass, “Analyzing hidden representations in end-to-end automatic speech recognition systems,” in NIPS , 2017
2017
Cited alongside, same era.
2017
Cited alongside, same era.
S. Bansal, H. Kamper, K. Livescu, A. Lopez, and S. Goldwater, “Low-resource speech-to-text translation,” in Proc. Interspeech , 2018
2018
Cited alongside, same era.
A. Bérard, L. Besacier, A. C. Kocabiyikoglu, and O. Pietquin, “End-to-end automatic speech translation of audiobooks,” in Proc. ICASSP , 2018
2018
Cited alongside, same era.
P. Godard, G. Adda, M. Adda-Decker et al. , “A very low resource language speech corpus for computational language documentation experiments,” in Proc. LREC , 2018
2018
Cited alongside, same era.
S. Dalmia, R. Sanabria, F. Metze, and A. W. Black, “Sequence-based multi-lingual low resource speech recognition,” in ICASSP , 2018
2018
Later among the works it cites.
2018
Later among the works it cites.
E. Hermann and S. Goldwater, “Multilingual bottleneck features for subword modeling in zero-resource languages,” in Proc. Interspeech , 2018
2018
Later among the works it cites.
S. Bansal, H. Kamper, K. Livescu, A. Lopez, and S. Goldwater, “Pre-training on high-resource speech recognition improves low-resource speech-to-text translation,” in Proc. NAACL , 2019
2019
Closest in time.
M. Sperber, G. Neubig, J. Niehues, and A. Waibel, “Attention-passing models for robust and data-efficient end-to-end speech translation,” in Trans. ACL , 2019
2019
Closest in time.
E. Salesky, M. Sperber, and A. Waibel, “Fluent translations from disfluent speech in end-to-end speech translation,” in Proc. NAACL , 2019
2019
Closest in time.
Y. Tian, “How does pre-training improve low-resource speech-to-text translation? — a case study on a Swahili-English dataset,” Master’s thesis, University of Edinburgh, 2019
2019
Closest in time.
O. Adams, M. Wiesner, S. Watanabe, and D. Yarowsky, “Massively multilingual adversarial speech recognition,” in Proc. NAACL , 2019
2019
Closest in time.