Fetching the paper…
Reading the bibliography…
End-to-end Speech Translation (ST) models have many potential advantages when compared to the cascade of Automatic Speech Recognition (ASR) and text Machine Translation (MT) models, including lowered inference latency and the avoidance of error compounding.
“Signal estimation from modified short-time Fourier transform,”
D. Griffin and J. Lim, · 1984
Earlier work this paper cites.
“Speech translation: Coupling of recognition and translation,”
H. Ney, · 1999
Earlier work this paper cites.
“BLEU: A method for automatic evaluation of machine translation,”
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu, · 2002
Earlier work this paper cites.
“On the integration of speech recognition and statistical machine translation,”
E. Matusov, S. Kanthak, and H. Ney, · 2005
Earlier work this paper cites.
“Japanese and Korean voice search,”
M. Schuster and K. Nakajima, · 2012
Earlier work this paper cites.
“Improved speech-to-text translation with the Fisher and Callhome Spanish–English speech translation corpus,”
M. Post, G. Kumar, A. Lopez, D. Karakos, C. Callison-Burch, and S. Khudanpur, · 2013
Earlier work this paper cites.
“Sequence to sequence learning with neural networks,”
I. Sutskever, O. Vinyals, and Q. V. Le, · 2014
Earlier work this paper cites.
“On the properties of neural machine translation: Encoder-decoder approaches,”
K. Cho, B. Van Merriënboer, D. Bahdanau, and Y. Bengio, · 2014
Earlier work this paper cites.
“Neural machine translation by jointly learning to align and translate,”
D. Bahdanau, K. Cho, and Y. Bengio, · 2015
Earlier work this paper cites.
“Distilling the knowledge in a neural network,”
G. Hinton, O. Vinyals, and J. Dean, · 2015
Earlier work this paper cites.
“LibriSpeech: an ASR corpus based on public domain audio books,”
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, · 2015
Earlier work this paper cites.
“Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,”
W. Chan, N. Jaitly, Q. Le, and O. Vinyals, · 2016
Cited alongside, same era.
Y. Wu, M. Schuster, Z. Chen, Q. V. Le, M. Norouzi, W. Macherey, M. Krikun, Y. Cao, Q. Gao, K. Macherey, et al., · 2016
Cited alongside, same era.
“Listen and translate: A proof of concept for end-to-end speech-to-text translation,”
A. Bérard, O. Pietquin, C. Servan, and L. Besacier, · 2016
Cited alongside, same era.
“Improving neural machine translation models with monolingual data,”
R. Sennrich, B. Haddow, and A. Birch, · 2016
Cited alongside, same era.
“Attention is all you need,”
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, · 2017
Cited alongside, same era.
“End-to-end automatic speech translation of audiobooks,”
A. Bérard, L. Besacier, A. C. Kocabiyikoglu, and O. Pietquin, · 2018
Closest in time.
“Pre-training on high-resource speech recognition improves low-resource speech-to-text translation,”
S. Bansal, H. Kamper, K. Livescu, A. Lopez, and S. Goldwater, · 2018
Closest in time.
“Machine speech chain with one-shot speaker adaptation,”
A. Tjandra, S. Sakti, and S. Nakamura, · 2018
Closest in time.
“Multi-modal data augmentation for end-to-end ASR,”
A. Renduchintala, S. Ding, M. Wiesner, and S. Watanabe, · 2018
Closest in time.
“Back-translation-style data augmentation for end-to-end ASR,”
T. Hayashi, S. Watanabe, Y. Zhang, T. Toda, T. Hori, R. Astudillo, and K. Takeda, · 2018
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Sequence-to-sequence models can directly translate foreign speech,”
R. J. Weiss, J. Chorowski, N. Jaitly, Y. Wu, and Z. Chen, · 2017
Cited alongside, same era.
“Structured-based curriculum learning for end-to-end english-japanese speech translation,”
T. Kano, S. Sakti, and S. Nakamura, · 2017
Cited alongside, same era.
“Natural TTS synthesis by conditioning wavenet on mel spectrogram predictions,”
J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. Skerry-Ryan, et al., · 2017
Cited alongside, same era.
“Tacotron: Towards end-to-end speech synthesis,”
Y. Wang, R. Skerry-Ryan, D. Stanton, Y. Wu, R. J. Weiss, N. Jaitly, Z. Yang, Y. Xiao, Z. Chen, S. Bengio, Q. Le, Y. Agiomyrgiannakis, R. Clark, et al., · 2017
Cited alongside, same era.
“State-of-the-art speech recognition with sequence-to-sequence models,”
C.-C. Chiu, T. N. Sainath, Y. Wu, R. Prabhavalkar, P. Nguyen, Z. Chen, A. Kannan, R. J. Weiss, K. Rao, K. Gonina, et al., · 2018
Cited alongside, same era.
“Tied multitask learning for neural speech translation,”
A. Anastasopoulos and D. Chiang, · 2018
Cited alongside, same era.
“The best of both worlds: Combining recent advances in neural machine translation,”
M. X. Chen, O. Firat, A. Bapna, M. Johnson, W. Macherey, G. Foster, L. Jones, N. Parmar, M. Schuster, Z. Chen, Y. Wu, and M. Hughes, · 2018
Closest in time.
“Deep voice 3: 2000-speaker neural text-to-speech,”
W. Ping, K. Peng, A. Gibiansky, S. O. Arik, A. Kannan, S. Narang, J. Raiman, and J. Miller, · 2018
Closest in time.
“Transfer learning from speaker verification to multispeaker text-to-speech synthesis,”
Y. Jia, Y. Zhang, R. J. Weiss, Q. Wang, J. Shen, F. Ren, Z. Chen, P. Nguyen, R. Pang, I. L. Moreno, and Y. Wu, · 2018
Closest in time.
“Towards end-to-end prosody transfer for expressive speech synthesis with Tacotron,”
R. Skerry-Ryan, E. Battenberg, Y. Xiao, Y. Wang, D. Stanton, J. Shor, R. J. Weiss, R. Clark, and R. A. Saurous, · 2018
Closest in time.
“Style tokens: Unsupervised style modeling, control and transfer in end-to-end speech synthesis,”
Y. Wang, D. Stanton, Y. Zhang, R. Skerry-Ryan, E. Battenberg, J. Shor, Y. Xiao, F. Ren, Y. Jia, and R. A. Saurous, · 2018
Closest in time.
“Hierarchical generative modeling for controllable speech synthesis,”
W.-N. Hsu, Y. Zhang, R. J. Weiss, H. Zen, Y. Wu, Y. Wang, Y. Cao, Y. Jia, Z. Chen, J. Shen, P. Nguyen, and R. Pang, · 2019
Closest in time.