Fetching the paper…
Reading the bibliography…
Attention-based sequence-to-sequence modeling provides a powerful and elegant solution for applications that need to map one sequence to a different sequence.
“Attention-based models for speech recognition,”
J. Chorowski, D. Bahdanau, D. Serdyuk, K. Cho, and Y. Bengio, · 2015
Earlier work this paper cites.
“Librispeech: An ASR corpus based on public domain audio books,”
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, · 2015
Earlier work this paper cites.
“Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,”
W. Chan, N. Jaitly, Q. V. Le, and O. Vinyals, · 2016
Earlier work this paper cites.
“Listen and translate: A proof of concept for end-to-end speech-to-text translation,”
A. Berard, O. Pietquin, C. Servan, and L. Besacier, · 2016
Earlier work this paper cites.
“Towards better decoding and language model integration in sequence to sequence models,”
Jan Chorowski and Navdeep Jaitly, · 2016
Earlier work this paper cites.
“Sequence-to-sequence models can directly translate foreign speech,”
R. Weiss, J. Chorowski, N. Jaitly, Y. Wu, and Z. Chen, · 2017
Earlier work this paper cites.
“A study on data augmentation of reverberant speech for robust speech recognition,”
T. Ko, V. Peddinti, Daniel P., M. Seltzer, and S. Khudanpur, · 2017
Earlier work this paper cites.
“Cold fusion: Training seq2seq models together with language models,”
A. Sriram, H. Jun, S. Satheesh, and A. Coates, · 2017
Earlier work this paper cites.
“Attention is all you need,”
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, · 2017
Earlier work this paper cites.
“Leveraging weakly supervised data to improve end-to-end speech-to-text translation,”
Y. Jia, M. Johnson, W. Macherey, R. Weiss, Y. Cao, C. Chiu, N. Ari, S. Laurenzo, and Y. Wu, · 2018
Earlier work this paper cites.
“Multi-modal data augmentation for end-to-end asr,”
A. Renduchintala, S. Ding, M. Wiesner, and S. Watanabe, · 2018
Cited alongside, same era.
“Tied multitask learning for neural speech translation,”
Antonios Anastasopoulos and David Chiang, · 2018
Cited alongside, same era.
“Unsupervised machine translation using monolingual corpora only,”
G. Lample, L. Denoyer, and M. Ranzato, · 2018
Cited alongside, same era.
“Learning pronunciation from a foreign language in speech synthesis networks,”
Y. Lee and T. Kim, · 2018
Cited alongside, same era.
“RWTH ASR systems for librispeech: Hybrid vs attention - w/o data augmentation,”
C. Lüscher, E. Beck, K. Irie, M. Kitza, W. Michel, A. Zeyer, R. Schlüter, and H. Ney, · 2019
Cited alongside, same era.
“The IWSLT 2019 evaluation campaign,” 2019
J. Niehues, R. Cattoni, S. Stüker, M. Negri, M. Turchi, E. Salesky, R. Sanabria, L. Barrault, L. Specia, and M. Federico, · 2019
“On layer normalization in the transformer architecture,” 2019
R. Xiong, Y. Yang, D. He, K. Zheng, S. Zheng, H. Zhang, Y. Lan, L. Wang, and T. Liu, · 2019
Later among the works it cites.
“MuST-C: a multilingual speech translation corpus,”
M. Gangi, R. Cattoni, L. Bentivogli, M. Negri, and M. Turchi, · 2019
Later among the works it cites.
“One-to-many multilingual end-to-end speech translation,”
Mattia Antonino Di Gangi, Matteo Negri, and Marco Turchi, · 2019
Later among the works it cites.
“Espnet-st: All-in-one speech translation toolkit,”
H. Inaguma, S. Kiyono, K. Duh, S. Karita, N. Soplin, T. Hayashi, and S. Watanabe, · 2020
Closest in time.
“fairseq s2t: Fast speech-to-text modeling with fairseq,”
C. Wang, Y. Tang, X. Ma, A. Wu, D. Okhonko, and J. Pino, · 2020
Closest in time.
“Phone features improve speech translation,”
Elizabeth Salesky and Alan W Black, · 2020
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“Specaugment: A simple data augmentation method for automatic speech recognition,”
D. Park, W. Chan, Y. Zhang, C. Chiu, B. Zoph, E. Cubuk, and Q. Le, · 2019
Cited alongside, same era.
“Speech recognition with augmented synthesized speech,”
A. Rosenberg, Y. Zhang, B. Ramabhadran, Y. Jia, P. Moreno, Y. Wu, and Z. Wu, · 2019
Cited alongside, same era.
“Learn spelling from teachers: Transferring knowledge from language models to sequence-to-sequence speech recognition,”
Y. Bai, J. Yi, J. Tao, Z. Tian, and Z. Wen, · 2019
Cited alongside, same era.
“A comparative study on end-to-end speech to text translation,”
P. Bahar, T. Bieschke, and H. Ney, · 2019
Cited alongside, same era.
“Leveraging unpaired text data for training end-to-end speech-to-intent systems,”
Y. Huang, H. Kuo, S. Thomas, Z. Kons, K. Audhkhasi, B. Kingsbury, R. Hoory, and M. Picheny, · 2020
Closest in time.
“Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension,”
M. Lewis, Y. Liu, N. Goyal, M. Ghazvininejad, A. Mohamed, O. Levy, V. Stoyanov, and L. Zettlemoyer, · 2020
Closest in time.
“End-to-end asr: from supervised to semi-supervised learning with modern architectures,”
G. Synnaeve, Q. Xu, J. Kahn, E. Grave, T. Likhomanenko, V. Pratap, A. Sriram, V. Liptchinsky, and R. Collobert, · 2020
Closest in time.
“Self-training for end-to-end speech translation,”
J. Pino, Q. Xu, X. Ma, M. Dousti, and Y. Tang, · 2020
Closest in time.