Fetching the paper…
Reading the bibliography…
Conventional spoken language translation (SLT) systems are pipeline based systems, where we have an Automatic Speech Recognition (ASR) system to convert the modality of source from speech to text and a Machine Translation (MT) systems to translate source text to text in target language.
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Speech-transformer: a no-recurrence sequence-to-sequence model for speech recognition
Linhao Dong, Shuang Xu, and Bo Xu · 2018
Earlier work this paper cites.
The jhu/kyotou speech translation system for iwslt 2018
Hirofumi Inaguma, Xuan Zhang, Zhiqi Wang, Adithya Renduchintala, Shinji Watanabe, and Kevin Duh · 2018
Earlier work this paper cites.
How2: a large-scale dataset for multimodal language understanding
Ramon Sanabria, Ozan Caglayan, Shruti Palaskar, Desmond Elliott, Loïc Barrault, Lucia Specia, and Florian Metze · 2018
Earlier work this paper cites.
Data augmentation for end-to-end speech translation: Fbk@ iwslt’19
M Di Gangi, Matteo Negri, Viet Nhat Nguyen, Amirhossein Tebbifakhr, and Marco Turchi · 2019
Earlier work this paper cites.
Adapting transformer to end-to-end spoken language translation
Mattia A Di Gangi, Matteo Negri, and Marco Turchi · 2019
Cited alongside, same era.
Adapting transformer to end-to-end spoken language translation
Mattia A Di Gangi, Matteo Negri, and Marco Turchi · 2019
Cited alongside, same era.
Self-training for end-to-end speech recognition
Jacob Kahn, Ann Lee, and Awni Hannun · 2019
Cited alongside, same era.
A comparative study on transformer vs rnn in speech applications
Shigeki Karita, Nanxin Chen, Tomoki Hayashi, Takaaki Hori, Hirofumi Inaguma, Ziyan Jiang, Masao Someki, Nelson Enrique Yalta Soplin, Ryuichi Yamamoto, Xiaofei Wang, et al · 2019
Cited alongside, same era.
Improving transformer-based end-to-end speech recognition with connectionist temporal classification and language model integration
Tomohiro Nakatani · 2019
Later among the works it cites.
Very deep self-attention networks for end-to-end speech recognition
Ngoc-Quan Pham, Thai-Son Nguyen, Jan Niehues, Markus Muller, and Alex Waibel · 2019
Later among the works it cites.
On leveraging the visual modality for neural machine translation
Vikas Raunak, Sang Keun Choe, Quanyang Lu, Yi Xu, and Florian Metze · 2019
Later among the works it cites.
Vectorized beam search for ctc-attention-based speech recognition
Hiroshi Seki, Takaaki Hori, Shinji Watanabe, Niko Moritz, and Jonathan Le Roux · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…