Fetching the paper…
Reading the bibliography…
End-to-end speech translation poses a heavy burden on the encoder, because it has to transcribe, understand, and learn cross-lingual semantics simultaneously.
End-to-end speech translation with knowledge distillation
Yuchen Liu, Hao Xiong, Zhongjun He, Jiajun Zhang, Hua Wu, Haifeng Wang, and Chengqing Zong. 2019 · 1904
Earlier work this paper cites.
Specaugment: A simple data augmentation method for automatic speech recognition
Daniel S. Park, William Chan, Yu Zhang, Chung-Cheng Chiu, Barret Zoph, Ekin D. Cubuk, and Quoc V. Le. 2019 · 1904
Earlier work this paper cites.
Bridging the gap between pre-training and fine-tuning for end-to-end speech translation
Chengyi Wang, Yu Wu, Shujie Liu, Zhenglu Yang, and Ming Zhou. 2019b · 1909
Earlier work this paper cites.
Improving transformer-based speech recognition using unsupervised pre-training
Dongwei Jiang, Xiaoning Lei, Wubo Li, Ne Luo, Yuxuan Hu, Wei Zou, and Xiangang Li. 2019 · 1910
Earlier work this paper cites.
On using specaugment for end-to-end speech translation
Parnia Bahar, Albert Zeyer, Ralf Schlüter, and Hermann Ney. 2019 · 1911
Earlier work this paper cites.
Semantic mask for transformer based end-to-end speech recognition
Chengyi Wang, Yu Wu, Yujiao Du, Jinyu Li, Shujie Liu, Liang Lu, Shuo Ren, Guoli Ye, Sheng Zhao, and Ming Zhou. 2019a · 1912
Earlier work this paper cites.
Speech translation: coupling of recognition and translation
Hermann Ney. 1999 · 1999
Earlier work this paper cites.
Alignment templates: the RWTH SMT system
Oliver Bender, Richard Zens, Evgeny Matusov, and Hermann Ney. 2004 · 2004
Earlier work this paper cites.
On the integration of speech recognition and statistical machine translation
Evgeny Matusov, Stephan Kanthak, and Hermann Ney. 2005 · 2005
Earlier work this paper cites.
Statistical phrase-based speech translation
Lambert Mathias and William Byrne. 2006 · 2006
Earlier work this paper cites.
Curriculum learning
Yoshua Bengio, Jérôme Louradour, Ronan Collobert, and Jason Weston. 2009 · 2009
Earlier work this paper cites.
Lium spkdiarization: an open source toolkit for diarization
Sylvain Meignier and Teva Merlin. 2010 · 2010
Earlier work this paper cites.
Enhancing the TED-LIUM corpus with selected data for language modeling and more TED talks
Anthony Rousseau, Paul Deléglise, and Yannick Estève. 2014 · 2014
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2015 · 2015
Cited alongside, same era.
Effective approaches to attention-based neural machine translation
Thang Luong, Hieu Pham, and Christopher D. Manning. 2015 · 2015
Cited alongside, same era.
Librispeech: An ASR corpus based on public domain audio books
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur. 2015 · 2015
Cited alongside, same era.
An unsupervised probability model for speech-to-translation alignment of low-resource languages
Antonios Anastasopoulos, David Chiang, and Long Duong. 2016 · 2016
Cited alongside, same era.
Listen and translate: A proof of concept for end-to-end speech-to-text translation
Alexandre Berard, Olivier Pietquin, Christophe Servan, and Laurent Besacier. 2016 · 2016
Cited alongside, same era.
Low-resource speech-to-text translation
Sameer Bansal, Herman Kamper, Karen Livescu, Adam Lopez, and Sharon Goldwater. 2018 · 2018
Later among the works it cites.
End-to-end automatic speech translation of audiobooks
Alexandre Berard, Laurent Besacier, Ali Can Kocabiyikoglu, and Olivier Pietquin. 2018 · 2018
Later among the works it cites.
State-of-the-art speech recognition with sequence-to-sequence models
Chung-Cheng Chiu, Tara N. Sainath, Yonghui Wu, Rohit Prabhavalkar, Patrick Nguyen, Zhifeng Chen, Anjuli Kannan, Ron J. Weiss, Kanishka Rao, Ekaterina Gonina, Navdeep Jaitly, Bo Li, Jan Chorowski, and Michiel Bacchiani. 2018 · 2018
Later among the works it cites.
Augmenting librispeech with french translations: A multimodal corpus for direct speech translation evaluation
Ali Can Kocabiyikoglu, Laurent Besacier, and Olivier Kraif. 2018 · 2018
Later among the works it cites.
Sentencepiece: A simple and language independent subword tokenizer and detokenizer for neural text processing
Taku Kudo and John Richardson. 2018 · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Listen, attend and spell: A neural network for large vocabulary conversational speech recognition
William Chan, Navdeep Jaitly, Quoc V. Le, and Oriol Vinyals. 2016 · 2016
Cited alongside, same era.
An attentional model for speech translation without transcription
Long Duong, Antonios Anastasopoulos, David Chiang, Steven Bird, and Trevor Cohn. 2016 · 2016
Cited alongside, same era.
Multi-modal curriculum learning for semi-supervised image classification
Chen Gong, Dacheng Tao, Stephen J. Maybank, Wei Liu, Guoliang Kang, and Jie Yang. 2016 · 2016
Cited alongside, same era.
Automated curriculum learning for neural networks
Alex Graves, Marc G. Bellemare, Jacob Menick, Rémi Munos, and Koray Kavukcuoglu. 2017 · 2017
Cited alongside, same era.
Structured-based curriculum learning for end-to-end english-japanese speech translation
Takatomo Kano, Sakriani Sakti, and Satoshi Nakamura. 2017 · 2017
Cited alongside, same era.
Montreal forced aligner: Trainable text-speech alignment using kaldi
Michael McAuliffe, Michaela Socolof, Sarah Mihuc, Michael Wagner, and Morgan Sonderegger. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
The iwslt 2018 evaluation campaign
Jan Niehues, Ronaldo Cattoni, Sebastian Stüker, Mauro Cettolo, Marco Turchi, and Marcello Federico. 2018 · 2018
Later among the works it cites.
Espnet: End-to-end speech processing toolkit
Shinji Watanabe, Takaaki Hori, Shigeki Karita, Tomoki Hayashi, Jiro Nishitoba, Yuya Unno, Nelson Enrique Yalta Soplin, Jahn Heymann, Matthew Wiesner, Nanxin Chen, Adithya Renduchintala, and Tsubasa Ochiai. 2018 · 2018
Later among the works it cites.
Pre-training on high-resource speech recognition improves low-resource speech-to-text translation
Sameer Bansal, Herman Kamper, Karen Livescu, Adam Lopez, and Sharon Goldwater. 2019 · 2019
Later among the works it cites.
BERT: pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Later among the works it cites.
Multilingual end-to-end speech translation
Hirofumi Inaguma, Kevin Duh, Tatsuya Kawahara, and Shinji Watanabe. 2019 · 2019
Later among the works it cites.
Leveraging weakly supervised data to improve end-to-end speech-to-text translation
Ye Jia, Melvin Johnson, Wolfgang Macherey, Ron J. Weiss, Yuan Cao, Chung-Cheng Chiu, Naveen Ari, Stella Laurenzo, and Yonghui Wu. 2019 · 2019
Later among the works it cites.
A comparative study on transformer vs RNN in speech applications
Shigeki Karita, Xiaofei Wang, Shinji Watanabe, Takenori Yoshimura, Wangyou Zhang, Nanxin Chen, Tomoki Hayashi, Takaaki Hori, Hirofumi Inaguma, Ziyan Jiang, Masao Someki, Nelson Enrique Yalta Soplin, and Ryuichi Yamamoto. 2019 · 2019
Later among the works it cites.
Attention-passing models for robust and data-efficient end-to-end speech translation
Matthias Sperber, Graham Neubig, Jan Niehues, and Alex Waibel. 2019 · 2019
Later among the works it cites.