Fetching the paper…
Reading the bibliography…
We present ESPnet-ST, which is designed for the quick development of speech-to-speech translation systems in a single framework.
Lingvo: a modular and scalable framework for sequence-to-sequence modeling
Jonathan Shen, Patrick Nguyen, Yonghui Wu, Zhifeng Chen, Mia X Chen, Ye Jia, Anjuli Kannan, Tara Sainath, Yuan Cao, Chung-Cheng Chiu, et al. 2019 · 1902
Earlier work this paper cites.
NeMo: a toolkit for building AI applications using Neural Modules
Oleksii Kuchaiev, Jason Li, Huyen Nguyen, Oleksii Hrinchuk, Ryan Leary, Boris Ginsburg, Samuel Kriman, Stanislav Beliaev, Vitaly Lavrukhin, Jack Cook, et al. 2019 · 1909
Earlier work this paper cites.
Speech translation: Coupling of recognition and translation
Hermann Ney. 1999 · 1999
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Moses: Open source toolkit for statistical machine translation
Philipp Koehn, Hieu Hoang, Alexandra Birch, Chris Callison-Burch, Marcello Federico, Nicola Bertoldi, Brooke Cowan, Wade Shen, Christine Moran, Richard Zens, Chris Dyer, Ondřej Bojar, Alexandra Constantin, and Evan Herbst. 2007 · 2007
Earlier work this paper cites.
Recent efforts in spoken language translation
F. Casacuberta, M. Federico, H. Ney, and E. Vidal. 2008 · 2008
Earlier work this paper cites.
Modeling punctuation prediction as machine translation
Stephan Peitz, Markus Freitag, Arne Mauser, and Hermann Ney. 2011 · 2011
Earlier work this paper cites.
The kaldi speech recognition toolkit
Daniel Povey, Arnab Ghoshal, Gilles Boulianne, Lukas Burget, Ondrej Glembek, Nagendra Goel, Mirko Hannemann, Petr Motlicek, Yanmin Qian, Petr Schwarz, et al. 2011 · 2011
Earlier work this paper cites.
Chainer: A deep learning framework for accelerating the research cycle
Seiya Tokui, Ryosuke Okuta, Takuya Akiba, Yusuke Niitani, Toru Ogawa, Shunta Saito, Shuji Suzuki, Kota Uenishi, Brian Vogel, and Hiroyuki Yamazaki Vincent. 2019 · 2011
Earlier work this paper cites.
Wit3: Web inventory of transcribed and translated talks
Mauro Cettolo, Christian Girardi, and Marcello Federico. 2012 · 2012
Earlier work this paper cites.
Improved speech-to-text translation with the Fisher and Callhome Spanish–English speech translation corpus
Matt Post, Gaurav Kumar, Adam Lopez, Damianos Karakos, Chris Callison-Burch, and Sanjeev Khudanpur. 2013 · 2013
Earlier work this paper cites.
Collecting bilingual audio in remote indigenous communities
Steven Bird, Lauren Gawne, Katie Gelbart, and Isaac McAlister. 2014 · 2014
Earlier work this paper cites.
Some insights from translating conversational telephone speech
Gaurav Kumar, Matt Post, Daniel Povey, and Sanjeev Khudanpur. 2014 · 2014
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2015 · 2015
Earlier work this paper cites.
Listen and translate: A proof of concept for end-to-end speech-to-text translation
Alexandre Bérard, Olivier Pietquin, Christophe Servan, and Laurent Besacier. 2016 · 2016
Earlier work this paper cites.
Listen, attend and spell: A neural network for large vocabulary conversational speech recognition
William Chan, Navdeep Jaitly, Quoc Le, and Oriol Vinyals. 2016 · 2016
Earlier work this paper cites.
Wavenet: A generative model for raw audio
Aaron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew W. Senior, and Koray Kavukcuoglu. 2016 · 2016
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Earlier work this paper cites.
MuST-C: a Multilingual Speech Translation Corpus
Mattia A. Di Gangi, Roldano Cattoni, Luisa Bentivogli, Matteo Negri, and Marco Turchi. 2019a · 2017
Earlier work this paper cites.
An analysis of incorporating an external language model into a sequence-to-sequence model
Anjuli Kannan, Yonghui Wu, Patrick Nguyen, Tara N Sainath, Zhifeng Chen, and Rohit Prabhavalkar. 2017 · 2017
Cited alongside, same era.
OpenNMT: Open-source toolkit for neural machine translation
Guillaume Klein, Yoon Kim, Yuntian Deng, Jean Senellart, and Alexander Rush. 2017 · 2017
Cited alongside, same era.
Natural TTS synthesis by conditioning WaveNet on Mel spectrogram predictions
Jonathan Shen, Ruoming Pang, Ron J. Weiss, Mike Schuster, Navdeep Jaitly, Zongheng Yang, Zhifeng Chen, Yu Zhang, Yuxuan Wang, R. J. Skerry-Ryan, Rif A. Saurous, Yannis Agiomyrgiannakis, and Yonghui Wu. 2018 · 2017
Cited alongside, same era.
Hybrid CTC/attention architecture for end-to-end speech recognition
Shinji Watanabe, Takaaki Hori, Suyoun Kim, John R Hershey, and Tomoki Hayashi. 2017 · 2017
Cited alongside, same era.
Sequence-to-sequence models can directly translate foreign speech
Ron J Weiss, Jan Chorowski, Navdeep Jaitly, Yonghui Wu, and Zhifeng Chen. 2017 · 2017
Cited alongside, same era.
A comparative study on end-to-end speech to text translation
Parnia Bahar, Tobias Bieschke, and Hermann Ney. 2019a · 2019
Later among the works it cites.
On using SpecAugment for end-to-end speech translation
Parnia Bahar, Albert Zeyer, Ralf Schlüter, and Hermann Ney. 2019b · 2019
Later among the works it cites.
Pre-training on high-resource speech recognition improves low-resource speech-to-text translation
Sameer Bansal, Herman Kamper, Karen Livescu, Adam Lopez, and Sharon Goldwater. 2019 · 2019
Later among the works it cites.
Adapting transformer to end-to-end spoken language translation
Mattia A Di Gangi, Matteo Negri, and Marco Turchi. 2019b · 2019
Later among the works it cites.
One-to-many multilingual end-to-end speech translation
Mattia Antonino Di Gangi, Matteo Negri, and Marco Turchi. 2019c · 2019
Later among the works it cites.
Multilingual end-to-end speech translation
Hirofumi Inaguma, Kevin Duh, Tatsuya Kawahara, and Shinji Watanabe. 2019 · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Tied multitask learning for neural speech translation
Antonios Anastasopoulos and David Chiang. 2018 · 2018
Cited alongside, same era.
End-to-end automatic speech translation of audiobooks
Alexandre Bérard, Laurent Besacier, Ali Can Kocabiyikoglu, and Olivier Pietquin. 2018 · 2018
Cited alongside, same era.
A very low resource language speech corpus for computational language documentation experiments
Pierre Godard, Gilles Adda, Martine Adda-Decker, Juan Benjumea, Laurent Besacier, Jamison Cooper-Leavitt, Guy-Noel Kouarata, Lori Lamel, Hélène Maynard, Markus Mueller, Annie Rialland, Sebastian Stueker, François Yvon, and Marcely Zanon-Boito. 2018 · 2018
Cited alongside, same era.
The IWSLT 2018 evaluation campaign
Niehues Jan, Roldano Cattoni, Stüker Sebastian, Mauro Cettolo, Marco Turchi, and Marcello Federico. 2018 · 2018
Cited alongside, same era.
Augmenting Librispeech with French translations: A multimodal corpus for direct speech translation evaluation
Ali Can Kocabiyikoglu, Laurent Besacier, and Olivier Kraif. 2018 · 2018
Cited alongside, same era.
OpenSeq2Seq: Extensible toolkit for distributed and mixed precision training of sequence-to-sequence models
Oleksii Kuchaiev, Boris Ginsburg, Igor Gitman, Vitaly Lavrukhin, Carl Case, and Paulius Micikevicius. 2018 · 2018
Cited alongside, same era.
Subword regularization: Improving neural network translation models with multiple subword candidates
Taku Kudo. 2018 · 2018
Cited alongside, same era.
Later among the works it cites.
A comparative study on Transformer vs RNN in speech applications
Shigeki Karita, Nanxin Chen, Tomoki Hayashi, Takaaki Hori, Hirofumi Inaguma, Ziyan Jiang, Masao Someki, Nelson Enrique Yalta Soplin, Ryuichi Yamamoto, Xiaofei Wang, et al. 2019 · 2019
Later among the works it cites.
Neural speech synthesis with transformer network
Naihan Li, Shujie Liu, Yanqing Liu, Sheng Zhao, and Ming Liu. 2019 · 2019
Later among the works it cites.
End-to-end speech translation with knowledge distillation
Yuchen Liu, Hao Xiong, Zhongjun He, Jiajun Zhang, Hua Wu, Haifeng Wang, and Chengqing Zong. 2019 · 2019
Later among the works it cites.
The IWSLT 2019 evaluation campaign
J. Niehues, R. Cattoni, S. Stüker, M. Negri, M. Turchi, E. Salesky, R. Sanabria, L. Barrault, L. Specia, and M Federico. 2019 · 2019
Later among the works it cites.
Fairseq: A fast, extensible toolkit for sequence modeling
Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli. 2019 · 2019
Later among the works it cites.
PyTorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al. 2019 · 2019
Later among the works it cites.
Wav2Letter++: A fast open-source speech recognition system
Vineel Pratap, Awni Hannun, Qiantong Xu, Jeff Cai, Jacob Kahn, Gabriel Synnaeve, Vitaliy Liptchinsky, and Ronan Collobert. 2019 · 2019
Later among the works it cites.
Fastspeech: Fast, robust and controllable text to speech
Yi Ren, Yangjun Ruan, Xu Tan, Tao Qin, Sheng Zhao, Zhou Zhao, and Tie-Yan Liu. 2019 · 2019
Later among the works it cites.
Vectorized Beam Search for CTC-Attention-Based Speech Recognition
Hiroshi Seki, Takaaki Hori, Shinji Watanabe, Niko Moritz, and Jonathan Le Roux. 2019 · 2019
Later among the works it cites.
Attention-passing models for robust and data-efficient end-to-end speech translation
Matthias Sperber, Graham Neubig, Jan Niehues, and Alex Waibel. 2019 · 2019
Later among the works it cites.
ESPnet-TTS: Unified, reproducible, and integratable open source end-to-end text-to-speech toolkit
Tomoki Hayashi, Ryuichi Yamamoto, Katsuki Inoue, Takenori Yoshimura, Shinji Watanabe, Tomoki Toda, Kazuya Takeda, Yu Zhang, and Xu Tan. 2020 · 2020
Closest in time.
Bridging the gap between pre-training and fine-tuning for end-to-end speech translation
Chengyi Wang, Yu Wu, Shujie Liu, Zhenglu Yang, and Ming Zhou. 2020 · 2020
Closest in time.
Parallel WaveGAN: A fast waveform generation model based on generative adversarial networks with multi-resolution spectrogram
Ryuichi Yamamoto, Eunwoo Song, and Jae-Min Kim. 2020 · 2020
Closest in time.