Fetching the paper…
Reading the bibliography…
Sequence-to-sequence models have been widely used in end-to-end speech processing, for example, automatic speech recognition (ASR), speech translation (ST), and text-to-speech (TTS).
“FastSpeech: Fast, Robust and Controllable Text to Speech”
Yi Ren et al · 1905
Earlier work this paper cites.
“Improving Transformer-Based End-to-End Speech Recognition with Connectionist Temporal Classification and Language Model Integration”
Shigeki Karita et al · 1938
Earlier work this paper cites.
“SWITCHBOARD: telephone speech corpus for research and development”
J. Godfrey, E. Holliman and J. McDaniel · 1992
Earlier work this paper cites.
“The Design for the Wall Street Journal-based CSR Corpus”
Douglas. Paul and Janet. Baker · 1992
Earlier work this paper cites.
“Spontaneous Speech Corpus of Japanese”
Kikuo Maekawa, Hanae Koiso, Sadaoki Furui and Hitoshi Isahara · 2000
Earlier work this paper cites.
“Aurora working group: DSR front end LVCSR evaluation AU/384/02”
David Pearce and J Picone · 2002
Earlier work this paper cites.
“Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks”
Alex Graves, Santiago Fernández, Faustino. Gomez and Jürgen Schmidhuber · 2006
Earlier work this paper cites.
“HKUST/MTS: A Very Large Scale Mandarin Telephone Speech Corpus”
Yi Liu et al · 2006
Earlier work this paper cites.
“Recurrent Neural Network based Language Model”
T Mikolov et al · 2010
Earlier work this paper cites.
“The Kaldi Speech Recognition Toolkit”
Daniel Povey et al · 2011
Earlier work this paper cites.
“SentencePiece: A simple and language independent subword tokenizer and detokenizer for Neural Text Processing”
Taku Kudo and John Richardson · 2012
Earlier work this paper cites.
“TED-LIUM: an automatic speech recognition dedicated corpus”
Anthony Rousseau, Paul Deleglise and Yannick Esteve · 2012
Earlier work this paper cites.
“ADADELTA: An Adaptive Learning Rate Method”
Matthew. Zeiler · 2012
Earlier work this paper cites.
“Improved Speech-to-Text Translation with the Fisher and Callhome Spanish–English Speech Translation Corpus”
Matt Post et al · 2013
Earlier work this paper cites.
“Sequence to Sequence Learning with Neural Networks”
Ilya Sutskever, Oriol Vinyals and Quoc Le · 2014
Earlier work this paper cites.
“A pitch extraction algorithm tuned for automatic speech recognition”
P. Ghahremani et al · 2014
Earlier work this paper cites.
“Neural Machine Translation by Jointly Learning to Align and Translate”
Dzmitry Bahdanau, Kyunghyun Cho and Yoshua Bengio · 2015
Earlier work this paper cites.
“Effective Approaches to Attention-based Neural Machine Translation”
Thang Luong, Hieu Pham and Christopher. Manning · 2015
Cited alongside, same era.
“Audio augmentation for speech recognition”
Tom Ko, Vijayaditya Peddinti, Daniel Povey and Sanjeev Khudanpur · 2015
Cited alongside, same era.
“LibriSpeech: An ASR corpus based on public domain audio books”
Vassil Panayotov, Guoguo Chen, Daniel Povey and Sanjeev Khudanpur · 2015
Cited alongside, same era.
“Listen, attend and spell: A neural network for large vocabulary conversational speech recognition”
William Chan, Navdeep Jaitly, Quoc Le and Oriol Vinyals · 2016
Cited alongside, same era.
“Attention is All you Need”
Ashish Vaswani et al · 2017
Cited alongside, same era.
“Advances in Joint CTC-Attention Based End-to-End Speech Recognition with a Deep CNN Encoder and RNN-LM”
Takaaki Hori, Shinji Watanabe, Yu Zhang and William Chan · 2017
“Tensor2Tensor for Neural Machine Translation”
Ashish Vaswani et al · 2018
Later among the works it cites.
“ESPnet: End-to-End Speech Processing Toolkit”
Shinji Watanabe et al · 2018
Later among the works it cites.
“Training Tips for the Transformer Model”
Martin Popel and Ondrej Bojar · 2018
Later among the works it cites.
“A Comparison of Transformer and Recurrent Neural Networks on Multilingual Neural Machine Translation”
Surafel Lakew, Mauro Cettolo and Marcello Federico · 2018
Later among the works it cites.
“Syllable-Based Sequence-to-Sequence Speech Recognition with the Transformer in Mandarin Chinese”
Shiyu Zhou, Linhao Dong, Shuang Xu and Bo Xu · 2018
Later among the works it cites.
“Natural TTS Synthesis by Conditioning Wavenet on MEL Spectrogram Predictions”
Jonathan Shen et al · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“Sequence-to-Sequence Models Can Directly Translate Foreign Speech”
Ron. Weiss et al · 2017
Cited alongside, same era.
“JSUT corpus: free large-scale Japanese speech corpus for end-to-end speech synthesis”
Ryosuke Sonobe, Shinnosuke Takamichi and Hiroshi Saruwatari · 2017
Cited alongside, same era.
“AISHELL-1: An open-source Mandarin speech corpus and a speech recognition baseline”
H. Bu et al · 2017
Cited alongside, same era.
“The third CHiME speech separation and recognition challenge: Analysis and outcomes”
Jon Barker, Ricard Marxer, Emmanuel Vincent and Shinji Watanabe · 2017
Cited alongside, same era.
“The REVERB Challenge: A Benchmark Task for Reverberation-Robust ASR Techniques”
Keisuke Kinoshita et al · 2017
Cited alongside, same era.
“Language independent end-to-end architecture for joint language identification and speech recognition”
Shinji Watanabe, Takaaki Hori and John Hershey · 2017
Cited alongside, same era.
Later among the works it cites.
“End-to-end Speech Recognition With Word-Based Rnn Language Models”
Takaaki Hori, Jaejin Cho and Shinji Watanabe · 2018
Later among the works it cites.
“Efficiently trainable text-to-speech system based on deep convolutional networks with guided attention”
Hideyuki Tachibana, Katsuya Uenoyama and Shunsuke Aihara · 2018
Later among the works it cites.
“The Fifth ’CHiME’ Speech Separation and Recognition Challenge: Dataset, Task and Baselines”
Jon Barker, Shinji Watanabe, Emmanuel Vincent and Jan Trmal · 2018
Later among the works it cites.
“TED-LIUM 3: Twice as Much Data and Corpus Repartition for Experiments on Speaker Adaptation”
Francois Hernandez et al · 2018
Later among the works it cites.
“Improved Training of End-to-end Attention Models for Speech Recognition”
Albert Zeyer, Kazuki Irie, Ralf Schluter and Hermann Ney · 2018
Later among the works it cites.
“Neural Speech Synthesis with Transformer Network”
Naihan Li et al · 2019
Closest in time.
“SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition”
D.. Park et al · 2019
Closest in time.
“Language Modeling with Deep Transformers”
Kazuki Irie, Albert Zeyer, Ralf Schlüter and Hermann Ney · 2019
Closest in time.
“RWTH ASR Systems for LibriSpeech: Hybrid vs Attention-w/o Data Augmentation”
Christoph Lüscher et al · 2019
Closest in time.
“The M-AILABS Speech Dataset”, https://www.caito.de/2019/01/the-m-ailabs-speech-dataset/ , 2019
Imdat Solak · 2019
Closest in time.