Fetching the paper…
Reading the bibliography…
This paper introduces a new end-to-end text-to-speech (E2E-TTS) toolkit named ESPnet-TTS, which is an extension of the open-source speech processing toolkit ESPnet.
“JNAS: Japanese speech corpus for large vocabulary continuous speech recognition research”
Katunobu Itou et al · 1999
Earlier work this paper cites.
“The CMU Arctic speech databases”
John Kominek and Alan Black · 2004
Earlier work this paper cites.
“The Kaldi Speech Recognition Toolkit”
Daniel Povey et al · 2011
Earlier work this paper cites.
“Speech synthesis based on hidden Markov models”
Keiichi Tokuda et al · 2013
Earlier work this paper cites.
“Statistical Parametric Speech Synthesis Using Deep Neural Networks”
Heiga Zen, Andrew Senior and Mike Schuster · 2013
Earlier work this paper cites.
“A fast Griffin-Lim algorithm”
Nathana“”el Perraudin, Peter Balazs and Peter Sndergaard · 2013
Earlier work this paper cites.
“Deep mixture density networks for acoustic modeling in statistical parametric speech synthesis”
Heiga Zen and Andrew Senior · 2014
Earlier work this paper cites.
“Empirical evaluation of gated recurrent neural networks on sequence modeling”
Junyoung Chung et al · 2014
Earlier work this paper cites.
“Attention-based models for speech recognition”
Jan Chorowski et al · 2015
Earlier work this paper cites.
“LibriSpeech: an ASR corpus based on public domain audio books”
Vassil Panayotov et al · 2015
Earlier work this paper cites.
“HMM/DNN-based Speech Synthesis System (HTS)”, http://hts.sp.nitech.ac.jp/ , 2016
HTS working group · 2016
Earlier work this paper cites.
“Merlin: An Open Source Neural Network Speech Synthesis System”
Zhizheng Wu, Oliver Watts and Simon King · 2016
Earlier work this paper cites.
“WaveNet: A Generative Model for Raw Audio”
Aaron van Oord et al · 2016
Earlier work this paper cites.
“Investigating gated recurrent networks for speech synthesis”
Zhizheng Wu and Simon King · 2016
Earlier work this paper cites.
“WORLD: A vocoder-based high-quality speech synthesis system for real-time applications”
Masanori Morise, Fumiya Yokomori and Kenji Ozawa · 2016
Earlier work this paper cites.
“Tacotron: Towards End-to-End Speech Synthesis”
Yuxuan Wang et al · 2017
Earlier work this paper cites.
“Speaker-Dependent WaveNet Vocoder”
Akira Tamamori et al · 2017
Earlier work this paper cites.
“Automatic Differentiation in PyTorch”
Adam Paszke et al · 2017
Earlier work this paper cites.
“The Blizzard Challenge 2017”
Simon King, Lovisa Wihlborg and Wei Guo · 2017
Earlier work this paper cites.
“Chinese Standard Mandarin Speech Copus”, https://www.data-baker.com/open_source.html , 2017
Data Baker · 2017
Cited alongside, same era.
“JSUT corpus: free large-scale Japanese speech corpus for end-to-end speech synthesis”
Ryosuke Sonobe, Shinnosuke Takamichi and Hiroshi Saruwatari · 2017
Cited alongside, same era.
“The LJ Speech Dataset”, https://keithito.com/LJ-Speech-Dataset/ , 2017
Keith Ito · 2017
Cited alongside, same era.
“The World English Bible: A large, single-speaker speech datasaet in English”, https://www.kaggle.com/bryanpark/the-world-english-bible-speech-dataset , 2017
Kyubyong Park · 2017
Cited alongside, same era.
“VAIS-1000: A Vietnamese Speech Synthesis Corpus”, http://dx.doi.org/10.21227/H2B887 , 2017
Quoc Do and Chi Luong · 2017
Cited alongside, same era.
“A Comparative Study on Transformer vs RNN in Speech Applications”
Shigeki Karita et al · 2019
Closest in time.
“r9y9/deepvoice3_pytorch”
Ryuichi Yamamoto · 2019
Closest in time.
“Kyubuyong/tacotron”
Kyubyong Park · 2019
Closest in time.
“Rayhane-mamah/Tacotron-2”
Rayhane Mama · 2019
Closest in time.
“NVIDIA/tacotron2”
Rafael Valle · 2019
Closest in time.
“mozilla/TTS”
Eren G“”olge · 2019
Closest in time.
“Semi-supervised End-to-end Speech Recognition Using Text-to-speech and Autoencoders”
Shigeki Karita et al · 2019
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Natural TTS Synthesis by Conditioning WaveNet on Mel Spectrogram Predictions”
Jonathan Shen et al · 2018
Cited alongside, same era.
“Deep Voice 3: Scaling Text-to-Speech with Convolutional Sequence Learning”
Wei Ping et al · 2018
Cited alongside, same era.
“Close to Human Quality TTS with Transformer”
Naihan Li et al · 2018
Cited alongside, same era.
“Efficient Neural Audio Synthesis”
Nal Kalchbrenner et al · 2018
Cited alongside, same era.
“Style tokens: Unsupervised style modeling, control and transfer in end-to-end speech synthesis”
Yuxuan Wang et al · 2018
Cited alongside, same era.
“ESPnet: End-to-End Speech Processing Toolkit”
Shinji Watanabe et al · 2018
Cited alongside, same era.
“Mixed-precision training for NLP and speech recognition with OpenSeq2Seq”
Oleksii Kuchaiev et al · 2018
Cited alongside, same era.
Murali Baskar et al · 2019
Closest in time.
“Almost Unsupervised Text to Speech and Automatic Speech Recognition”
Yi Ren et al · 2019
Closest in time.
“r9y9/wavenet_vocoder”
Ryuichi Yamamoto · 2019
Closest in time.
“kan-bayashi/PytorchWaveNetVocoder”
Tomoki Hayashi · 2019
Closest in time.
“kan-bayashi/ParallelWaveGAN”
Tomoki Hayashi · 2019
Closest in time.
“JVS corpus: free Japanese multi-speaker voice corpus”
Shinnosuke Takamichi et al · 2019
Closest in time.
“LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech”
Heiga Zen et al · 2019
Closest in time.
“The M-AILABS speech dataset”, https://www.caito.de/2019/01/the-m-ailabs-speech-dataset/ , 2019
Imdat Solak · 2019
Closest in time.
“WaveGlow: A flow-based generative network for speech synthesis”
Ryan Prenger, Rafael Valle and Bryan Catanzaro · 2019
Closest in time.
“r9y9/nnmnkwii”
Ryuichi Yamamoto · 2019
Closest in time.
“Parallel WaveGAN: A fast waveform generation model based on generative adversarial networks with multi-resolution spectrogram”
Ryuichi Yamamoto, Eunwoo Song and Jae-Min Kim · 2020
Closest in time.