Fetching the paper…
Reading the bibliography…
This paper describes ESPnet2-TTS, an end-to-end text-to-speech (E2E-TTS) toolkit.
“Corpus of spontaneous Japanese: Its design and evaluation”
Kikuo Maekawa · 2003
Earlier work this paper cites.
“Toward accurate dynamic time warping in linear time and space”
Stan Salvador and Philip Chan · 2007
Earlier work this paper cites.
“The Kaldi Speech Recognition Toolkit”
Daniel Povey et al · 2011
Earlier work this paper cites.
“Chainer: A Deep Learning Framework for Accelerating the Research Cycle”
Seiya Tokui et al · 2011
Earlier work this paper cites.
“Speech synthesis based on hidden Markov models”
Keiichi Tokuda et al · 2013
Earlier work this paper cites.
“Deep mixture density networks for acoustic modeling in statistical parametric speech synthesis”
Heiga Zen and Andrew Senior · 2014
Earlier work this paper cites.
“LibriSpeech: an ASR corpus based on public domain audio books”
Vassil Panayotov et al · 2015
Earlier work this paper cites.
“Tacotron: Towards End-to-End Speech Synthesis”
Yuxuan Wang et al · 2017
Earlier work this paper cites.
“The LJ Speech Dataset”, https://keithito.com/LJ-Speech-Dataset/ , 2017
Keith Ito · 2017
Earlier work this paper cites.
“JSUT corpus: free large-scale Japanese speech corpus for end-to-end speech synthesis”
Ryosuke Sonobe, Shinnosuke Takamichi and Hiroshi Saruwatari · 2017
Earlier work this paper cites.
“Natural TTS Synthesis by Conditioning WaveNet on Mel Spectrogram Predictions”
Jonathan Shen et al · 2018
Earlier work this paper cites.
“Close to Human Quality TTS with Transformer”
Naihan Li et al · 2018
Earlier work this paper cites.
“Transfer learning from speaker verification to multispeaker text-to-speech synthesis”
Ye Jia et al · 2018
Earlier work this paper cites.
“Style tokens: Unsupervised style modeling, control and transfer in end-to-end speech synthesis”
Yuxuan Wang et al · 2018
Earlier work this paper cites.
“ESPnet: End-to-end speech processing toolkit”
Shinji Watanabe et al · 2018
Earlier work this paper cites.
“OpenSeq2Seq: extensible toolkit for distributed and mixed precision training of sequence-to-sequence models”
Oleksii Kuchaiev et al · 2018
Earlier work this paper cites.
“Efficiently trainable text-to-speech system based on deep convolutional networks with guided attention”
Hideyuki Tachibana, Katsuya Uenoyama and Shunsuke Aihara · 2018
Cited alongside, same era.
“X-vectors: Robust DNN embeddings for speaker recognition”
David Snyder et al · 2018
Cited alongside, same era.
“webMUSHRA – A comprehensive framework for web-based listening tests”
Michael Schoeffler et al · 2018
Cited alongside, same era.
“An Investigation of Noise Shaping with Perceptual Weighting for WaveNet-Based Speech Generation”
Kentaro Tachibana et al · 2018
Cited alongside, same era.
“FastSpeech: Fast, Robust and Controllable Text to Speech”
Yi Ren et al · 2019
Cited alongside, same era.
“PyTorch: An imperative style, high-performance deep learning library”
Adam Paszke et al · 2019
“ESPnet-TTS: Unified, reproducible, and integratable open source end-to-end text-to-speech toolkit”
Tomoki Hayashi et al · 2020
Later among the works it cites.
“Voice Conversion Challenge 2020: Intra-lingual semi-parallel and cross-lingual voice conversion”
Yi Zhao et al · 2020
Later among the works it cites.
“Conformer: Convolution-augmented Transformer for speech recognition”
Anmol Gulati et al · 2020
Later among the works it cites.
“Asteroid: the PyTorch-based audio source separation toolkit for researchers”
Manuel Pariente et al · 2020
Later among the works it cites.
“Transformers: State-of-the-Art Natural Language Processing”
Thomas Wolf et al · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Wen-Chin Huang et al · 2019
Cited alongside, same era.
“NeMo: a toolkit for building AI applications using neural modules”
Oleksii Kuchaiev et al · 2019
Cited alongside, same era.
“fairseq: A Fast, Extensible Toolkit for Sequence Modeling”
Myle Ott et al · 2019
Cited alongside, same era.
“MelGAN: Generative Adversarial Networks for Conditional Waveform Synthesis”
Kundan Kumar et al · 2019
Cited alongside, same era.
“High fidelity speech synthesis with adversarial networks”
Mikołaj Bińkowski et al · 2019
Cited alongside, same era.
“CSTR VCTK corpus: English multi-speaker corpus for CSTR voice cloning toolkit (version 0.92)”, 2019
Junichi Yamagishi, Christophe Veaux and Kirsten MacDonald · 2019
Cited alongside, same era.
“Parallel WaveGAN: A fast waveform generation model based on generative adversarial networks with multi-resolution spectrogram”
Ryuichi Yamamoto, Eunwoo Song and Jae-Min Kim · 2020
Later among the works it cites.
“HiFi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis”
Jungil Kong, Jaehyeon Kim and Jaekyoung Bae · 2020
Later among the works it cites.
“FastSpeech 2: Fast and High-Quality End-to-End Text to Speech”
Yi Ren et al · 2021
Closest in time.
“fairseq S^2: A Scalable and Integrable Speech Synthesis Toolkit”
Changhan Wang et al · 2021
Closest in time.
“Recent developments on ESPnet toolkit boosted by Conformer”
Pengcheng Guo et al · 2021
Closest in time.
“FastPitch: Parallel text-to-speech with pitch prediction”
Adrian Łańcucki · 2021
Closest in time.
“StyleMelGAN: An efficient high-fidelity adversarial vocoder with temporal adaptive normalization”
Ahmed Mustafa, Nicola Pia and Guillaume Fuchs · 2021
Closest in time.
“Multi-band MelGAN: Faster waveform generation for high-quality text-to-speech”
Geng Yang et al · 2021
Closest in time.
“WaveGrad 2: Iterative Refinement for Text-to-Speech Synthesis”
Nanxin Chen et al · 2021
Closest in time.
“Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-Speech”
Jaehyeon Kim, Jungil Kong and Juhee Son · 2021
Closest in time.
“Prosodic Features Control by Symbols as Input of Sequence-to-Sequence Acoustic Modeling for Neural TTS”
Kiyoshi Kurihara, Nobumasa Seiyama and Tadashi Kumano · 2021
Closest in time.