Fetching the paper…
Reading the bibliography…
This paper proposes a new end-to-end text-to-speech (E2E-TTS) model based on neural machine translation (NMT).
“Corpus of Spontaneous Japanese: Its design and evaluation”
Kikuo Maekawa · 2003
Earlier work this paper cites.
“Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks”
Alex Graves, Santiago Fernández, Faustino Gomez and Jürgen Schmidhuber · 2006
Earlier work this paper cites.
“Text-to-speech synthesis”, 2009
Paul Taylor · 2009
Earlier work this paper cites.
“Statistical parametric speech synthesis”
Heiga Zen, Keiichi Tokuda and Alan Black · 2009
Earlier work this paper cites.
“Recent development of open-source speech recognition engine julius”
Akinobu Lee and Tatsuya Kawahara · 2009
Earlier work this paper cites.
“Speech synthesis based on hidden Markov models”
Keiichi Tokuda et al · 2013
Earlier work this paper cites.
“WaveNet: A Generative Model for Raw Audio”
Aaron Oord et al · 2016
Earlier work this paper cites.
“Investigating gated recurrent networks for speech synthesis”
Zhizheng Wu and Simon King · 2016
Earlier work this paper cites.
“Tacotron: Towards End-to-End Speech Synthesis”
Yuxuan Wang et al · 2017
Earlier work this paper cites.
“Speaker-dependent WaveNet vocoder.”
Akira Tamamori, Tomoki Hayashi, Kazuhiro Kobayashi, Kazuya Takeda and Tomoki Toda · 2017
Earlier work this paper cites.
“Neural discrete representation learning”
Aaron van Oord and Oriol Vinyals · 2017
Earlier work this paper cites.
“The zero resource speech challenge 2017”
Ewan Dunbar et al · 2017
Earlier work this paper cites.
“Attention is all you need”
Ashish Vaswani et al · 2017
Earlier work this paper cites.
“Hybrid CTC/attention architecture for end-to-end speech recognition”
Shinji Watanabe, Takaaki Hori, Suyoun Kim, John Hershey and Tomoki Hayashi · 2017
Cited alongside, same era.
“JSUT corpus: free large-scale Japanese speech corpus for end-to-end speech synthesis”
Ryosuke Sonobe, Shinnosuke Takamichi and Hiroshi Saruwatari · 2017
Cited alongside, same era.
“Natural TTS Synthesis by Conditioning WaveNet on Mel Spectrogram Predictions”
Jonathan Shen et al · 2018
Cited alongside, same era.
“Deep Voice 3: Scaling Text-to-Speech with Convolutional Sequence Learning”
Wei Ping et al · 2018
Cited alongside, same era.
“Close to Human Quality TTS with Transformer”
Naihan Li et al · 2018
Cited alongside, same era.
“WaveGlow: A flow-based generative network for speech synthesis”
Ryan Prenger, Rafael Valle and Bryan Catanzaro · 2019
Later among the works it cites.
“Melgan: Generative adversarial networks for conditional waveform synthesis”
Kundan Kumar et al · 2019
Later among the works it cites.
“Group Latent Embedding for Vector Quantized Variational Autoencoder in Non-Parallel Voice Conversion”
Shaojin Ding and Ricardo Gutierrez-Osuna · 2019
Later among the works it cites.
“Speech-to-speech Translation between Untranscribed Unknown Languages”
Andros Tjandra, Sakriani Sakti and Satoshi Nakamura · 2019
Later among the works it cites.
“The zero resource speech challenge 2019: TTS without T”
Ewan Dunbar et al · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Gustav Henter, Jaime Lorenzo-Trueba, Xin Wang and Junichi Yamagishi · 2018
Cited alongside, same era.
Taku Kudo and John Richardson · 2018
Cited alongside, same era.
“Fast spectrogram inversion using multi-head convolutional neural networks”
SercanÖ Arık, Heewoo Jun and Gregory Diamos · 2018
Cited alongside, same era.
“Open JTalk (version 1.11)”,
2018
Cited alongside, same era.
“ESPnet: End-to-End Speech Processing Toolkit”
Shinji Watanabe et al · 2018
Cited alongside, same era.
“webMUSHRA-A comprehensive framework for web-based listening tests”
Michael Schoeffler et al · 2018
Cited alongside, same era.
“ClariNet: Parallel Wave Generation in End-to-End Text-to-Speech”
Wei Ping, Kainan Peng and Jitong Chen · 2019
Cited alongside, same era.
Alexei Baevski, Steffen Schneider and Michael Auli · 2019
Later among the works it cites.
“When does label smoothing help?”
Rafael Müller, Simon Kornblith and Geoffrey Hinton · 2019
Later among the works it cites.
“On the variance of the adaptive learning rate and beyond”
Liyuan Liu et al · 2019
Later among the works it cites.
“Parallel WaveGAN: A fast waveform generation model based on generative adversarial networks with multi-resolution spectrogram”
Ryuichi Yamamoto, Eunwoo Song and Jae-Min Kim · 2020
Closest in time.
“ESPnet-TTS: Unified, Reproducible, and Integratable Open Source End-to-End Text-to-Speech Toolkit”
Tomoki Hayashi et al · 2020
Closest in time.
“DiscreTalk audio sample”,
2020
Closest in time.
“Unsupervised speech representation learning using wavenet autoencoders”
Jan Chorowski, Ron Weiss, Samy Bengio and Aäron van Oord · 2053
Closest in time.