Fetching the paper…
Reading the bibliography…
The recent progress in non-autoregressive text-to-speech (NAR-TTS) has made fast and high-quality speech synthesis possible.
Learning prosodic features using a tree representation
Julia Hirschberg and Owen Rambow · 2001
Earlier work this paper cites.
Learning phrase representations using rnn encoder–decoder for statistical machine translation
Kyunghyun Cho, Bart van Merrienboer, Çaglar Gülçehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio · 2014
Earlier work this paper cites.
Intonational phrase break prediction for text-to-speech synthesis using dependency relations
Taniya Mishra, Yeon-jun Kim, and Srinivas Bangalore · 2015
Earlier work this paper cites.
Gated graph sequence neural networks
Yujia Li, Richard Zemel, Marc Brockschmidt, and Daniel Tarlow · 2016
Earlier work this paper cites.
WaveNet: A Generative Model for Raw Audio
Aäron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior, and Koray Kavukcuoglu · 2016
Earlier work this paper cites.
The lj speech dataset
Keith Ito and Linda Johnson · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Tacotron: Towards end-to-end speech synthesis
Yuxuan Wang, R. J. Skerry-Ryan, Daisy Stanton, Yonghui Wu, Ron J. Weiss, Navdeep Jaitly, Zongheng Yang, Ying Xiao, Z. Chen, Samy Bengio, Quoc V. Le, Yannis Agiomyrgiannakis, Robert A. J. Clark, and Rif A. Saurous · 2017
Earlier work this paper cites.
Deep voice 3: Scaling text-to-speech with convolutional sequence learning
Wei Ping, Kainan Peng, Andrew Gibiansky, Sercan O Arik, Ajay Kannan, Sharan Narang, Jonathan Raiman, and John Miller · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Cited alongside, same era.
Fastspeech: fast, robust and controllable text to speech
Yi Ren, Yangjun Ruan, Xu Tan, Tao Qin, Sheng Zhao, Zhou Zhao, and Tie-Yan Liu · 2019
Cited alongside, same era.
Libritts: A corpus derived from librispeech for text-to-speech
Heiga Zen, Viet Dang, Rob Clark, Yu Zhang, Ron J Weiss, Ye Jia, Zhifeng Chen, and Yonghui Wu · 2019
Cited alongside, same era.
High fidelity speech synthesis with adversarial networks
Mikołaj Bińkowski, Jeff Donahue, Sander Dieleman, Aidan Clark, Erich Elsen, Norman Casagrande, Luis C. Cobo, and Karen Simonyan · 2020
Cited alongside, same era.
Hifisinger: Towards high-fidelity neural singing voice synthesis
Jiawei Chen, Xu Tan, Jian Luan, Tao Qin, and Tie-Yan Liu · 2020
Cited alongside, same era.
Graphtts: graph-to-sequence modelling in neural text-to-speech
Aolan Sun, Jianzong Wang, Ning Cheng, Huayi Peng, Zhen Zeng, and Jing Xiao · 2020
Later among the works it cites.
Parallel wavegan: A fast waveform generation model based on generative adversarial networks with multi-resolution spectrogram
Ryuichi Yamamoto, Eunwoo Song, and Jae-Min Kim · 2020
Later among the works it cites.
End-to-end adversarial text-to-speech
Jeff Donahue, Sander Dieleman, Mikolaj Binkowski, Erich Elsen, and Karen Simonyan · 2021
Later among the works it cites.
Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech
Jaehyeon Kim, Jungil Kong, and Juhee Son · 2021
Later among the works it cites.
Graphspeech: Syntax-aware graph attention network for neural speech synthesis
Rui Liu, Berrak Sisman, and Haizhou Li · 2021
Later among the works it cites.
Fastspeech 2: Fast and high-quality end-to-end text to speech
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Glow-tts: A generative flow for text-to-speech via monotonic alignment search
Jaehyeon Kim, Sungwon Kim, Jungil Kong, and Sungroh Yoon · 2020
Cited alongside, same era.
Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis
Jungil Kong, Jaehyeon Kim, and Jaekyoung Bae · 2020
Cited alongside, same era.
Non-autoregressive neural text-to-speech
Kainan Peng, Wei Ping, Zhao Song, and Kexin Zhao · 2020
Cited alongside, same era.
Stanza: A Python natural language processing toolkit for many human languages
Peng Qi, Yuhao Zhang, Yuhui Zhang, Jason Bolton, and Christopher D. Manning · 2020
Cited alongside, same era.
Yi Ren, Chenxu Hu, Xu Tan, Tao Qin, Sheng Zhao, Zhou Zhao, and Tie-Yan Liu · 2021
Later among the works it cites.
Portaspeech: Portable and high-quality generative text-to-speech
Yi Ren, Jinglin Liu, and Zhou Zhao · 2021
Later among the works it cites.
Graphpb: Graphical representations of prosody boundary in speech synthesis
Aolan Sun, Jianzong Wang, Ning Cheng, Huayi Peng, Zhen Zeng, Lingwei Kong, and Jing Xiao · 2021
Later among the works it cites.
Yixuan Zhou, Changhe Song, Jingbei Li, Zhiyong Wu, and Helen Meng · 2021
Later among the works it cites.