Fetching the paper…
Reading the bibliography…
Some recent studies have demonstrated the feasibility of single-stage neural text-to-speech, which does not need to generate mel-spectrograms but generates the raw waveforms directly from the text.
Signal estimation from modified short-time fourier transform
Daniel Griffin and Jae Lim · 1984
Earlier work this paper cites.
Auto-encoding variational bayes
Diederik P Kingma and Max Welling · 2013
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Earlier work this paper cites.
Variational inference with normalizing flows
Danilo Rezende and Shakir Mohamed · 2015
Earlier work this paper cites.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton · 2016
Earlier work this paper cites.
Xi Chen, Diederik P Kingma, Tim Salimans, Yan Duan, Prafulla Dhariwal, John Schulman, Ilya Sutskever, and Pieter Abbeel · 2016
Earlier work this paper cites.
Wavenet: A generative model for raw audio
Aaron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior, and Koray Kavukcuoglu · 2016
Earlier work this paper cites.
Deep voice: Real-time neural text-to-speech
Sercan Ö Arık, Mike Chrzanowski, Adam Coates, Gregory Diamos, Andrew Gibiansky, Yongguo Kang, Xian Li, John Miller, Andrew Ng, Jonathan Raiman, et al · 2017
Earlier work this paper cites.
Density estimation using real NVP
Laurent Dinh, Jascha Sohl-Dickstein, and Samy Bengio · 2017
Earlier work this paper cites.
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Tacotron: Towards end-to-end speech synthesis
Yuxuan Wang, RJ Skerry-Ryan, Daisy Stanton, Yonghui Wu, Ron J Weiss, Navdeep Jaitly, Zongheng Yang, Ying Xiao, Zhifeng Chen, Samy Bengio, et al · 2017
Earlier work this paper cites.
Expressive speech synthesis via modeling expressions with variational autoencoder
Kei Akuzawa, Yusuke Iwasawa, and Yutaka Matsuo · 2018
Earlier work this paper cites.
Hierarchical generative modeling for controllable speech synthesis
Wei-Ning Hsu, Yu Zhang, Ron J Weiss, Heiga Zen, Yonghui Wu, Yuxuan Wang, Yuan Cao, Ye Jia, Zhifeng Chen, Jonathan Shen, et al · 2018
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2018
Earlier work this paper cites.
Deep voice 3: 2000-speaker neural text-to-speech
Wei Ping, Kainan Peng, Andrew Gibiansky, Sercan O Arik, Ajay Kannan, Sharan Narang, Jonathan Raiman, and John Miller · 2018
Cited alongside, same era.
Natural tts synthesis by conditioning wavenet on mel spectrogram predictions
Jonathan Shen, Ruoming Pang, Ron J Weiss, Mike Schuster, Navdeep Jaitly, Zongheng Yang, Zhifeng Chen, Yu Zhang, Yuxuan Wang, Rj Skerrv-Ryan, et al · 2018
Cited alongside, same era.
Towards end-to-end prosody transfer for expressive speech synthesis with tacotron
RJ Skerry-Ryan, Eric Battenberg, Ying Xiao, Yuxuan Wang, Daisy Stanton, Joel Shor, Ron Weiss, Rob Clark, and Rif A Saurous · 2018
Cited alongside, same era.
Style tokens: Unsupervised style modeling, control and transfer in end-to-end speech synthesis
Yuxuan Wang, Daisy Stanton, Yu Zhang, RJ-Skerry Ryan, Eric Battenberg, Joel Shor, Ying Xiao, Ye Jia, Fei Ren, and Rif A Saurous · 2018
Cited alongside, same era.
Melgan: Generative adversarial networks for conditional waveform synthesis
Kundan Kumar, Rithesh Kumar, Thibault de Boissiere, Lucas Gestin, Wei Zhen Teoh, Jose Sotelo, Alexandre de Brébisson, Yoshua Bengio, and Aaron Courville · 2019
Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis
Jungil Kong, Jaehyeon Kim, and Jaekyoung Bae · 2020
Later among the works it cites.
Non-autoregressive neural text-to-speech
Kainan Peng, Wei Ping, Zhao Song, and Kexin Zhao · 2020
Later among the works it cites.
Fully-hierarchical fine-grained prosody modeling for interpretable speech synthesis
Guangzhi Sun, Yu Zhang, Ron J Weiss, Yuan Cao, Heiga Zen, and Yonghui Wu · 2020
Later among the works it cites.
Qiao Tian, Yi Chen, Zewang Zhang, Heng Lu, Linghui Chen, Lei Xie, and Shan Liu · 2020
Later among the works it cites.
Flowtron: an autoregressive flow-based generative network for text-to-speech synthesis
Rafael Valle, Kevin J Shih, Ryan Prenger, and Bryan Catanzaro · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Neural speech synthesis with transformer network
Naihan Li, Shujie Liu, Yanqing Liu, Sheng Zhao, and Ming Liu · 2019
Cited alongside, same era.
Waveglow: A flow-based generative network for speech synthesis
Ryan Prenger, Rafael Valle, and Bryan Catanzaro · 2019
Cited alongside, same era.
Fastspeech: Fast, robust and controllable text to speech
Yi Ren, Yangjun Ruan, Xu Tan, Tao Qin, Sheng Zhao, Zhou Zhao, and Tie-Yan Liu · 2019
Cited alongside, same era.
Learning latent representations for style control and transfer in end-to-end speech synthesis
Ya-Jie Zhang, Shifeng Pan, Lei He, and Zhen-Hua Ling · 2019
Cited alongside, same era.
Latent normalizing flows for discrete sequences
Zachary Ziegler and Alexander Rush · 2019
Cited alongside, same era.
Location-relative attention mechanisms for robust long-form speech synthesis
Eric Battenberg, R.J. Skerry-Ryan, Soroosh Mariooryad, Daisy Stanton, David Kao, Matt Shannon, and Tom Bagby · 2020
Cited alongside, same era.
End-to-end adversarial text-to-speech
Jeff Donahue, Sander Dieleman, Mikolaj Binkowski, Erich Elsen, and Karen Simonyan · 2020
Cited alongside, same era.
Parallel wavegan: A fast waveform generation model based on generative adversarial networks with multi-resolution spectrogram
Ryuichi Yamamoto, Eunwoo Song, and Jae-Min Kim · 2020
Later among the works it cites.
Rich prosody diversity modelling with phone-level mixture density network
Chenpeng Du and Kai Yu · 2021
Later among the works it cites.
Parallel tacotron: Non-autoregressive and controllable tts
Isaac Elias, Heiga Zen, Jonathan Shen, Yu Zhang, Ye Jia, Ron J Weiss, and Yonghui Wu · 2021
Later among the works it cites.
Neural dubber: Dubbing for videos according to scripts
Chenxu Hu, Qiao Tian, Tingle Li, Wang Yuping, Yuxuan Wang, and Hang Zhao · 2021
Later among the works it cites.
Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech
Jaehyeon Kim, Jungil Kong, and Juhee Son · 2021
Later among the works it cites.
Stylemelgan: An efficient high-fidelity adversarial vocoder with temporal adaptive normalization
Ahmed Mustafa, Nicola Pia, and Guillaume Fuchs · 2021
Later among the works it cites.
Fastspeech 2: Fast and high-quality end-to-end text to speech
Yi Ren, Chenxu Hu, Xu Tan, Tao Qin, Sheng Zhao, Zhou Zhao, and Tie-Yan Liu · 2021
Later among the works it cites.
Unsupervised word-level prosody tagging for controllable speech synthesis
Yiwei Guo, Chenpeng Du, and Kai Yu · 2022
Closest in time.