Fetching the paper…
Reading the bibliography…
Several recent end-to-end text-to-speech (TTS) models enabling single-stage training and parallel sampling have been proposed, but their sample quality does not match that of two-stage TTS systems.
Yin, a fundamental frequency estimator for speech and music
De Cheveigné, A. and Kawahara, H · 2002
Earlier work this paper cites.
Generating sequences with recurrent neural networks
Graves, A · 2013
Earlier work this paper cites.
Generative Adversarial Nets
Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y · 2014
Earlier work this paper cites.
Auto-Encoding Variational Bayes
Kingma, D. P. and Welling, M · 2014
Earlier work this paper cites.
Variational inference with normalizing flows
Rezende, D. and Mohamed, S · 2015
Earlier work this paper cites.
Improved variational inference with inverse autoregressive flow
Kingma, D. P., Salimans, T., Jozefowicz, R., Chen, X., Sutskever, I., and Welling, M · 2016
Earlier work this paper cites.
Autoencoding beyond pixels using a learned similarity metric
Larsen, A. B. L., Sønderby, S. K., Larochelle, H., and Winther, O · 2016
Earlier work this paper cites.
Wavenet: A generative model for raw audio
Oord, A. v. d., Dieleman, S., Zen, H., Simonyan, K., Vinyals, O., Graves, A., Kalchbrenner, N., Senior, A., and Kavukcuoglu, K · 2016
Earlier work this paper cites.
Variational lossy autoencoder
Chen, X., Kingma, D. P., Salimans, T., Duan, Y., Dhariwal, P., Schulman, J., Sutskever, I., and Abbeel, P · 2017
Earlier work this paper cites.
Density estimation using Real NVP
Dinh, L., Sohl-Dickstein, J., and Bengio, S · 2017
Earlier work this paper cites.
The LJ Speech Dataset
Ito, K · 2017
Earlier work this paper cites.
Least squares generative adversarial networks
Mao, X., Li, Q., Xie, H., Lau, R. Y., Wang, Z., and Paul Smolley, S · 2017
Earlier work this paper cites.
Neural discrete representation learning
van den Oord, A., Vinyals, O., and Kavukcuoglu, K · 2017
Earlier work this paper cites.
Attention is All you Need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
CSTR VCTK corpus: English multi-speaker corpus for CSTR voice cloning toolkit
Veaux, C., Yamagishi, J., MacDonald, K., et al · 2017
Earlier work this paper cites.
Transfer Learning from Speaker Verification to Multispeaker Text-To-Speech Synthesis
Jia, Y., Zhang, Y., Weiss, R. J., Wang, Q., Shen, J., Ren, F., Chen, Z., Nguyen, P., Pang, R., Lopez-Moreno, I., et al · 2018
Cited alongside, same era.
Efficient neural audio synthesis
Kalchbrenner, N., Elsen, E., Simonyan, K., Noury, S., Casagrande, N., Lockhart, E., Stimberg, F., Oord, A., Dieleman, S., and Kavukcuoglu, K · 2018
Cited alongside, same era.
Deep Voice 3: 2000-Speaker Neural Text-to-Speech
Ping, W., Peng, K., Gibiansky, A., Arik, S. O., Kannan, A., Narang, S., Raiman, J., and Miller, J · 2018
Cited alongside, same era.
Self-Attention with Relative Position Representations
Shaw, P., Uszkoreit, J., and Vaswani, A · 2018
Cited alongside, same era.
Natural tts synthesis by conditioning wavenet on mel spectrogram predictions
Shen, J., Pang, R., Weiss, R. J., Schuster, M., Jaitly, N., Yang, Z., Chen, Z., Zhang, Y., Wang, Y., Skerrv-Ryan, R., et al · 2018
Cited alongside, same era.
FastSpeech: Fast, Robust and Controllable Text to Speech
Ren, Y., Ruan, Y., Tan, X., Qin, T., Zhao, S., Zhao, Z., and Liu, T.-Y · 2019
Later among the works it cites.
Learning latent representations for style control and transfer in end-to-end speech synthesis
Zhang, Y.-J., Pan, S., He, L., and Ling, Z.-H · 2019
Later among the works it cites.
Latent normalizing flows for discrete sequences
Ziegler, Z. and Rush, A · 2019
Later among the works it cites.
Vflow: More expressive generative flows with variational data augmentation
Chen, J., Lu, C., Chenli, B., Zhu, J., and Tian, T · 2020
Later among the works it cites.
Glow-TTS: A Generative Flow for Text-to-Speech via Monotonic Alignment Search
Kim, J., Kim, S., Kong, J., and Yoon, S · 2020
Later among the works it cites.
HiFi-GAN: Generative Adversarial networks for Efficient and High Fidelity Speech Synthesis
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Taigman, Y., Wolf, L., Polyak, A., and Nachmani, E · 2018
Cited alongside, same era.
High Fidelity Speech Synthesis with Adversarial Networks
Bińkowski, M., Donahue, J., Dieleman, S., Clark, A., Elsen, E., Casagrande, N., Cobo, L. C., and Simonyan, K · 2019
Cited alongside, same era.
Neural Spline Flows
Durkan, C., Bekasov, A., Murray, I., and Papamakarios, G · 2019
Cited alongside, same era.
Flow++: Improving flow-based generative models with variational dequantization and architecture design
Ho, J., Chen, X., Srinivas, A., Duan, Y., and Abbeel, P · 2019
Cited alongside, same era.
Hierarchical Generative Modeling for Controllable Speech Synthesis
Hsu, W.-N., Zhang, Y., Weiss, R., Zen, H., Wu, Y., Cao, Y., and Wang, Y · 2019
Cited alongside, same era.
MelGAN: Generative Adversarial Networks for Conditional waveform synthesis
Kumar, K., Kumar, R., de Boissiere, T., Gestin, L., Teoh, W. Z., Sotelo, J., de Brébisson, A., Bengio, Y., and Courville, A. C · 2019
Cited alongside, same era.
Neural speech synthesis with transformer network
Li, N., Liu, S., Liu, Y., Zhao, S., and Liu, M · 2019
Cited alongside, same era.
Kong, J., Kim, J., and Bae, J · 2020
Later among the works it cites.
Flow-TTS: A non-autoregressive network for text to speech based on flow
Miao, C., Liang, S., Chen, M., Ma, J., Wang, S., and Xiao, J · 2020
Later among the works it cites.
Non-autoregressive neural text-to-speech
Peng, K., Ping, W., Song, Z., and Zhao, K · 2020
Later among the works it cites.
Wave-Tacotron: Spectrogram-free end-to-end text-to-speech synthesis
Weiss, R. J., Skerry-Ryan, R., Battenberg, E., Mariooryad, S., and Kingma, D. P · 2020
Later among the works it cites.
Aligntts: Efficient feed-forward text-to-speech system without explicit alignment
Zeng, Z., Wang, J., Cheng, N., Xia, T., and Xiao, J · 2020
Later among the works it cites.
Phonemizer
Bernard, M · 2021
Closest in time.
End-to-end Adversarial Text-to-Speech
Donahue, J., Dieleman, S., Binkowski, M., Elsen, E., and Simonyan, K · 2021
Closest in time.
Bidirectional Variational Inference for Non-Autoregressive Text-to-speech
Lee, Y., Shin, J., and Jung, K · 2021
Closest in time.
FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
Ren, Y., Hu, C., Tan, X., Qin, T., Zhao, S., Zhao, Z., and Liu, T.-Y · 2021
Closest in time.
Flowtron: an Autoregressive Flow-based Generative Network for Text-to-Speech Synthesis
Valle, R., Shih, K. J., Prenger, R., and Catanzaro, B · 2021
Closest in time.