Improved variational inference with inverse autoregressive flow
Kingma, D. P., Salimans, T., Jozefowicz, R., Chen, X., Sutskever, I., and Welling, M · 2016
Cited alongside, same era.
Singing voice synthesis based on deep neural networks
Nishimura, M., Hashimoto, K., Oura, K., Nankaku, Y., and Tokuda, K · 2016
Cited alongside, same era.
Deep voice 2: Multi-speaker neural text-to-speech
Gibiansky, A., Arik, S., Diamos, G., Miller, J., Peng, K., Ping, W., Raiman, J., and Zhou, Y · 2017
Cited alongside, same era.
The LJ speech dataset, 2017
Ito, K. et al · 2017
Cited alongside, same era.
Deep voice 3: 2000-speaker neural text-to-speech
Original
Ping, W., Peng, K., Gibiansky, A., Arik, S. O., Kannan, A., Narang, S., Raiman, J., and Miller, J · 2017
Cited alongside, same era.
Natural tts synthesis by conditioning wavenet on mel spectrogram predictions
Original
Shen, J., Pang, R., Weiss, R. J., Schuster, M., Jaitly, N., Yang, Z., Chen, Z., Zhang, Y., Wang, Y., Skerry-Ryan, R., et al · 2017
Cited alongside, same era.
Tacotron: A fully end-to-end text-to-speech synthesis model
Original
Wang, Y., Skerry-Ryan, R., Stanton, D., Wu, Y., Weiss, R. J., Jaitly, N., Yang, Z., Xiao, Y., Chen, Z., Bengio, S., et al · 2017
Cited alongside, same era.
Expressive speech synthesis via modeling expressions with variational autoencoder
Original
Akuzawa, K., Iwasawa, Y., and Matsuo, Y · 2018
Cited alongside, same era.
Hierarchical generative modeling for controllable speech synthesis
Original
Hsu, W.-N., Zhang, Y., Weiss, R. J., Zen, H., Wu, Y., Wang, Y., Cao, Y., Jia, Y., Chen, Z., Shen, J., et al · 2018
Cited alongside, same era.
Deep voice 2: Multi-speaker neural text-to-speech
Original
Arik, S., Diamos, G., Gibiansky, A., Miller, J., Peng, K., Ping, W., Raiman, J., and Zhou, Y
Cited in the paper.
Deep voice: Real-time neural text-to-speech
Original
Arik, S. O., Chrzanowski, M., Coates, A., Diamos, G., Gibiansky, A., Kang, Y., Li, X., Miller, J., Ng, A., Raiman, J., et al
Cited in the paper.
Mellotron github repo, 2019a
Valle, R., Li, J., Prenger, R., and Catanzaro, B
Cited in the paper.