Reverse-time diffusion equation models
Anderson, B. D. O · 1982
Earlier work this paper cites.
Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks
Graves, A., Fernández, S., Gomez, F., and Schmidhuber, J · 2006
Earlier work this paper cites.
MNIST handwritten digit database
LeCun, Y. and Cortes, C · 2010
Earlier work this paper cites.
The blizzard challenge 2013
King, S. J. and Karaiskos, V · 2013
Earlier work this paper cites.
Generative adversarial nets
Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y · 2014
Earlier work this paper cites.
Auto-encoding variational Bayes
Kingma, D. P. and Welling, M · 2014
Earlier work this paper cites.
Librispeech: An asr corpus based on public domain audio books
Panayotov, V., Chen, G., Povey, D., and Khudanpur, S · 2015
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation
Ronneberger, O., Fischer, P., and Brox, T · 2015
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
Sohl-Dickstein, J., Weiss, E., Maheswaranathan, N., and Ganguli, S · 2015
Earlier work this paper cites.
Wavenet: A generative model for raw audio
Original
van den Oord, A., Dieleman, S., Zen, H., Simonyan, K., Vinyals, O., Graves, A., Kalchbrenner, N., Senior, A. W., and Kavukcuoglu, K · 2016
Earlier work this paper cites.
The lj speech dataset
Ito, K · 2017
Earlier work this paper cites.
Montreal forced aligner: Trainable text-speech alignment using kaldi
McAuliffe, M., Socolof, M., Mihuc, S., Wagner, M., and Sonderegger, M · 2017
Earlier work this paper cites.
Neural discrete representation learning
van den Oord, A., Vinyals, O., et al · 2017
Earlier work this paper cites.
Tacotron: Towards End-to-End Speech Synthesis
Wang, Y., Skerry-Ryan, R., Stanton, D., Wu, Y., Weiss, R. J., Jaitly, N., Yang, Z., Xiao, Y., Chen, Z., Bengio, S., Le, Q., Agiomyrgiannakis, Y., Clark, R., and Saurous, R. A · 2017
Earlier work this paper cites.
Neural voice cloning with a few samples
Arik, S., Chen, J., Peng, K., Ping, W., and Zhou, Y · 2018
Earlier work this paper cites.
Voxceleb2: Deep speaker recognition
Chung, J. S., Nagrani, A., and Zisserman, A · 2018
Earlier work this paper cites.
Transfer learning from speaker verification to multispeaker text-to-speech synthesis
Jia, Y., Zhang, Y., Weiss, R., Wang, Q., Shen, J., Ren, F., Chen, z., Nguyen, P., Pang, R., Lopez Moreno, I., and Wu, Y · 2018
Earlier work this paper cites.
Efficient neural audio synthesis
Kalchbrenner, N., Elsen, E., Simonyan, K., Noury, S., Casagrande, N., Lockhart, E., Stimberg, F., Oord, A., Dieleman, S., and Kavukcuoglu, K · 2018
Earlier work this paper cites.
Glow: Generative flow with invertible 1x1 convolutions
Kingma, D. P. and Dhariwal, P · 2018
Earlier work this paper cites.