Fetching the paper…
Reading the bibliography…
We present a novel generative model that combines state-of-the-art neural text-to-speech (TTS) with semi-supervised probabilistic latent variable models.
A circumplex model of affect
J. A. Russell · 1980
Earlier work this paper cites.
Mel-cepstral distance measure for objective speech quality assessment
R Kubichek · 1993
Earlier work this paper cites.
Nonlinear independent component analysis: Existence and uniqueness results
A. Hyvärinen and P. Pajunen · 1999
Earlier work this paper cites.
Emotional speech synthesis: A review
M. Schröder · 2001
Earlier work this paper cites.
Yin, a fundamental frequency estimator for speech and music
A. De Cheveigné and H. Kawahara · 2002
Earlier work this paper cites.
Generating sequences with recurrent neural networks
A. Graves · 2013
Earlier work this paper cites.
Auto-encoding variational bayes
D. P. Kingma and M. Welling · 2013
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
D. Bahdanau, K. Cho, and Y. Bengio · 2014
Earlier work this paper cites.
Semi-supervised learning with deep generative models
D. P. Kingma, S. Mohamed, D. J. Rezende, and M. Welling · 2014
Earlier work this paper cites.
Stochastic backpropagation and approximate inference in deep generative models
D. J. Rezende, S. Mohamed, and D. Wierstra · 2014
Earlier work this paper cites.
Generating sentences from a continuous space
S. Bowman, L. Vilnis, O. Vinyals, A. M. Dai, R. Jozefowicz, and S. Bengio · 2015
Earlier work this paper cites.
Disentangling factors of variation in deep representation using adversarial training
M. F. Mathieu, J. J. Zhao, J. Zhao, A. Ramesh, P. Sprechmann, and Y. LeCun · 2016
Earlier work this paper cites.
Wavenet: A generative model for raw audio
A. van den Oord, S. Dieleman, H. Zen, K. Simonyan, O. Vinyals, A. Graves, N. Kalchbrenner, A. Senior, and K. Kavukcuoglu · 2016
Cited alongside, same era.
Deep voice: Real-time neural text-to-speech
S. Ö. Arik, M. Chrzanowski, A. Coates, G. Diamos, A. Gibiansky, Y. Kang, X. Li, J. Miller, A. Ng, and J. Raiman · 2017
Cited alongside, same era.
Deep voice 2: Multi-speaker neural text-to-speech
A. Gibiansky, S. Arik, G. Diamos, J. Miller, K. Peng, W. Ping, J. Raiman, and Y. Zhou · 2017
Cited alongside, same era.
beta-vae: Learning basic visual concepts with a constrained variational framework
I. Higgins, L. Matthey, A. Pal, C. Burgess, X. Glorot, M. Botvinick, S. Mohamed, and A. Lerchner · 2017
Cited alongside, same era.
Online measuring of available resources
E. Munoz-de Escalona and J. J. Canas · 2017
Cited alongside, same era.
Efficient neural audio synthesis
N. Kalchbrenner, E. Elsen, K. Simonyan, S. Noury, N. Casagrande, E. Lockhart, F. Stimberg, A. Van den Oord, S. Dieleman, and K. Kavukcuoglu · 2018
Later among the works it cites.
H. Kim and A. Mnih · 2018
Later among the works it cites.
Natural tts synthesis by conditioning wavenet on mel spectrogram predictions
J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, and R. Skerrv-Ryan · 2018
Later among the works it cites.
Towards end-to-end prosody transfer for expressive speech synthesis with tacotron
R. Skerry-Ryan, E. Battenberg, Y. Xiao, Y. Wang, D. Stanton, J. Shor, R. Weiss, R. Clark, and R. A. Saurous · 2018
Later among the works it cites.
Style tokens: Unsupervised style modeling, control and transfer in end-to-end speech synthesis
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Learning disentangled representations with semi-supervised deep generative models
S. Narayanaswamy, B. T. Paige, J. Van de Meent, A. Desmaison, N. Goodman, P. Kohli, F. Wood, and P. Torr · 2017
Cited alongside, same era.
Deep voice 3: Scaling text-to-speech with convolutional sequence learning
W. Ping, K. Peng, A. Gibiansky, S. O. Arik, A. Kannan, S. Narang, J. Raiman, and J. Miller · 2017
Cited alongside, same era.
Voiceloop: Voice fitting and synthesis via a phonological loop
Y. Taigman, L. Wolf, A. Polyak, and E. Nachmani · 2017
Cited alongside, same era.
Tacotron: Towards end-to-end speech synthesis
Y. Wang, RJ. Skerry-Ryan, D. Stanton, Y. Wu, R. J. Weiss, N. Jaitly, Z. Yang, Y. Xiao, Z. Chen, S. Bengio, et al · 2017
Cited alongside, same era.
Expressive speech synthesis via modeling expressions with variational autoencoder
K. Akuzawa, Y. Iwasawa, and Y. Matsuo · 2018
Cited alongside, same era.
Hierarchical generative modeling for controllable speech synthesis
W. Hsu, Y. Zhang, R. J. Weiss, H. Zen, Y. Wu, Y. Wang, Y. Cao, Y. Jia, Z. Chen, J. Shen, et al · 2018
Cited alongside, same era.
Y. Wang, D. Stanton, Y. Zhang, R. J. Skerry-Ryan, E. Battenberg, J. Shor, Y. Xiao, F Ren, Y. Jia, and R. A. Saurous · 2018
Later among the works it cites.
End-to-end emotional speech synthesis using style tokens and semi-supervised training
P. Wu, Z. Ling, L. Liu, Y. Jiang, H. Wu, and L. Dai · 2018
Later among the works it cites.
Effective use of variational embedding capacity in expressive end-to-end speech synthesis
E. Battenberg, S. Mariooryad, D. Stanton, R. Skerry-Ryan, M. Shannon, D. Kao, and T. Bagby · 2019
Closest in time.
Diva: Domain invariant variational autoencoders
M. Ilse, J. M. Tomczak, C. Louizos, and M. Welling · 2019
Closest in time.
Fastspeech: Fast, robust and controllable text to speech
Y. Ren, Y. Ruan, X. Tan, T. Qin, S. Zhao, Z. Zhao, and T. Liu · 2019
Closest in time.
Melnet: A generative model for audio in the frequency domain
S. Vasquez and M. Lewis · 2019
Closest in time.
V. Wan, C. Chan, T. Kenter, J. Vit, and R. Clark · 2019
Closest in time.