Fetching the paper…
Reading the bibliography…
Recent work has explored sequence-to-sequence latent variable models for expressive speech synthesis (supporting control and transfer of prosody and style), but has not presented a coherent framework for understanding the trade-offs between the competing methods.
Mel-cepstral distance measure for objective speech quality assessment
R Kubichek · 1993
Earlier work this paper cites.
Dynamic time warping
Meinard Müller · 2007
Earlier work this paper cites.
Experimental and theoretical advances in prosody: A review
Michael Wagner and Duane G Watson · 2010
Earlier work this paper cites.
Generating sequences with recurrent neural networks
Alex Graves · 2013
Earlier work this paper cites.
Auto-Encoding Variational Bayes
Diederik P Kingma and Max Welling · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Alireza Makhzani, Jonathon Shlens, Navdeep Jaitly, Ian Goodfellow, and Brendan Frey · 2015
Earlier work this paper cites.
Variational inference with normalizing flows
Danilo Rezende and Shakir Mohamed · 2015
Earlier work this paper cites.
Elbo surgery: yet another way to carve up the variational evidence lower bound
Matthew D Hoffman and Matthew J Johnson · 2016
Earlier work this paper cites.
Ladder variational autoencoders
Casper Kaae Sønderby, Tapani Raiko, Lars Maaløe, Søren Kaae Sønderby, and Ole Winther · 2016
Earlier work this paper cites.
Wavenet: A generative model for raw audio
Aäron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior, and Koray Kavukcuoglu · 2016
Cited alongside, same era.
beta-vae: Learning basic visual concepts with a constrained variational framework
Irina Higgins, Loic Matthey, Arka Pal, Christopher Burgess, Xavier Glorot, Matthew Botvinick, Shakir Mohamed, and Alexander Lerchner · 2017
Cited alongside, same era.
Char2wav: End-to-end speech synthesis
Jose Sotelo, Soroush Mehri, Kundan Kumar, Joao Felipe Santos, Kyle Kastner, Aaron Courville, and Yoshua Bengio · 2017
Cited alongside, same era.
Tacotron: Towards end-to-end speech synthesis
Yuxuan Wang, RJ Skerry-Ryan, Daisy Stanton, Yonghui Wu, Ron J. Weiss, Navdeep Jaitly, Zongheng Yang, Ying Xiao, Zhifeng Chen, Samy Bengio, Quoc Le, Yannis Agiomyrgiannakis, Rob Clark, and Rif A. Saurous · 2017
Cited alongside, same era.
Fixing a broken elbo
Alexander Alemi, Ben Poole, Ian Fischer, Joshua Dillon, Rif A Saurous, and Kevin Murphy · 2018
Cited alongside, same era.
Towards end-to-end prosody transfer for expressive speech synthesis with Tacotron
RJ Skerry-Ryan, Eric Battenberg, Ying Xiao, Yuxuan Wang, Daisy Stanton, Joel Shor, Ron Weiss, Rob Clark, and Rif A. Saurous · 2018
Later among the works it cites.
Predicting expressive speaking style from text in end-to-end speech synthesis
Daisy Stanton, Yuxuan Wang, and RJ Skerry-Ryan · 2018
Later among the works it cites.
Voiceloop: Voice fitting and synthesis via a phonological loop
Yaniv Taigman, Lior Wolf, Adam Polyak, and Eliya Nachmani · 2018
Later among the works it cites.
Style tokens: Unsupervised style modeling, control and transfer in end-to-end speech synthesis
Yuxuan Wang, Daisy Stanton, Yu Zhang, RJ-Skerry Ryan, Eric Battenberg, Joel Shor, Ying Xiao, Ye Jia, Fei Ren, and Rif A Saurous · 2018
Later among the works it cites.
Location-relative attention mechanisms for robust long-form speech synthesis
Eric Battenberg, RJ Skerry-Ryan, Soroosh Mariooryad, Daisy Stanton, David Kao, Matt Shannon, and Tom Bagby · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Christopher P Burgess, Irina Higgins, Arka Pal, Loic Matthey, Nick Watters, Guillaume Desjardins, and Alexander Lerchner · 2018
Cited alongside, same era.
Deep encoder-decoder models for unsupervised learning of controllable speech synthesis
Gustav Eje Henter, Jaime Lorenzo-Trueba, Xin Wang, and Junichi Yamagishi · 2018
Cited alongside, same era.
Deep voice 3: 2000-speaker neural text-to-speech
Wei Ping, Kainan Peng, Andrew Gibiansky, Sercan O. Arik, Ajay Kannan, Sharan Narang, Jonathan Raiman, and John Miller · 2018
Cited alongside, same era.
Danilo Jimenez Rezende and Fabio Viola · 2018
Cited alongside, same era.
Natural tts synthesis by conditioning wavenet on mel spectrogram predictions
Jonathan Shen, Ruoming Pang, Ron J Weiss, Mike Schuster, Navdeep Jaitly, Zongheng Yang, Zhifeng Chen, Yu Zhang, Yuxuan Wang, Rj Skerrv-Ryan, et al · 2018
Cited alongside, same era.
Closest in time.
Hierarchical generative modeling for controllable speech synthesis
Wei-Ning Hsu, Yu Zhang, Ron Weiss, Heiga Zen, Yonghui Wu, Yuan Cao, and Yuxuan Wang · 2019
Closest in time.
Robust and fine-grained prosody control of end-to-end speech synthesis
Younggun Lee and Taesu Kim · 2019
Closest in time.
A generative adversarial network for style modeling in a text-to-speech system
Shuang Ma, Daniel Mcduff, and Yale Song · 2019
Closest in time.
Learning latent representations for style control and transfer in end-to-end speech synthesis
Ya-Jie Zhang, Shifeng Pan, Lei He, and Zhen-Hua Ling · 2019
Closest in time.