Fetching the paper…
Reading the bibliography…
In this work, we propose "global style tokens" (GSTs), a bank of embeddings that are jointly trained within Tacotron, a state-of-the-art end-to-end speech synthesis system.
Signal estimation from modified short-time fourier transform
Griffin, Daniel and Lim, Jae · 1984
Earlier work this paper cites.
ToBI: A standard for labeling english prosody
Silverman, Kim, Beckman, Mary, Pitrelli, John, Ostendorf, Mori, Wightman, Colin, Price, Patti, Pierrehumbert, Janet, and Hirschberg, Julia · 1992
Earlier work this paper cites.
Tobi or not tobi?
Wightman, Colin W · 2002
Earlier work this paper cites.
Visualizing data using t-sne
Maaten, Laurens van der and Hinton, Geoffrey · 2008
Earlier work this paper cites.
Text-to-speech synthesis
Taylor, Paul · 2009
Earlier work this paper cites.
AuToBI-a tool for automatic ToBI annotation
Rosenberg, Andrew · 2010
Earlier work this paper cites.
Front-end factor analysis for speaker verification
Dehak, Najim, Kenny, Patrick J, Dehak, Réda, Dumouchel, Pierre, and Ouellet, Pierre · 2011
Earlier work this paper cites.
Unsupervised clustering of emotion and voice styles for expressive tts
Eyben, Florian, Buchholz, Sabine, and Braunschweiler, Norbert · 2012
Earlier work this paper cites.
Conditional restricted boltzmann machine for voice conversion
Wu, Zhizheng, Chng, Eng Siong, and Li, Haizhou · 2013
Earlier work this paper cites.
Graves, Alex, Wayne, Greg, and Danihelka, Ivo · 2014
Cited alongside, same era.
A note on the evaluation of generative models
Theis, Lucas, Oord, Aäron van den, and Bethge, Matthias · 2015
Cited alongside, same era.
Non-parallel training in voice conversion using an adaptive restricted boltzmann machine
Nakashika, Toru, Takiguchi, Tetsuya, Minami, Yasuhiro, Nakashika, Toru, Takiguchi, Tetsuya, and Minami, Yasuhiro · 2016
Cited alongside, same era.
Wavenet: A generative model for raw audio
van den Oord, Aäron, Dieleman, Sander, Zen, Heiga, Simonyan, Karen, Vinyals, Oriol, Graves, Alex, Kalchbrenner, Nal, Senior, Andrew, and Kavukcuoglu, Koray · 2016
Cited alongside, same era.
Fast, compact, and high quality LSTM-RNN based statistical parametric speech synthesizers for mobile devices
Non-parallel voice conversion using i-vector plda: Towards unifying speaker verification and transformation
Kinnunen, Tomi, Juvela, Lauri, Alku, Paavo, and Yamagishi, Junichi · 2017
Later among the works it cites.
Zoneout: Regularizing RNNs by randomly preserving hidden activations
Krueger, David, Maharaj, Tegan, Kramár, János, Pezeshki, Mohammad, Ballas, Nicolas, Ke, Nan Rosemary, Goyal, Anirudh, Bengio, Yoshua, Larochelle, Hugo, Courville, Aaron, et al · 2017
Later among the works it cites.
Adapting and controlling dnn-based speech synthesis using input codes
Luong, Hieu-Thi, Takaki, Shinji, Henter, Gustav Eje, and Yamagishi, Junichi · 2017
Later among the works it cites.
Deep voice 3: 2000-speaker neural text-to-speech
Ping, Wei, Peng, Kainan, Gibiansky, Andrew, Arik, Sercan O, Kannan, Ajay, Narang, Sharan, Raiman, Jonathan, and Miller, John · 2017
Later among the works it cites.
Natural tts synthesis by conditioning wavenet on mel spectrogram predictions
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zen, Heiga, Agiomyrgiannakis, Yannis, Egberts, Niels, Henderson, Fergus, and Szczepaniak, Przemysław · 2016
Cited alongside, same era.
Deep voice: Real-time neural text-to-speech
Arik, Sercan O, Chrzanowski, Mike, Coates, Adam, Diamos, Gregory, Gibiansky, Andrew, Kang, Yongguo, Li, Xian, Miller, John, Raiman, Jonathan, Sengupta, Shubho, et al · 2017
Cited alongside, same era.
Unsupervised learning of disentangled and interpretable representations from sequential data
Hsu, Wei-Ning, Zhang, Yu, and Glass, James · 2017
Cited alongside, same era.
Unsupervised learning for expressive speech synthesis
Jauk, Igor · 2017
Cited alongside, same era.
Generation of large-scale simulated utterances in virtual rooms to train deep-neural networks for far-field speech recognition in google home
Kim, Chanwoo, Misra, Ananya, Chin, Kean, Hughes, Thad, Narayanan, Arun, Sainath, Tara, and Bacchiani, Michiel · 2017
Cited alongside, same era.
Tacotron: Towards end-to-end speech synthesis
Wang, Yuxuan, Skerry-Ryan, RJ, Stanton, Daisy, Wu, Yonghui, Weiss, Ron J., Jaitly, Navdeep, Yang, Zongheng, Xiao, Ying, Chen, Zhifeng, Bengio, Samy, Le, Quoc, Agiomyrgiannakis, Yannis, Clark, Rob, and Saurous, Rif A
Cited in the paper.
Uncovering latent style factors for expressive speech synthesis
Wang, Yuxuan, Skerry-Ryan, RJ, Xiao, Ying, Stanton, Daisy, Shor, Joel, Battenberg, Eric, Clark, Rob, and Saurous, Rif A
Cited in the paper.
Shen, Jonathan, Pang, Ruoming, Weiss, Ron J, Schuster, Mike, Jaitly, Navdeep, Yang, Zongheng, Chen, Zhifeng, Zhang, Yu, Wang, Yuxuan, Skerry-Ryan, RJ, et al · 2017
Later among the works it cites.
Voice synthesis for in-the-wild speakers via a phonological loop
Taigman, Yaniv, Wolf, Lior, Polyak, Adam, and Nachmani, Eliya · 2017
Later among the works it cites.
Neural discrete representation learning
van den Oord, Aaron, Vinyals, Oriol, et al · 2017
Later among the works it cites.
Attention is all you need
Vaswani, Ashish, Shazeer, Noam, Parmar, Niki, Uszkoreit, Jakob, Jones, Llion, Gomez, Aidan N, Kaiser, Łukasz, and Polosukhin, Illia · 2017
Later among the works it cites.
Towards end-to-end prosody transfer for expressive speech synthesis with Tacotron
Skerry-Ryan, RJ, Battenberg, Eric, Xiao, Ying, Wang, Yuxuan, Stanton, Daisy, Shor, Joel, Weiss, Ron J., Clark, Rob, and Saurous, Rif A · 2018
Closest in time.