Fetching the paper…
Reading the bibliography…
We propose prosody embeddings for emotional and expressive speech synthesis networks.
“Signal estimation from modified short-time fourier transform,”
Daniel W. Griffin and Jae S. Lim, · 1983
Earlier work this paper cites.
“Tobi: a standard for labeling english prosody,”
Kim E. A. Silverman, Mary E. Beckman, John F. Pitrelli, Mari Ostendorf, Colin W. Wightman, Patti Price, Janet B. Pierrehumbert, and Julia Hirschberg, · 1992
Earlier work this paper cites.
“Tobi or not tobi?,”
Colin W Wightman, · 2002
Earlier work this paper cites.
“Learning phrase representations using rnn encoder–decoder for statistical machine translation,”
Kyunghyun Cho, Bart van Merrienboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio, · 2014
Earlier work this paper cites.
“Tacotron: Towards end-to-end speech synthesis,”
Yuxuan Wang, R.J. Skerry-Ryan, Daisy Stanton, Yonghui Wu, Ron J. Weiss, Navdeep Jaitly, Zongheng Yang, Ying Xiao, Zhifeng Chen, Samy Bengio, Quoc Le, Yannis Agiomyrgiannakis, Rob Clark, and Rif A. Saurous, · 2017
Earlier work this paper cites.
“Deep voice 2: Multi-speaker neural text-to-speech,”
Andrew Gibiansky, Sercan Arik, Gregory Diamos, John Miller, Kainan Peng, Wei Ping, Jonathan Raiman, and Yanqi Zhou, · 2017
Cited alongside, same era.
“Attention is all you need,”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin, · 2017
Cited alongside, same era.
“Fitting new speakers based on a short untranscribed sample,”
Eliya Nachmani, Adam Polyak, Yaniv Taigman, and Lior Wolf, · 2018
Cited alongside, same era.
“Towards end-to-end prosody transfer for expressive speech synthesis with tacotron,”
RJ Skerry-Ryan, Eric Battenberg, Ying Xiao, Yuxuan Wang, Daisy Stanton, Joel Shor, Ron Weiss, Rob Clark, and Rif A. Saurous, · 2018
Cited alongside, same era.
“Style tokens: Unsupervised style modeling, control and transfer in end-to-end speech synthesis,”
Yuxuan Wang, Daisy Stanton, Yu Zhang, RJ-Skerry Ryan, Eric Battenberg, Joel Shor, Ying Xiao, Ye Jia, Fei Ren, and Rif A. Saurous, · 2018
Closest in time.
“Predicting expressive speaking style from text in end-to-end speech synthesis,”
Daisy Stanton, Yuxuan Wang, and RJ Skerry-Ryan, · 2018
Closest in time.
“Natural TTS synthesis by conditioning wavenet on MEL spectrogram predictions,”
Jonathan Shen, Ruoming Pang, Ron J. Weiss, Mike Schuster, Navdeep Jaitly, Zongheng Yang, Zhifeng Chen, Yu Zhang, Yuxuan Wang, RJ-Skerrv Ryan, Rif A. Saurous, Yannis Agiomyrgiannakis, and Yonghui Wu, · 2018
Closest in time.
“An intriguing failing of convolutional neural networks and the coordconv solution,”
Rosanne Liu, Joel Lehman, Piero Molino, Felipe Petroski Such, Eric Frank, Alex Sergeev, and Jason Yosinski, · 2018
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…