Fetching the paper…
Reading the bibliography…
Emotional voice conversion (EVC) is one way to generate expressive synthetic speech.
“Signal estimation from modified short-time fourier transform,”
D. Griffin and J. Lim, · 1984
Earlier work this paper cites.
“Timit acoustic-phonetic continuous speech corpus, 1993,”
J. Garofolo, L. Lamel, W. Fisher, J. Fiscus, D. Pallett, N. Dahlgren, and V. Zue, · 1993
Earlier work this paper cites.
“Restructuring speech representations using a pitch-adaptive time–frequency smoothing and an instantaneous-frequency-based f0 extraction: Possible role of a repetitive structure in sounds1,”
H. Kawahara, I. Masuda-Katsuse, and A. De Cheveigne, · 1999
Earlier work this paper cites.
“Information, prosody, and modeling-with emphasis on tonal features of speech,”
H. Fujisaki, · 2004
Earlier work this paper cites.
“Prosody conversion from neutral speech to emotional speech,”
J. Tao, Y. Kang, and A. Li, · 2006
Earlier work this paper cites.
“A style control technique for hmm-based expressive speech synthesis,”
T. Nose, J. Yamagishi, T. Masuko, and T. Kobayashi, · 2007
Earlier work this paper cites.
“A system for transforming the emotion in speech: Combining data-driven conversion techniques for prosody and voice quality,”
Z. Inanoglu and S. Young, · 2007
Earlier work this paper cites.
“Multilevel parametric-base f0 model for speech synthesis,”
J. Latorre and M. Akamine, · 2008
Earlier work this paper cites.
“Hierarchical prosody conversion using regression-based clustering for emotional speech synthesis,”
C. Wu, C. Hsia, C. Lee, and M. Lin, · 2010
Earlier work this paper cites.
“Stylization and trajectory modelling of short and long term speech prosody variations,”
N. Obin, A. Lacheret, and X. Rodet, · 2011
Earlier work this paper cites.
“Improved prosody generation by maximizing joint probability of state and longer units,”
Y. Qian, Z. Wu, B. Gao, and F. Soong, · 2011
Earlier work this paper cites.
“Fundamental frequency modeling using wavelets for emotional voice conversion,”
H. Ming, D. Huang, M. Dong, H. Li, L. Xie, and S. Zhang, · 2015
Earlier work this paper cites.
“Deep bidirectional lstm modeling of timbre and prosody for emotional voice conversion,”
H. Ming, D. Huang, L. Xie, J. Wu, M. Dong, and H. Li, · 2016
Cited alongside, same era.
“Emotional voice conversion using neural networks with different temporal scales of f0 based on wavelet transform.,”
Z. Luo, T. Takiguchi, and Y. Ariki, · 2016
Cited alongside, same era.
“World: a vocoder-based high-quality speech synthesis system for real-time applications,”
M. Morise, F. Yokomori, and K. Ozawa, · 2016
Cited alongside, same era.
“Wavenet: A generative model for raw audio.,”
A. Van Den Oord, S. Dieleman, H. Zen, K. Simonyan, O. Vinyals, A. Graves, N. Kalchbrenner, A. Senior, and K. Kavukcuoglu, · 2016
Cited alongside, same era.
“Phonetic posteriorgrams for many-to-one voice conversion without parallel data training,”
L. Sun, K. Li, H. Wang, S. Kang, and H. Meng, · 2016
Cited alongside, same era.
“Towards end-to-end prosody transfer for expressive speech synthesis with tacotron,”
R. Skerry-Ryan, E. Battenberg, Y. Xiao, Y. Wang, D. Stanton, J. Shor, R. Weiss, R. Clark, and R. Saurous, · 2018
Later among the works it cites.
“Feature based adaptation for speaking style synthesis,”
X. Wu, L. Sun, S. Kang, S. Liu, Z. Wu, X. Liu, and H. Meng, · 2018
Later among the works it cites.
“Rapid style adaptation using residual error embedding for expressive speech synthesis,”
X. Wu, Y. Cao, M. Wang, S. Liu, S. Kang, Z. Wu, X. Liu, D. Su, D. Yu, and H. Meng, · 2018
Later among the works it cites.
“Natural tts synthesis by conditioning wavenet on mel spectrogram predictions,”
J. Shen, R. Pang, R. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. Skerrv-Ryan, et al., · 2018
Later among the works it cites.
“Waveglow: A flow-based generative network for speech synthesis,”
R. Prenger, R. Valle, and B. Catanzaro, · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
L. Dinh, J. Sohl-Dickstein, and S. Bengio, · 2016
Cited alongside, same era.
“Principles for learning controllable tts from annotated and latent variation,”
G. Henter, J. Lorenzo-Trueba, X. Wang, and J. Yamagishi, · 2017
Cited alongside, same era.
“Emotional voice conversion with adaptive scales f0 based on wavelet transform using limited amount of emotional data,”
Z. Luo, J. Chen, T. Takiguchi, and Y. Ariki, · 2017
Cited alongside, same era.
“Speaker-dependent wavenet vocoder,”
A. Tamamori, T. Hayashi, K. Kobayashi, K. Takeda, and T. Toda, · 2017
Cited alongside, same era.
“Investigating different representations for modeling and controlling multiple emotions in dnn-based speech synthesis,”
J. Lorenzo-Trueba, G. Henter, S. Takaki, J. Yamagishi, Y. Morino, and Y. Ochiai, · 2018
Cited alongside, same era.
“Style tokens: Unsupervised style modeling, control and transfer in end-to-end speech synthesis,”
Y. Wang, D. Stanton, Y. Zhang, R. Skerry-Ryan, E. Battenberg, J. Shor, Y. Xiao, F. Ren, Y. Jia, and R. Saurous, · 2018
Cited alongside, same era.
Later among the works it cites.
“Flowavenet: A generative flow for raw audio,”
S. Kim, S. Lee, J. Song, and S. Yoon, · 2018
Later among the works it cites.
“Voice conversion across arbitrary speakers based on a single target-speaker utterance,”
S. Liu, J. Zhong, L. Sun, X. Wu, X. Liu, and H. Meng, · 2018
Later among the works it cites.
“The hccl-cuhk system for the voice conversion challenge 2018,”
S. Liu, L. Sun, X. Wu, X. Liu, and H. Meng, · 2018
Later among the works it cites.
“Sequence-to-sequence acoustic modeling for voice conversion,”
J. Zhang, Z. Ling, L. Dai, L. Liu, and Y. Jiang, · 2018
Later among the works it cites.
“Glow: Generative flow with invertible 1x1 convolutions,”
D. Kingma and P. Dhariwal, · 2018
Later among the works it cites.
“Jointly trained conversion model and wavenet vocoder for non-parallel voice conversion using mel-spectrograms and phonetic posteriorgrams,”
S. Liu, Y. Cao, X. Wu, L. Sun, X. Liu, and H. Meng, · 2019
Later among the works it cites.