Fetching the paper…
Reading the bibliography…
Emotional text-to-speech synthesis (ETTS) has seen much progress in recent years.
G. D and L. J. S, “Signal estimation from modified short-time fourier transform,”
1984
Earlier work this paper cites.
R. J. Williams, “Simple statistical gradient-following algorithms for connectionist reinforcement learning,”
1992
Earlier work this paper cites.
K. Tokuda, H. Zen, and A. W. Black, “An hmm-based speech synthesis system applied to english,” in
2002
Earlier work this paper cites.
J. Yamagishi, K. Onishi, T. Masuko, and T. Kobayashi, “Modeling of various speaking styles and emotions for hmm-based speech synthesis,” in
2003
Earlier work this paper cites.
M. El Ayadi, M. S. Kamel, and F. Karray, “Survey on speech emotion recognition: Features, classification schemes, and databases,”
2011
Earlier work this paper cites.
F. Eyben, S. Buchholz, N. Braunschweiler, J. Latorre, V. Wan, M. J. Gales, and K. Knill, “Unsupervised clustering of emotion and voice styles for expressive tts,” in
2012
Earlier work this paper cites.
A. Graves, “Generating sequences with recurrent neural networks,”
2013
Earlier work this paper cites.
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski
2015
Earlier work this paper cites.
Y. Wang, R. Skerry-Ryan, D. Stanton, Y. Wu, R. J. Weiss, N. Jaitly, Z. Yang, Y. Xiao, Z. Chen, S. Bengio
2017
Earlier work this paper cites.
J. Lorenzo-Trueba, G. E. Henter, S. Takaki, J. Yamagishi, Y. Morino, and Y. Ochiai, “Investigating different representations for modeling and controlling multiple emotions in dnn-based speech synthesis,”
2018
Earlier work this paper cites.
W.-N. Hsu, Y. Zhang, R. J. Weiss, H. Zen, Y. Wu, Y. Wang, Y. Cao, Y. Jia, Z. Chen, J. Shen
2018
Earlier work this paper cites.
R. S. Sutton and A. G. Barto,
2018
Cited alongside, same era.
Y. Koizumi, K. Niwa, Y. Hioka, K. Kobayashi, and Y. Haneda, “Dnn-based source enhancement to increase objective sound quality assessment score,”
2018
Cited alongside, same era.
T. Kala and T. Shinozaki, “Reinforcement learning of speech recognition system based on policy gradient and hypothesis selection,” in
2018
Cited alongside, same era.
J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. Skerrv-Ryan
2018
Cited alongside, same era.
Y. Wang, D. Stanton, Y. Zhang, R.-S. Ryan, E. Battenberg, J. Shor, Y. Xiao, Y. Jia, F. Ren, and R. A. Saurous, “Style tokens: Unsupervised style modeling, control and transfer in end-to-end speech synthesis,” in
2018
Cited alongside, same era.
O. Kwon, I. Jang, C. Ahn, and H.-G. Kang, “An effective style token weight control technique for end-to-end emotional speech synthesis,”
2019
Later among the works it cites.
S.-Y. Um, S. Oh, K. Byun, I. Jang, C. Ahn, and H.-G. Kang, “Emotional speech synthesis with rich and granularized control,” in
2020
Later among the works it cites.
M. Seurin, F. Strub, P. Preux, and O. Pietquin, “A machine of few words: Interactive speaker recognition with reinforcement learning,”
2020
Later among the works it cites.
R. Liu, B. Sisman, J. Li, F. Bao, G. Gao, and H. Li, “Teacher-student training for robust tacotron-based TTS,” in
2020
Later among the works it cites.
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
M. Chen, X. He, J. Yang, and H. Zhang, “3-d convolutional recurrent neural networks with attention model for speech emotion recognition,”
2018
Cited alongside, same era.
H. Choi, S. Park, J. Park, and M. Hahn, “Multi-speaker emotional acoustic modeling for cnn-based speech synthesis,” in
2019
Cited alongside, same era.
Y. Zhang, S. Pan, L. He, and Z. Ling, “Learning latent representations for style control and transfer in end-to-end speech synthesis,” in
2019
Cited alongside, same era.
P. Wu, Z. Ling, L. Liu, Y. Jiang, H. Wu, and L. Dai, “End-to-end emotional speech synthesis using style tokens and semi-supervised training,” in
2019
Cited alongside, same era.
Y. Shen, C. Huang, S. Wang, Y. Tsao, H. Wang, and T. Chi, “Reinforcement learning based speech enhancement for robust speech recognition,” in
2019
Cited alongside, same era.
F. Luo, P. Li, J. Zhou, P. Yang, B. Chang, Z. Sui, and X. Sun, “A dual reinforcement learning framework for unsupervised text style transfer,” in
2019
Cited alongside, same era.
Y. Zhou, X. Tian, and H. Li, “Multi-task wavernn with an integrated architecture for cross-lingual voice conversion,”
2020
Later among the works it cites.
X. Cai, D. Dai, Z. Wu, X. Li, J. Li, and H. Meng, “Emotion controllable speech synthesis using emotion-unlabeled dataset with the assistance of cross-domain speech emotion recognition,” in
2021
Closest in time.
T. Li, S. Yang, L. Xue, and L. Xie, “Controllable emotion transfer for end-to-end speech synthesis,” in
2021
Closest in time.
R. Liu, B. Sisman, G. Gao, and H. Li, “Expressive tts training with frame and style reconstruction loss,”
2021
Closest in time.
2021
Closest in time.
B. Sisman, S. King, J. Yamagishi, and H. Li, “An overview of voice conversion and its challenges: From statistical modeling to deep learning,”
2021
Closest in time.