Fetching the paper…
Reading the bibliography…
While modern TTS technologies have made significant advancements in audio quality, there is still a lack of behavior naturalness compared to conversing with people.
K. R. Scherer, R. Banse, H. G. Wallbott, and T. Goldbeck, “Vocal cues in emotion encoding and decoding,”
1991
Earlier work this paper cites.
R. Banse and K. R. Scherer, “Acoustic profiles in vocal emotion expression.”
1996
Earlier work this paper cites.
Y. Gao, “Demo for ’interactive text-to-speech via semi-supervised style transfer learning’,”
2002
Earlier work this paper cites.
J. Yamagishi, K. Onishi, T. Masuko, and T. Kobayashi, “Modeling of various speaking styles and emotions for HMM-based speech synthesis,” in
2003
Earlier work this paper cites.
M. Tachibana, J. Yamagishi, K. Onishi, T. Masuko, and T. Kobayashi, “HMM-based speech synthesis with various speaking styles using model interpolation,” in
2004
Earlier work this paper cites.
J. Yamagishi, M. Tachibana, T. Masuko, and T. Kobayashi, “Speaking style adaptation using context clustering decision tree for HMM-based speech synthesis,” in
2004
Earlier work this paper cites.
C. Busso, M. Bulut, C.-C. Lee, A. Kazemzadeh, E. Mower, S. Kim, J. N. Chang, S. Lee, and S. S. Narayanan, “IEMOCAP: Interactive emotional dyadic motion capture database,”
2008
Earlier work this paper cites.
E. Mower, M. J. Mataric, and S. Narayanan, “A framework for automatic human emotion classification using emotion profiles,”
2010
Earlier work this paper cites.
C.-C. Lee, E. Mower, C. Busso, S. Lee, and S. Narayanan, “Emotion recognition using a hierarchical binary decision tree approach,”
2011
Earlier work this paper cites.
F. Eyben, F. Weninger, F. Gross, and B. Schuller, “Recent developments in opensmile, the munich open-source multimedia feature extractor,” in
2013
Earlier work this paper cites.
K. Han, D. Yu, and I. Tashev, “Speech emotion recognition using deep neural network and extreme learning machine,” in
2014
Cited alongside, same era.
D. Bahdanau, K. Cho, and Y. Bengio, “Neural machine translation by jointly learning to align and translate,”
2015
Cited alongside, same era.
2015
Cited alongside, same era.
2016
Cited alongside, same era.
G. Trigeorgis, F. Ringeval, R. Brueckner, E. Marchi, M. A. Nicolaou, B. Schuller, and S. Zafeiriou, “Adieu features? end-to-end speech emotion recognition using a deep convolutional recurrent network,” in
Z. Hodari, O. Watts, S. Ronanki, and S. King, “Learning interpretable control dimensions for speech synthesis by using external data.” in
2018
Later among the works it cites.
S. Yoon, S. Byun, and K. Jung, “Multimodal speech emotion recognition using audio and text,” in
2018
Later among the works it cites.
R. Skerry-Ryan, E. Battenberg, Y. Xiao, Y. Wang, D. Stanton, J. Shor, R. J. Weiss, R. Clark, and R. A. Saurous, “Towards end-to-end prosody transfer for expressive speech synthesis with Tacotron,”
2018
Later among the works it cites.
Y. Wang, D. Stanton, Y. Zhang, R. Skerry-Ryan, E. Battenberg, J. Shor, Y. Xiao, F. Ren, Y. Jia, and R. A. Saurous, “Style tokens: Unsupervised style modeling, control and transfer in end-to-end speech synthesis,”
2018
Later among the works it cites.
F. Locatello, S. Bauer, M. Lucic, S. Gelly, B. Schölkopf, and O. Bachem, “Challenging common assumptions in the unsupervised learning of disentangled representations,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2016
Cited alongside, same era.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in
2017
Cited alongside, same era.
2017
Cited alongside, same era.
G. E. Henter, J. Lorenzo-Trueba, X. Wang, and J. Yamagishi, “Principles for learning controllable TTS from annotated and latent variation.” in
2017
Cited alongside, same era.
S. King and V. Karaiskos, “The Blizzard challenge 2017,” in
2017
Cited alongside, same era.
J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. Skerrv-Ryan
2018
Cited alongside, same era.
2018
Later among the works it cites.
2018
Later among the works it cites.
Y. Li, N. Wang, J. Shi, X. Hou, and J. Liu, “Adaptive batch normalization for practical domain adaptation,”
2018
Later among the works it cites.
N. Tits, F. Wang, K. E. Haddad, V. Pagel, and T. Dutoit, “Visualization and interpretation of latent spaces for controlling expressive speech synthesis through audio analysis,” in
2019
Later among the works it cites.
A. Rabiee, T.-H. Kim, and S.-Y. Lee, “Adjusting pleasure-arousal-dominance for continuous emotional text-to-speech synthesizer,” in
2019
Later among the works it cites.