Fetching the paper…
Reading the bibliography…
Global Style Tokens (GSTs) are a recently-proposed method to learn latent disentangled representations of high-dimensional data.
D. J. Hirst, “La représentation linguistique des systèmes prosodiques: une approche cognitive,” Ph.D. dissertation, Aix-Marseille 1, 1987
1987
Earlier work this paper cites.
K. Silverman, M. Beckman, J. Pitrelli, M. Ostendorf, C. Wightman, P. Price, J. Pierrehumbert, and J. Hirschberg, “ToBI: A standard for labeling english prosody,” in Second International Conference on Spoken Language Processing , 1992
1992
Earlier work this paper cites.
D. Hirst and R. Espesser, “Automatic modelling of fundamental frequency using a quadratic spline function,” Travaux de l’Institut de Phonétique d’Aix , vol. 15, pp. 71–85, 1993. [Online]. Available: http://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.40.3623
1993
Earlier work this paper cites.
S. A. Liu, “Landmark detection for distinctive feature-based speech recognition,” The Journal of the Acoustical Society of America , vol. 100, no. 5, pp. 3417–3430, 1996
1996
Earlier work this paper cites.
P. Taylor, “The tilt intonation model,” in ICSLP . International Speech Communication Association, 1998
1998
Earlier work this paper cites.
A. Rosenberg, “AuToBI-a tool for automatic ToBI annotation.” in Interspeech , 2010, pp. 146–149. [Online]. Available: http://eniac.cs.qc.cuny.edu/andrew/autobi/
2010
Earlier work this paper cites.
N. Dehak, P. J. Kenny, R. Dehak, P. Dumouchel, and P. Ouellet, “Front-end factor analysis for speaker verification,” IEEE Transactions on Audio, Speech, and Language Processing , vol. 19, no. 4, pp. 788–798, 2011
2011
Earlier work this paper cites.
N. Obin, J. Beliao, C. Veaux, and A. Lacheret, “SLAM: Automatic stylization and labelling of speech melody,” in Speech Prosody , 2014, pp. 246–250
2014
Earlier work this paper cites.
K. Cho, B. van Merrienboer, Ã. Gulcehre, D. Bahdanau, F. Bougares, H. Schwenk, and Y. Bengio, “Learning phrase representations using RNN encoder-decoder for statistical machine translation.” ACL, 2014, pp. 1724–1734
2014
Earlier work this paper cites.
2014
Cited alongside, same era.
J. Lee and I. Tashev, “High-level feature representation using recurrent neural network for speech emotion recognition,” in Interspeech 2015 , September 2015
2015
Cited alongside, same era.
2015
Cited alongside, same era.
2015
Cited alongside, same era.
I. Jauk, “Unsupervised learning for expressive speech synthesis,” Ph.D. dissertation, Universitat Politècnica de Catalunya, 2017
2017
Later among the works it cites.
2017
Later among the works it cites.
2018
Closest in time.
J. Lorenzo-Trueba, G. E. Henter, S. Takaki, J. Yamagishi, Y. Morino, and Y. Ochiai, “Investigating different representations for modeling and controlling multiple emotions in DNN-based speech synthesis,” Speech Communication , vol. 99, pp. 135–143, 2018
2018
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2017
Cited alongside, same era.
Z.-Q. Wang and I. Tashev, “Learning utterance-level representations for speech emotion and age/gender recognition using deep neural networks,” in Acoustics, Speech and Signal Processing (ICASSP), 2017 IEEE International Conference on . IEEE, 2017, pp. 5150–5154
2017
Cited alongside, same era.
2017
Cited alongside, same era.
2017
Cited alongside, same era.
J. Deng, X. Xu, Z. Zhang, S. Fruhholz, and B. Schuller, “Semisupervised autoencoders for speech emotion recognition,” IEEE/ACM Transactions on Audio, Speech and Language Processing (TASLP) , vol. 26, no. 1, pp. 31–43, 2018
2018
Closest in time.
2018
Closest in time.
K. Akuzawa, Y. Iwasawa, and Y. Matsuo, “VAELoopDemo: audio samples generated by VAE-Loop,” https://akuzeee.github.io/VAELoopDemo/ , 2018
2018
Closest in time.
2018
Closest in time.