Fetching the paper…
Reading the bibliography…
The increasing adoption of text-to-speech technologies has led to a growing demand for natural and emotive voices that adapt to a conversation's context and emotional tone.
1903
Earlier work this paper cites.
1912
Earlier work this paper cites.
P. Shaver, J. Schwartz, D. Kirson, and C. O’Connor, “Emotion knowledge: Further exploration of a prototype approach.” Journal of Personality and Social Psychology , vol. 52, no. 6, pp. 1061–1086, 1987. [Online]. Available: http://doi.apa.org/getdoi.cfm?doi=10.1037/0022-3514.52.6.1061
1987
Earlier work this paper cites.
T. Dalgleish and M. J. Power, Eds., Handbook of cognition and emotion . Chichester, England ; New York: Wiley, 1999
1999
Earlier work this paper cites.
J. Kominek and A. W. Black, “The CMU arctic speech databases,” in Fifth ISCA ITRW on Speech Synthesis, Pittsburgh, PA, USA, June 14-16, 2004 , A. W. Black and K. A. Lenzo, Eds. ISCA, 2004, pp. 223–224. [Online]. Available: http://www.isca-speech.org/archive_open/ssw5/ssw5_223.html
2004
Earlier work this paper cites.
2004
Earlier work this paper cites.
2006
Earlier work this paper cites.
T. Bänziger, M. Mortillaro, and K. R. Scherer, “Introducing the geneva multimodal expression corpus for experimental research on emotion perception.” Emotion , vol. 12, no. 5, pp. 1161–1179, 2012. [Online]. Available: https://doi.org/10.1037/a0025827
2012
Cited alongside, same era.
H. Cao, D. G. Cooper, M. K. Keutmann, R. C. Gur, A. Nenkova, and R. Verma, “Crema-d: Crowd-sourced emotional multimodal actors dataset,” IEEE Transactions on Affective Computing , vol. 5, no. 4, pp. 377–390, 2014
2014
Cited alongside, same era.
C. Busso, S. Parthasarathy, A. Burmania, M. AbdelWahab, N. Sadoughi, and E. M. Provost, “MSP-IMPROV: An acted corpus of dyadic interactions to study emotion perception,” IEEE Transactions on Affective Computing , vol. 8, no. 1, pp. 67–80, Jan. 2017. [Online]. Available: https://doi.org/10.1109/taffc.2016.2515617
2016
Cited alongside, same era.
K. Ito and L. Johnson, “The lj speech dataset,” https://keithito.com/LJ-Speech-Dataset/ , 2017
2017
2018
Later among the works it cites.
H. Zen, V. Dang, R. Clark, Y. Zhang, R. J. Weiss, Y. Jia, Z. Chen, and Y. Wu, “Libritts: A corpus derived from librispeech for text-to-speech,” Interspeech 2019 , 2019
2019
Later among the works it cites.
Y. Yan, X. Tan, B. Li, T. Qin, S. Zhao, Y. Shen, and T.-Y. Liu, “Adaspeech 2: Adaptive Text to Speech with Untranscribed Data,” in ICASSP 2021 - 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . Toronto, ON, Canada: IEEE, Jun. 2021, pp. 6613–6617. [Online]. Available: https://ieeexplore.ieee.org/document/9414872/
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
M. McAuliffe, M. Socolof, S. Mihuc, M. Wagner, and M. Sonderegger, “Montreal Forced Aligner: Trainable Text-Speech Alignment Using Kaldi,” in Proc. Interspeech 2017 , 2017, pp. 498–502
2017
Cited alongside, same era.
S. R. Livingstone and F. A. Russo, “The ryerson audio-visual database of emotional speech and song (RAVDESS): A dynamic, multimodal set of facial and vocal expressions in north american english,” PLOS ONE , vol. 13, no. 5, p. e0196391, May 2018. [Online]. Available: https://doi.org/10.1371/journal.pone.0196391
2018
Cited alongside, same era.
2021
Later among the works it cites.
2021
Later among the works it cites.
K. Zhou, B. Sisman, R. Liu, and H. Li, “Seen and unseen emotional style transfer for voice conversion with a new emotional speech dataset,” in ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2021, pp. 920–924
2021
Later among the works it cites.