Fetching the paper…
Reading the bibliography…
Recognizing emotions in spoken communication is crucial for advanced human-machine interaction.
P. Shaver, J. Schwartz, D. Kirson, and C. O’connor, “Emotion knowledge: further exploration of a prototype approach.”
1987
Earlier work this paper cites.
F. Burkhardt, A. Paeschke, M. Rolfes, W. F. Sendlmeier, B. Weiss
2005
Earlier work this paper cites.
T. Wu, Y. Yang, Z. Wu, and D. Li, “Masc: A speech corpus in mandarin for emotion analysis and affective speaker recognition,” in
2006
Earlier work this paper cites.
C. Busso, M. Bulut, C.-C. Lee, A. Kazemzadeh, E. Mower, S. Kim, J. N. Chang, S. Lee, and S. S. Narayanan, “Iemocap: Interactive emotional dyadic motion capture database,”
2008
Earlier work this paper cites.
K. Dupuis and M. K. Pichora-Fuller, “Toronto emotional speech set (tess)-younger talker_happy,” 2010
2010
Earlier work this paper cites.
H. Cao, D. G. Cooper, M. K. Keutmann, R. C. Gur, A. Nenkova, and R. Verma, “Crema-d: Crowd-sourced emotional multimodal actors dataset,”
2014
Earlier work this paper cites.
P. Jackson and S. Haq, “Surrey audio-visual expressed emotion (savee) database,”
2014
Earlier work this paper cites.
D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,”
2014
Earlier work this paper cites.
Y. Ganin, E. Ustinova, H. Ajakan, P. Germain, H. Larochelle, F. Laviolette, M. Marchand, and V. Lempitsky, “Domain-adversarial training of neural networks,”
2016
Earlier work this paper cites.
R. Lotfian and C. Busso, “Building naturalistic emotionally balanced speech corpus by retrieving emotional speech from existing podcast recordings,”
2017
Earlier work this paper cites.
C. Busso, S. Parthasarathy, A. Burmania, M. Abdelwahab, N. Sadoughi, and E. M. Provost, “Msp-improv: An acted corpus of dyadic interactions to study emotion perception,”
2017
Cited alongside, same era.
I. Loshchilov and F. Hutter, “Decoupled weight decay regularization,”
2017
Cited alongside, same era.
A. Zadeh, P. P. Liang, S. Poria, P. Vij, E. Cambria, and L.-P. Morency, “Multi-attention recurrent network for human communication comprehension,” in
2018
Cited alongside, same era.
2018
Cited alongside, same era.
J. James, L. Tian, and C. I. Watson, “An open source emotional speech corpus for human robot interaction applications,” vol. 2018-September. International Speech Communication Association, 2018, pp. 2768–2772
D. S. Park, W. Chan, Y. Zhang, C.-C. Chiu, B. Zoph, E. D. Cubuk, and Q. V. Le, “Specaugment: A simple data augmentation method for automatic speech recognition,”
2019
Later among the works it cites.
L. N. Smith and N. Topin, “Super-convergence: Very fast training of neural networks using large learning rates,” in
2019
Later among the works it cites.
D. Tientcheu, T. Landry, Q. He, H. Yan, and Y. Li, “Asvp-esd: A dataset and its benchmark for emotion recognition using both speech and non-speech utterances,” 2020. [Online]. Available:
2020
Later among the works it cites.
2020
Later among the works it cites.
“Emotional voice conversion: Theory, databases and esd,”
2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2018
Cited alongside, same era.
S. R. Livingstone and F. A. Russo, “The ryerson audio-visual database of emotional speech and song (ravdess): A dynamic, multimodal set of facial and vocal expressions in north american english,” 2018. [Online]. Available:
2018
Cited alongside, same era.
N. Vryzas, R. Kotsakis, A. Liatsou, C. A. Dimoulas, and G. Kalliris, “Speech emotion recognition for performance interaction,”
2018
Cited alongside, same era.
S. Latif, A. Qayyum, M. Usman, and J. Qadir, “Cross lingual speech emotion recognition: Urdu vs. western languages,” in
2018
Cited alongside, same era.
J. Sager, R. Shankar, J. Reinhold, and A. Venkataraman, “Vesus: A crowd-annotated database to study emotion production and perception in spoken english,” vol. 2019-September. International Speech Communication Association, 2019, pp. 316–320
2019
Cited alongside, same era.
[Online]. Available:
Cited in the paper.
2022
Later among the works it cites.
M. Chen and Z. Yu, “Pre-finetuning for few-shot emotional speech recognition,”
2023
Closest in time.
A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever, “Robust speech recognition via large-scale weak supervision,” in
2023
Closest in time.